Can DeepSeek V4 Flash 0731 run on Mac Studio M5 Ultra 96GB?
What this means
DeepSeek V4 Flash 0731 fits on the Mac Studio M5 Ultra · 96GB at IQ1_S. We estimate it uses 87.6GB of the conservative 88GB working budget.
Estimated decode speed is 22.2–69.9 tokens/s. A roughly 300-word answer may take around 9 seconds.
See how fast it feels
Using the midpoint of our 22.2–69.9 tokens/s estimate for this demo.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | IQ1_S GGUF (82.5 GB) + mmap overhead | 86.6 |
| KV cache | 4K context window | 0.2 |
| Runtime | macOS inference app + compute buffers | 0.8 |
| Total needed | at IQ1_S, 4K context | 87.6 |
| Working budget | 96 GB unified memory − conservative macOS reserve | 88 |
| Headroom | remaining inside the working budget | 0.4 |
Pick your quant
| Quant | Download | Memory | Estimated speed | Verdict |
|---|---|---|---|---|
| IQ1_S ★ | 82.5 GB | 87.6 GB | ~22.2–69.9 tok/s | ! Runs, barely |
| IQ1_M | 86.9 GB | 92.2 GB | — | ✕ Won't fit |
| IQ2_M | 90.9 GB | 96.4 GB | — | ✕ Won't fit |
| IQ2_XXS | 90.9 GB | 96.4 GB | — | ✕ Won't fit |
| Q2_K_XL | 96.8 GB | 102.6 GB | — | ✕ Won't fit |
| IQ3_XXS | 104.2 GB | 110.4 GB | — | ✕ Won't fit |
| IQ3_S | 116.1 GB | 122.9 GB | — | ✕ Won't fit |
| Q3_K_M | 128.1 GB | 135.5 GB | — | ✕ Won't fit |
| Q3_K_XL | 128.2 GB | 135.6 GB | — | ✕ Won't fit |
| IQ4_NL | 136.7 GB | 144.5 GB | — | ✕ Won't fit |
| IQ4_XS | 136.7 GB | 144.5 GB | — | ✕ Won't fit |
| Q4_K_XL | 155.1 GB | 163.8 GB | — | ✕ Won't fit |
| Q8_K_XL | 161.9 GB | 171 GB | — | ✕ Won't fit |
Get it running on this Mac
Recommended app: Unsloth Desktop. Follow the point-and-click steps below.
Download it from the publisher's official macOS page. This is the graphical route documented alongside the quant files; the exact Mac/app pairing is still Guided until AICanRun completes its own offline check.
Use unsloth/DeepSeek-V4-Flash-0731-GGUF, then choose IQ1_S. The tracked download is about 82.5 GB.
Keep at least 91GB of free storage for the model and download overhead. This page estimates 87.6GB of working memory for the selected quant.
Use a 4096-token context for the first run. If DSpark loading is unstable, use the publisher's documented latest llama.cpp route without a draft module before increasing context. Very long Think Max sessions need materially more memory than this 4K compatibility estimate.
A reply with Wi-Fi disabled confirms local inference. The displayed 22.2–69.9 tokens/s range is the conservative ordinary-GGUF estimate. Custom DS4 and oMLX Q2 runtimes have reported materially higher results, but they use different artifacts and do not turn this into an AICanRun measurement of Unsloth Desktop.
The quant publisher documents Unsloth Desktop on macOS and recommends UD-IQ3_XXS for 128GB devices, with at least 110GB RAM. AICanRun has not completed an exact app-level offline test, so the path remains Guided and the displayed speed remains a calibrated formula estimate.
The quant publisher documents this exact repository, recommends IQ3_XXS for a 128GB device, and provides a macOS app that can download and run it without a manual build.
The publisher-documented command-line route for the same Unsloth GGUF shards.
Use a current build with Metal support; this is a Terminal workflow, not a beginner app check.A DeepSeek V4-specific Metal runtime with independent M4 Max 128GB reports.
DS4 uses its own custom Q2/Q4 artifacts, so its memory and speed do not describe the tracked IQ3_XXS file exactly.- Model not listed: paste the exact repository unsloth/DeepSeek-V4-Flash-0731-GGUF, not only the model nickname.
- Download stalls: confirm there is enough free storage, reconnect to Wi-Fi, and restart the download inside the app.
- Load fails or the app closes: quit memory-heavy apps. In Unsloth Desktop, confirm IQ1_S is the model currently loaded. Otherwise use a smaller model from this Mac's results.
- No reply while offline: make sure the downloaded local model—not a remote or cloud model—is selected in the chat.
Your anonymous feedback helps us prioritize which Mac and model paths to retest.
Other models on this Mac
DeepSeek V4 Flash 0731 on other Mac Studio M5 Ultra configurations
FAQ
Can the Mac Studio M5 Ultra · 96GB run DeepSeek V4 Flash 0731?
Yes at IQ1_S. We estimate about 87.6GB of working memory and 22.2–69.9 tokens/s at 4K context.
Which DeepSeek V4 Flash 0731 quant should I use on this Mac?
IQ1_S. It is a 82.5GB download and leaves about 0.4GB inside our conservative working budget.
Which app should I use for DeepSeek V4 Flash 0731 on this Mac?
Start with Unsloth Desktop. This page gives the complete point-and-click walkthrough.
Are these speeds measured on a Mac Studio M5 Ultra?
No. The displayed range remains a formula estimate for the tracked Unsloth IQ3 GGUF. Public M4 Max 128GB records range from about 8 tokens/s for stock llama.cpp on that IQ3 path to roughly 23–47 tokens/s for different, custom Q2 runtimes. Those are calibration evidence, not AICanRun verified benchmarks.