Can DeepSeek V4 Flash 0731 run on Mac Studio M5 Ultra 96GB?

YES — Runs, barely
IQ1_S · 4K context · formula estimate

What this means

DeepSeek V4 Flash 0731 fits on the Mac Studio M5 Ultra · 96GB at IQ1_S. We estimate it uses 87.6GB of the conservative 88GB working budget.

Estimated decode speed is 22.269.9 tokens/s. A roughly 300-word answer may take around 9 seconds.

22.2–69.9tokens/s
estimated · instant
87.6GB
needed at IQ1_S
IQ1_S
quant selected
0.4GB
working headroom
needs 87.6 GBworking budget 88 GB
Apple M5 Ultra96 GB unified memoryMetalActive cooling

See how fast it feels

Using the midpoint of our 22.269.9 tokens/s estimate for this demo.

Live demo · 46.1 tokens/s

Where the memory goes

ComponentDetailGB
Model weightsIQ1_S GGUF (82.5 GB) + mmap overhead86.6
KV cache4K context window0.2
RuntimemacOS inference app + compute buffers0.8
Total neededat IQ1_S, 4K context87.6
Working budget96 GB unified memory − conservative macOS reserve88
Headroomremaining inside the working budget0.4

Pick your quant

QuantDownloadMemoryEstimated speedVerdict
IQ1_S82.5 GB87.6 GB~22.2–69.9 tok/s! Runs, barely
IQ1_M86.9 GB92.2 GB Won't fit
IQ2_M90.9 GB96.4 GB Won't fit
IQ2_XXS90.9 GB96.4 GB Won't fit
Q2_K_XL96.8 GB102.6 GB Won't fit
IQ3_XXS104.2 GB110.4 GB Won't fit
IQ3_S116.1 GB122.9 GB Won't fit
Q3_K_M128.1 GB135.5 GB Won't fit
Q3_K_XL128.2 GB135.6 GB Won't fit
IQ4_NL136.7 GB144.5 GB Won't fit
IQ4_XS136.7 GB144.5 GB Won't fit
Q4_K_XL155.1 GB163.8 GB Won't fit
Q8_K_XL161.9 GB171 GB Won't fit

Get it running on this Mac

Recommended app: Unsloth Desktop. Follow the point-and-click steps below.

1
Install Unsloth Desktop for macOS.

Download it from the publisher's official macOS page. This is the graphical route documented alongside the quant files; the exact Mac/app pairing is still Guided until AICanRun completes its own offline check.

2
Search for the exact DeepSeek repository.

Use unsloth/DeepSeek-V4-Flash-0731-GGUF, then choose IQ1_S. The tracked download is about 82.5 GB.

3
Leave enough disk space and close memory-heavy apps.

Keep at least 91GB of free storage for the model and download overhead. This page estimates 87.6GB of working memory for the selected quant.

4
Load the model and start at 4K context.

Use a 4096-token context for the first run. If DSpark loading is unstable, use the publisher's documented latest llama.cpp route without a draft module before increasing context. Very long Think Max sessions need materially more memory than this 4K compatibility estimate.

5
Send a short prompt, then repeat it offline.

A reply with Wi-Fi disabled confirms local inference. The displayed 22.269.9 tokens/s range is the conservative ordinary-GGUF estimate. Custom DS4 and oMLX Q2 runtimes have reported materially higher results, but they use different artifacts and do not turn this into an AICanRun measurement of Unsloth Desktop.

The quant publisher documents Unsloth Desktop on macOS and recommends UD-IQ3_XXS for 128GB devices, with at least 110GB RAM. AICanRun has not completed an exact app-level offline test, so the path remains Guided and the displayed speed remains a calibrated formula estimate.

Why Unsloth Desktop?

The quant publisher documents this exact repository, recommends IQ3_XXS for a 128GB device, and provides a macOS app that can download and run it without a manual build.

Latest llama.cpp

The publisher-documented command-line route for the same Unsloth GGUF shards.

Use a current build with Metal support; this is a Terminal workflow, not a beginner app check.
DS4

A DeepSeek V4-specific Metal runtime with independent M4 Max 128GB reports.

DS4 uses its own custom Q2/Q4 artifacts, so its memory and speed do not describe the tracked IQ3_XXS file exactly.
If it does not work
  • Model not listed: paste the exact repository unsloth/DeepSeek-V4-Flash-0731-GGUF, not only the model nickname.
  • Download stalls: confirm there is enough free storage, reconnect to Wi-Fi, and restart the download inside the app.
  • Load fails or the app closes: quit memory-heavy apps. In Unsloth Desktop, confirm IQ1_S is the model currently loaded. Otherwise use a smaller model from this Mac's results.
  • No reply while offline: make sure the downloaded local model—not a remote or cloud model—is selected in the chat.
Did the first offline reply work?

Your anonymous feedback helps us prioritize which Mac and model paths to retest.

Other models on this Mac

DeepSeek V4 Flash 0731 on other Mac Studio M5 Ultra configurations

FAQ

Can the Mac Studio M5 Ultra · 96GB run DeepSeek V4 Flash 0731?

Yes at IQ1_S. We estimate about 87.6GB of working memory and 22.2–69.9 tokens/s at 4K context.

Which DeepSeek V4 Flash 0731 quant should I use on this Mac?

IQ1_S. It is a 82.5GB download and leaves about 0.4GB inside our conservative working budget.

Which app should I use for DeepSeek V4 Flash 0731 on this Mac?

Start with Unsloth Desktop. This page gives the complete point-and-click walkthrough.

Are these speeds measured on a Mac Studio M5 Ultra?

No. The displayed range remains a formula estimate for the tracked Unsloth IQ3 GGUF. Public M4 Max 128GB records range from about 8 tokens/s for stock llama.cpp on that IQ3 path to roughly 23–47 tokens/s for different, custom Q2 runtimes. Those are calibration evidence, not AICanRun verified benchmarks.