AI Models for Mac Studio M3 Ultra — What runs on 96GB
Specs checked against manufacturer and public documentation on . Results below are estimates, not measurements.
What runs on the Mac Studio M3 Ultra
All 66 models at the best tracked quant for this 96GB configuration. Select any row for the full report and step-by-step setup guide.
| Model | Params | Quant | Needs | Speed | Verdict |
|---|---|---|---|---|---|
| Laguna S 2.1 | 118B | Q4_K_M | 80.7GB | ~88.6–124 tokens/s | ✓ Runs great |
| Llama 3.3 70B | 70B | Q4_K_M | 50.3GB | ~9.6–13.5 tokens/s | ✓ Runs great |
| Qwen 3.5 35B-A3B | 35B | Q4_K_M | 26.4GB | ~217.2–304 tokens/s | ✓ Runs great |
| Qwen 3.6 35B-A3B | 35B | Q4_K_M | 24.1GB | ~216.2–302.6 tokens/s | ✓ Runs great |
| Ornith 1.0 35B-A3B | 35B | Q4_K_M | 25.5GB | ~225.4–315.5 tokens/s | ✓ Runs great |
| KAT-Coder V2.5 Dev | 35B | Q4_K_M | 25.7GB | ~223.2–312.5 tokens/s | ✓ Runs great |
| Laguna XS 2.1 | 33B | Q4_K_M | 24.4GB | ~221.9–310.7 tokens/s | ✓ Runs great |
| Qwen3 32B | 32.8B | Q4_K_M | 23.9GB | ~20.7–29 tokens/s | ✓ Runs great |
| Gemma 4 31B | 31B | Q4_K_M | 22.2GB | ~22.4–31.3 tokens/s | ✓ Runs great |
| Qwen3 30B A3B | 30.5B | Q4_K_M | 22.5GB | ~203.5–284.9 tokens/s | ✓ Runs great |
| Nemotron 3 Nano 30B-A3B | 30B | Q4_K_M | 28.7GB | ~166.5–233 tokens/s | ✓ Runs great |
| Salience 1.5 Flash | 30B | Q4_K_M | 22.5GB | ~199.1–278.7 tokens/s | ✓ Runs great |
| Qwen 3.5 27B | 27B | Q4_K_M | 20.2GB | ~24.5–34.3 tokens/s | ✓ Runs great |
| Qwen 3.6 27B | 27B | Q4_K_M | 18.7GB | ~24.4–34.1 tokens/s | ✓ Runs great |
| Bonsai 27B (1-bit) | 27B | Q1_0 | 5GB | ~107.8–150.9 tokens/s | ✓ Runs great |
| Ternary Bonsai 27B | 27B | Q2_0 | 8.6GB | ~56.9–79.6 tokens/s | ✓ Runs great |
| Fara 1.5 27B | 27B | Q4_K_M | 21.1GB | ~23.4–32.8 tokens/s | ✓ Runs great |
| Gemma 4 26B-A4B | 26B | Q4_0 | 17.7GB | ~184.8–258.8 tokens/s | ✓ Runs great |
| GPT-OSS 20B | 21B | MXFP4 | 15GB | ~197.4–276.4 tokens/s | ✓ Runs great |
| Qwen3 14B | 14.8B | Q4_K_M | 11.3GB | ~45.5–63.7 tokens/s | ✓ Runs great |
| Phi-4 14B | 14.7B | Q4_K_M | 11.4GB | ~45–63 tokens/s | ✓ Runs great |
| Ministral 3 14B | 14B | Q4_K_M | 10.4GB | ~49.9–69.9 tokens/s | ✓ Runs great |
| Gemma 3 12B | 12.2B | Q4_K_M | 9.3GB | ~56.1–78.5 tokens/s | ✓ Runs great |
| Gemma 4 12B | 12B | Q4_0 | 9GB | ~58.5–81.9 tokens/s | ✓ Runs great |
| Grug 12B | 12B | Q4_K_M | 9.7GB | ~53.2–74.5 tokens/s | ✓ Runs great |
| grug 9B (ProCreations) | 9.4B | Q4_K_M | 7.7GB | ~69.4–97.2 tokens/s | ✓ Runs great |
| Fara 1.5 9B | 9.4B | Q4_K_M | 7.7GB | ~69.4–97.2 tokens/s | ✓ Runs great |
| Qwen 3.5 9B | 9B | Q4_K_M | 7.4GB | ~71.8–100.6 tokens/s | ✓ Runs great |
| Ornith 1.0 9B | 9B | Q4_K_M | 7.3GB | ~73.1–102.4 tokens/s | ✓ Runs great |
| Qwythos 9B v2 | 9B | Q4_K_M | 7.4GB | ~71.8–100.6 tokens/s | ✓ Runs great |
| Tini Cybersec 8B-A1B | 8.5B | Q4_K_M | 6.9GB | ~669.4–937.1 tokens/s | ✓ Runs great |
| Qwen3 8B | 8.2B | Q4_K_M | 6.6GB | ~81.9–114.7 tokens/s | ✓ Runs great |
| Llama 3.1 8B | 8B | Q4_K_M | 6.5GB | ~83.6–117 tokens/s | ✓ Runs great |
| Ministral 8B | 8B | Q4_K_M | 6.5GB | ~83.6–117 tokens/s | ✓ Runs great |
| Ternary Bonsai 8B | 8B | Q2_0 | 3.7GB | ~186.1–260.6 tokens/s | ✓ Runs great |
| Ministral 3 8B | 8B | Q4_K_M | 6.8GB | ~78.8–110.2 tokens/s | ✓ Runs great |
| LFM2.5 8B-A1B | 8B | Q4_K_M | 6.8GB | ~630–882 tokens/s | ✓ Runs great |
| DeepSeek R1 Distill 7B | 7.6B | Q4_K_M | 6.3GB | ~87.1–122 tokens/s | ✓ Runs great |
| Mistral 7B v0.3 | 7.2B | Q4_K_M | 5.9GB | ~93.1–130.3 tokens/s | ✓ Runs great |
| AREX Turbo 4B | 4.5B | Q4_K_M | 4.2GB | ~141.2–197.7 tokens/s | ✓ Runs great |
| Fara 1.5 4B | 4.5B | Q4_K_M | 4.2GB | ~141.2–197.7 tokens/s | ✓ Runs great |
| Gemma 3 4B | 4.3B | Q4_K_M | 3.7GB | ~163.8–229.3 tokens/s | ✓ Runs great |
| Qwen3 4B | 4B | Q4_K_M | 3.7GB | ~163.8–229.3 tokens/s | ✓ Runs great |
| Qwen 3.5 4B | 4B | Q4_K_M | 3.9GB | ~151.7–212.3 tokens/s | ✓ Runs great |
| Gemma 4 E4B | 4B | Q4_K_M | 6.3GB | ~81.9–114.7 tokens/s | ✓ Runs great |
| Ternary Bonsai 4B | 4B | Q2_0 | 2.2GB | ~372.3–521.2 tokens/s | ✓ Runs great |
| Nemotron 3 Nano 4B | 4B | Q4_K_M | 4GB | ~146.3–204.8 tokens/s | ✓ Runs great |
| Agents-A1 4B | 4B | Q4_K_M | 3.9GB | ~151.7–212.3 tokens/s | ✓ Runs great |
| Phi-4 Mini 3.8B | 3.8B | Q4_K_M | 3.7GB | ~163.8–229.3 tokens/s | ✓ Runs great |
| Llama 3.2 3B | 3.2B | Q4_K_M | 3.1GB | ~204.8–286.7 tokens/s | ✓ Runs great |
| SmolLM3 3B | 3.1B | Q4_K_M | 3GB | ~215.5–301.7 tokens/s | ✓ Runs great |
| Ministral 3 3B | 3B | Q4_K_M | 3.2GB | ~195–273 tokens/s | ✓ Runs great |
| G9v3 3B | 3B | Q4_K_M | 3GB | ~215.5–301.7 tokens/s | ✓ Runs great |
| Qwen 3.5 2B | 2B | Q4_K_M | 2.3GB | ~315–441 tokens/s | ✓ Runs great |
| Gemma 4 E2B | 2B | Q4_K_M | 4.2GB | ~132.1–184.9 tokens/s | ✓ Runs great |
| DeepSeek R1 Distill 1.5B | 1.8B | Q4_K_M | 2.1GB | ~372.3–521.2 tokens/s | ✓ Runs great |
| Qwen3 1.7B | 1.7B | Q8_0 | 2.8GB | ~227.5–318.5 tokens/s | ✓ Runs great |
| SmolLM2 1.7B | 1.7B | Q4_K_M | 2.1GB | ~372.3–521.2 tokens/s | ✓ Runs great |
| Ternary Bonsai 1.7B | 1.7B | Q2_0 | 1.4GB | ~819–1146.6 tokens/s | ✓ Runs great |
| Llama 3.2 1B | 1.2B | Q4_K_M | 1.7GB | ~511.9–716.6 tokens/s | ✓ Runs great |
| Gemma 3 1B | 1B | Q4_K_M | 1.7GB | ~511.9–716.6 tokens/s | ✓ Runs great |
| OvisOCR2 0.8B | 0.8B | Q4_K_M | 1.5GB | ~682.5–955.5 tokens/s | ✓ Runs great |
| Qwen3 0.6B | 0.6B | Q8_0 | 1.5GB | ~682.5–955.5 tokens/s | ✓ Runs great |
| Inkling | 952.4B | Q8_0 | 967.1GB | — | ✕ Won't fit |
| Ornith 1.0 397B | 397B | Q4_K_M | 285.8GB | — | ✕ Won't fit |
| Hunyuan 3 (Hy3) | 298.8B | Q4_K_M | 213GB | — | ✕ Won't fit |
~ = bandwidth-based estimate · open a row to see the recommended app and exact steps
Other Mac Studio M3 Ultra memory options
FAQ
What is the biggest local AI model the Mac Studio M3 Ultra · 96GB can run?
Laguna S 2.1 is the largest model in our 66-model comparison that fits at Q4_K_M. It needs about 80.7GB inside our 88GB working-memory budget, with an estimated 88.6–124 tokens/s decode range.
How much of the 96GB unified memory is available to a local LLM?
We use a conservative 88GB working budget, leaving room for macOS, the inference app, and normal background activity. Memory pressure, context length, and other open apps can change the real limit.
Can the Mac Studio M3 Ultra · 96GB run Llama 3.1 8B?
Yes. At Q4_K_M, our estimate uses 6.5GB and lands around 83.6–117 tokens/s.
Can the Mac Studio M3 Ultra · 96GB run Qwen 3.6 27B?
Yes at Q4_K_M: about 18.7GB of working memory and an estimated 24.4–34.1 tokens/s.
Which app should I use for local AI on the Mac Studio M3 Ultra?
For an exact GGUF repository and quant from this site, start with LM Studio's graphical Discover, download, load, and chat flow. Jan is the open-source GGUF alternative; Ollama is strongest when the exact model already has a trustworthy catalog package or you want coding/API integrations; Msty is useful for mixed GGUF, MLX, and document workflows. For the curated Bonsai 27B build, try Locally AI and first confirm that its catalog shows the model on this Mac.
Can I upgrade the Mac Studio M3 Ultra to more unified memory later?
No. Apple-silicon unified memory is integrated into the chip package and must be chosen at purchase. Storage upgrades or external SSDs do not increase the memory an LLM can use.
Are these Mac Studio M3 Ultra local AI speeds measured?
No. Every speed on this page is a formula range based on Apple’s published memory bandwidth, model size, active parameters, and a cooling-aware efficiency range. App, backend, context length, and thermals can change real results.