AI Models for Mac Studio M5 Ultra — What runs on 512GB

Specs checked against manufacturer and public documentation on . The 512GB unified-memory configuration is scheduled to become available in late October 2026.. Results below are estimates, not measurements.
DEEPSEEK V4 FLASH 0731
Fits at Q4_K_XL
This 512GB configuration uses the 155.1GB Q4_K_XL build. Estimated working memory is 163.8GB, with a formula-based 11.8–37.2 tokens/s range.
Open the exact DeepSeek report and setup guide →What runs on the Mac Studio M5 Ultra
All 89 models use the recommended quant when it fits, or the largest smaller tracked fallback for this 512GB configuration. Select any row for the full report and step-by-step setup guide.
| Model | Params | Quant | Needs | Speed | Verdict |
|---|---|---|---|---|---|
| Qwen3.8 2.4T-A95B | 2400B | Q1_0smaller version | 418.3GB | ~38.2–53.4 tokens/s | ✓ Runs great |
| Inkling | 952.4B | IQ1_Ssmaller version | 351.2GB | ~2.2–3.1 tokens/s | ✓ Runs great |
| LongCat Flash Chat | 561.9B | IQ1_S | 159.8GB | ~109.5–153.3 tokens/s | ✓ Runs great |
| Ornith 1.0 397B | 397B | Q4_K_M | 285.8GB | ~2.4–3.4 tokens/s | ✓ Runs great |
| Hunyuan 3 (Hy3) | 298.8B | Q4_K_M | 213GB | ~3.3–4.6 tokens/s | ✓ Runs great |
| DeepSeek V4 Flash 0731 | 284B | Q4_K_XL | 163.8GB | ~11.8–37.2 tokens/s | ✓ Runs great |
| Inkling Small | 276B | Q4_K_M | 190.7GB | ~84.9–118.9 tokens/s | ✓ Runs great |
| Ling 3.0 Flash | 124B | Q4_K_M | 82.5GB | ~187.5–262.5 tokens/s | ✓ Runs great |
| Laguna S 2.1 | 118B | Q4_K_M | 109.9GB | ~92.2–129.1 tokens/s | ✓ Runs great |
| GLM-4.5V | 106B | Q4_K_M | 75GB | ~83.3–116.7 tokens/s | ✓ Runs great |
| GLM-4.5-Air | 106B | Q4_K_M | 75GB | ~83.3–116.7 tokens/s | ✓ Runs great |
| Llama 3.3 70B | 70B | Q4_K_M | 50.3GB | ~14.1–19.8 tokens/s | ✓ Runs great |
| BigBang V1 36B-A3B | 36B | Q4_K_M | 26.3GB | ~328.8–460.3 tokens/s | ✓ Runs great |
| Qwen 3.5 35B-A3B | 35B | Q4_K_M | 26.4GB | ~318.2–445.5 tokens/s | ✓ Runs great |
| Qwen 3.6 35B-A3B | 35B | Q4_K_M | 24.1GB | ~316.7–443.4 tokens/s | ✓ Runs great |
| Ornith 1.0 35B-A3B | 35B | Q4_K_M | 25.5GB | ~330.2–462.3 tokens/s | ✓ Runs great |
| XYZ-Aquila mini 35B-A3B | 35B | Q4_K_M | 25.7GB | ~327.1–457.9 tokens/s | ✓ Runs great |
| KAT-Coder V2.5 Dev | 35B | Q4_K_M | 25.7GB | ~327.1–457.9 tokens/s | ✓ Runs great |
| Laguna XS 2.1 | 33B | Q4_K_M | 24.4GB | ~325.1–455.2 tokens/s | ✓ Runs great |
| Qwen3 32B | 32.8B | Q4_K_M | 23.9GB | ~30.3–42.4 tokens/s | ✓ Runs great |
| Nemotron 3 Nano Omni 30B-A3B | 31B | Q4_K_M | 26.5GB | ~276.8–387.5 tokens/s | ✓ Runs great |
| Gemma 4 31B | 31B | Q4_K_M | 22.2GB | ~32.8–45.9 tokens/s | ✓ Runs great |
| Qwen3 30B A3B | 30.5B | Q4_K_M | 22.5GB | ~298.1–417.4 tokens/s | ✓ Runs great |
| Nemotron 3 Nano 30B-A3B | 30B | Q4_K_M | 28.7GB | ~243.9–341.5 tokens/s | ✓ Runs great |
| Nemotron 3.5 Lightning 30B-A3B | 30B | Q8_0 | 36.1GB | ~178.6–250 tokens/s | ✓ Runs great |
| Salience 1.5 Flash | 30B | Q4_K_M | 22.5GB | ~291.7–408.4 tokens/s | ✓ Runs great |
| Granite 4.2 30B | 30B | Q4_K_M | 21.5GB | ~33.9–47.5 tokens/s | ✓ Runs great |
| Muse Glimmer 30B | 29.6B | K_QUANT_17GB | 20.5GB | ~35.7–50 tokens/s | ✓ Runs great |
| Qwen 3.5 27B | 27B | Q4_K_M | 20.2GB | ~35.9–50.3 tokens/s | ✓ Runs great |
| Qwen 3.6 27B | 27B | Q4_K_M | 18.7GB | ~35.7–50 tokens/s | ✓ Runs great |
| Bonsai 27B (1-bit) | 27B | Q1_0 | 5GB | ~157.9–221.1 tokens/s | ✓ Runs great |
| Ternary Bonsai 27B | 27B | Q2_0 | 8.6GB | ~83.3–116.7 tokens/s | ✓ Runs great |
| Fara 1.5 27B | 27B | Q4_K_M | 21.1GB | ~34.3–48 tokens/s | ✓ Runs great |
| Qwen3.8 27B | 27B | Q4_K_M | 21GB | ~31.6–44.2 tokens/s | ✓ Runs great |
| UI-Mate 27B | 27B | Q4_K_M | 21.1GB | ~34.3–48 tokens/s | ✓ Runs great |
| Gemma 4 26B-A4B | 26B | Q4_0 | 17.7GB | ~270.8–379.2 tokens/s | ✓ Runs great |
| GPT-OSS 20B | 21B | MXFP4 | 15GB | ~289.3–405 tokens/s | ✓ Runs great |
| Qwen3 14B | 14.8B | Q4_K_M | 11.3GB | ~66.7–93.3 tokens/s | ✓ Runs great |
| Phi-4 14B | 14.7B | Q4_K_M | 11.4GB | ~65.9–92.3 tokens/s | ✓ Runs great |
| Ministral 3 14B | 14B | Q4_K_M | 10.4GB | ~73.2–102.4 tokens/s | ✓ Runs great |
| Gemma 3 12B | 12.2B | Q4_K_M | 9.3GB | ~82.2–115.1 tokens/s | ✓ Runs great |
| Gemma 4 12B | 12B | Q4_0 | 9GB | ~85.7–120 tokens/s | ✓ Runs great |
| Grug 12B | 12B | Q4_K_M | 9.7GB | ~77.9–109.1 tokens/s | ✓ Runs great |
| GRM 3.2 Cliff 9B | 9.4B | Q4_K_M | 7.7GB | ~101.7–142.4 tokens/s | ✓ Runs great |
| grug 9B (ProCreations) | 9.4B | Q4_K_M | 7.7GB | ~101.7–142.4 tokens/s | ✓ Runs great |
| Fara 1.5 9B | 9.4B | Q4_K_M | 7.7GB | ~101.7–142.4 tokens/s | ✓ Runs great |
| Qwen 3.5 9B | 9B | Q4_K_M | 7.4GB | ~105.3–147.4 tokens/s | ✓ Runs great |
| Ornith 1.0 9B | 9B | Q4_K_M | 7.3GB | ~107.1–150 tokens/s | ✓ Runs great |
| Qwythos 9B v2 | 9B | Q4_K_M | 7.4GB | ~105.3–147.4 tokens/s | ✓ Runs great |
| UI-Mate 9B | 9B | Q4_K_M | 7.6GB | ~101.7–142.4 tokens/s | ✓ Runs great |
| Tini Cybersec 8B-A1B | 8.5B | Q4_K_M | 6.9GB | ~980.8–1373.1 tokens/s | ✓ Runs great |
| Qwen3 8B | 8.2B | Q4_K_M | 6.6GB | ~120–168 tokens/s | ✓ Runs great |
| Llama 3.1 8B | 8B | Q4_K_M | 6.5GB | ~122.4–171.4 tokens/s | ✓ Runs great |
| Ministral 8B | 8B | Q4_K_M | 6.5GB | ~122.4–171.4 tokens/s | ✓ Runs great |
| Ternary Bonsai 8B | 8B | Q2_0 | 3.7GB | ~272.7–381.8 tokens/s | ✓ Runs great |
| Ministral 3 8B | 8B | Q4_K_M | 6.8GB | ~115.4–161.5 tokens/s | ✓ Runs great |
| LFM2.5 8B-A1B | 8B | Q4_K_M | 6.8GB | ~923.1–1292.3 tokens/s | ✓ Runs great |
| Granite 4.2 8B | 8B | Q4_K_M | 6.9GB | ~113.2–158.5 tokens/s | ✓ Runs great |
| Ling 3.0 Tiny | 7.9B | Q4_K_M | 6.5GB | ~691–967.3 tokens/s | ✓ Runs great |
| DeepSeek R1 Distill 7B | 7.6B | Q4_K_M | 6.3GB | ~127.7–178.7 tokens/s | ✓ Runs great |
| Mistral 7B v0.3 | 7.2B | Q4_K_M | 5.9GB | ~136.4–190.9 tokens/s | ✓ Runs great |
| AREX Turbo 4B | 4.5B | Q4_K_M | 4.2GB | ~206.9–289.7 tokens/s | ✓ Runs great |
| Fara 1.5 4B | 4.5B | Q4_K_M | 4.2GB | ~206.9–289.7 tokens/s | ✓ Runs great |
| Gemma 3 4B | 4.3B | Q4_K_M | 3.7GB | ~240–336 tokens/s | ✓ Runs great |
| Nanbeige 4.2 3B | 4.2B | Q4_K_M | 3.9GB | ~222.2–311.1 tokens/s | ✓ Runs great |
| Qwen3 4B | 4B | Q4_K_M | 3.7GB | ~240–336 tokens/s | ✓ Runs great |
| Qwen 3.5 4B | 4B | Q4_K_M | 3.9GB | ~222.2–311.1 tokens/s | ✓ Runs great |
| Gemma 4 E4B | 4B | Q4_K_M | 6.3GB | ~120–168 tokens/s | ✓ Runs great |
| Ternary Bonsai 4B | 4B | Q2_0 | 2.2GB | ~545.5–763.6 tokens/s | ✓ Runs great |
| Nemotron 3 Nano 4B | 4B | Q4_K_M | 4GB | ~214.3–300 tokens/s | ✓ Runs great |
| Agents-A1 4B | 4B | Q4_K_M | 3.9GB | ~222.2–311.1 tokens/s | ✓ Runs great |
| Phi-4 Mini 3.8B | 3.8B | Q4_K_M | 3.7GB | ~240–336 tokens/s | ✓ Runs great |
| Llama 3.2 3B | 3.2B | Q4_K_M | 3.1GB | ~300–420 tokens/s | ✓ Runs great |
| SmolLM3 3B | 3.1B | Q4_K_M | 3GB | ~315.8–442.1 tokens/s | ✓ Runs great |
| LFM2.5-VL 3B | 3.1B | Q4_K_M | 2.8GB | ~352.9–494.1 tokens/s | ✓ Runs great |
| Ministral 3 3B | 3B | Q4_K_M | 3.2GB | ~285.7–400 tokens/s | ✓ Runs great |
| G9v3 3B | 3B | Q4_K_M | 3GB | ~315.8–442.1 tokens/s | ✓ Runs great |
| Granite 4.2 3B | 3B | Q4_K_M | 3.3GB | ~272.7–381.8 tokens/s | ✓ Runs great |
| LFM2.5 2.6B | 2.7B | Q4_K_M | 2.8GB | ~352.9–494.1 tokens/s | ✓ Runs great |
| Qwen 3.5 2B | 2B | Q4_K_M | 2.3GB | ~461.5–646.2 tokens/s | ✓ Runs great |
| Gemma 4 E2B | 2B | Q4_K_M | 4.2GB | ~193.5–271 tokens/s | ✓ Runs great |
| DeepSeek R1 Distill 1.5B | 1.8B | Q4_K_M | 2.1GB | ~545.5–763.6 tokens/s | ✓ Runs great |
| Qwen3 1.7B | 1.7B | Q8_0 | 2.8GB | ~333.3–466.7 tokens/s | ✓ Runs great |
| SmolLM2 1.7B | 1.7B | Q4_K_M | 2.1GB | ~545.5–763.6 tokens/s | ✓ Runs great |
| Ternary Bonsai 1.7B | 1.7B | Q2_0 | 1.4GB | ~1200–1680 tokens/s | ✓ Runs great |
| Llama 3.2 1B | 1.2B | Q4_K_M | 1.7GB | ~750–1050 tokens/s | ✓ Runs great |
| Gemma 3 1B | 1B | Q4_K_M | 1.7GB | ~750–1050 tokens/s | ✓ Runs great |
| OvisOCR2 0.8B | 0.8B | Q4_K_M | 1.5GB | ~1000–1400 tokens/s | ✓ Runs great |
| Qwen3 0.6B | 0.6B | Q8_0 | 1.5GB | ~1000–1400 tokens/s | ✓ Runs great |
~ = bandwidth-based estimate · open a row to see the recommended app and exact steps
Other Mac Studio M5 Ultra memory options
FAQ
What is the biggest local AI model the Mac Studio M5 Ultra · 512GB can run?
Qwen3.8 2.4T-A95B is the largest model in our 89-model comparison that fits at Q1_0. It needs about 418.3GB inside our 504GB working-memory budget, with an estimated 38.2–53.4 tokens/s decode range.
How much of the 512GB unified memory is available to a local LLM?
We use a conservative 504GB working budget, leaving room for macOS, the inference app, and normal background activity. Memory pressure, context length, and other open apps can change the real limit.
Can the Mac Studio M5 Ultra · 512GB run Llama 3.1 8B?
Yes. At Q4_K_M, our estimate uses 6.5GB and lands around 122.4–171.4 tokens/s.
Can the Mac Studio M5 Ultra · 512GB run Qwen 3.6 27B?
Yes at Q4_K_M: about 18.7GB of working memory and an estimated 35.7–50 tokens/s.
Can the Mac Studio M5 Ultra · 512GB run DeepSeek V4 Flash 0731?
Yes. Our per-Mac selector uses Q4_K_XL: about 163.8GB of working memory and an estimated 11.8–37.2 tokens/s. This is a formula estimate for the listed GGUF, not a benchmark.
Which app should I use for local AI on the Mac Studio M5 Ultra?
For most exact GGUF repositories and quants on this site, start with LM Studio's graphical Discover, download, load, and chat flow. Jan is the open-source GGUF alternative; Ollama is strongest when the exact model already has a trustworthy catalog package or you want coding/API integrations; Msty is useful for mixed GGUF, MLX, and document workflows. DeepSeek V4 Flash uses a separate Unsloth Desktop path because of its sharded GGUF packaging. For the curated Bonsai 27B build, try Locally AI and first confirm that its catalog shows the model on this Mac.
Can I upgrade the Mac Studio M5 Ultra to more unified memory later?
No. Apple-silicon unified memory is integrated into the chip package and must be chosen at purchase. Storage upgrades or external SSDs do not increase the memory an LLM can use.
Are these Mac Studio M5 Ultra local AI speeds measured?
No. Every speed on this page is a formula range based on Apple’s published memory bandwidth, model size, active parameters, and a cooling-aware efficiency range. App, backend, context length, and thermals can change real results.