AI Models for MacBook Neo A18 Pro — What runs on 8GB
Specs checked against manufacturer and public documentation on . Apple does not publish a MacBook Neo memory-bandwidth figure; the speed range uses the A18 Pro chip profile and should be treated as lower-confidence.. Results below are estimates, not measurements.
What runs on the MacBook Neo A18 Pro
All 66 models at the best tracked quant for this 8GB configuration. Select any row for the full report and step-by-step setup guide.
| Model | Params | Quant | Needs | Speed | Verdict |
|---|---|---|---|---|---|
| Ternary Bonsai 4B | 4B | Q2_0 | 2.2GB | ~22.9–31.6 tokens/s | ✓ Runs great |
| Llama 3.2 3B | 3.2B | Q4_K_M | 3.1GB | ~12.6–17.4 tokens/s | ✓ Runs great |
| SmolLM3 3B | 3.1B | Q4_K_M | 3GB | ~13.3–18.3 tokens/s | ✓ Runs great |
| Ministral 3 3B | 3B | Q4_K_M | 3.2GB | ~12–16.6 tokens/s | ✓ Runs great |
| G9v3 3B | 3B | Q4_K_M | 3GB | ~13.3–18.3 tokens/s | ✓ Runs great |
| Qwen 3.5 2B | 2B | Q4_K_M | 2.3GB | ~19.4–26.8 tokens/s | ✓ Runs great |
| DeepSeek R1 Distill 1.5B | 1.8B | Q4_K_M | 2.1GB | ~22.9–31.6 tokens/s | ✓ Runs great |
| Qwen3 1.7B | 1.7B | Q8_0 | 2.8GB | ~14–19.3 tokens/s | ✓ Runs great |
| SmolLM2 1.7B | 1.7B | Q4_K_M | 2.1GB | ~22.9–31.6 tokens/s | ✓ Runs great |
| Ternary Bonsai 1.7B | 1.7B | Q2_0 | 1.4GB | ~50.4–69.6 tokens/s | ✓ Runs great |
| Llama 3.2 1B | 1.2B | Q4_K_M | 1.7GB | ~31.5–43.5 tokens/s | ✓ Runs great |
| Gemma 3 1B | 1B | Q4_K_M | 1.7GB | ~31.5–43.5 tokens/s | ✓ Runs great |
| OvisOCR2 0.8B | 0.8B | Q4_K_M | 1.5GB | ~42–58 tokens/s | ✓ Runs great |
| Qwen3 0.6B | 0.6B | Q8_0 | 1.5GB | ~42–58 tokens/s | ✓ Runs great |
| Bonsai 27B (1-bit) | 27B | Q1_0 | 5GB | ~6.6–9.2 tokens/s | ! Runs, barely |
| Gemma 3 12B | 12.2B | IQ1_Ssmaller version | 4.9GB | ~8.1–11.2 tokens/s | ! Runs, barely |
| Tini Cybersec 8B-A1B | 8.5B | Q2_Ksmaller version | 4.7GB | ~69.1–95.4 tokens/s | ! Runs, barely |
| Llama 3.1 8B | 8B | Q2_Ksmaller version | 4.7GB | ~7.9–10.9 tokens/s | ! Runs, barely |
| Ministral 8B | 8B | Q2_Ksmaller version | 4.7GB | ~7.9–10.9 tokens/s | ! Runs, barely |
| Ternary Bonsai 8B | 8B | Q2_0 | 3.7GB | ~11.5–15.8 tokens/s | ! Runs, barely |
| DeepSeek R1 Distill 7B | 7.6B | Q2_Ksmaller version | 4.5GB | ~8.4–11.6 tokens/s | ! Runs, barely |
| Mistral 7B v0.3 | 7.2B | Q3_K_Msmaller version | 5GB | ~7.2–9.9 tokens/s | ! Runs, barely |
| AREX Turbo 4B | 4.5B | Q4_K_M | 4.2GB | ~8.7–12 tokens/s | ! Runs, barely |
| Fara 1.5 4B | 4.5B | Q4_K_M | 4.2GB | ~8.7–12 tokens/s | ! Runs, barely |
| Gemma 3 4B | 4.3B | Q4_K_M | 3.7GB | ~10.1–13.9 tokens/s | ! Runs, barely |
| Qwen3 4B | 4B | Q4_K_M | 3.7GB | ~10.1–13.9 tokens/s | ! Runs, barely |
| Qwen 3.5 4B | 4B | Q4_K_M | 3.9GB | ~9.3–12.9 tokens/s | ! Runs, barely |
| Nemotron 3 Nano 4B | 4B | Q4_K_M | 4GB | ~9–12.4 tokens/s | ! Runs, barely |
| Agents-A1 4B | 4B | Q4_K_M | 3.9GB | ~9.3–12.9 tokens/s | ! Runs, barely |
| Phi-4 Mini 3.8B | 3.8B | Q4_K_M | 3.7GB | ~10.1–13.9 tokens/s | ! Runs, barely |
| Gemma 4 E2B | 2B | Q4_K_M | 4.2GB | ~8.1–11.2 tokens/s | ! Runs, barely |
| Inkling | 952.4B | Q8_0 | 967.1GB | — | ✕ Won't fit |
| Ornith 1.0 397B | 397B | Q4_K_M | 285.8GB | — | ✕ Won't fit |
| Hunyuan 3 (Hy3) | 298.8B | Q4_K_M | 213GB | — | ✕ Won't fit |
| Laguna S 2.1 | 118B | Q4_K_M | 80.7GB | — | ✕ Won't fit |
| Llama 3.3 70B | 70B | Q4_K_M | 50.3GB | — | ✕ Won't fit |
| Qwen 3.5 35B-A3B | 35B | Q4_K_M | 26.4GB | — | ✕ Won't fit |
| Qwen 3.6 35B-A3B | 35B | Q4_K_M | 24.1GB | — | ✕ Won't fit |
| Ornith 1.0 35B-A3B | 35B | Q4_K_M | 25.5GB | — | ✕ Won't fit |
| KAT-Coder V2.5 Dev | 35B | Q4_K_M | 25.7GB | — | ✕ Won't fit |
| Laguna XS 2.1 | 33B | Q4_K_M | 24.4GB | — | ✕ Won't fit |
| Qwen3 32B | 32.8B | Q4_K_M | 23.9GB | — | ✕ Won't fit |
| Gemma 4 31B | 31B | Q4_K_M | 22.2GB | — | ✕ Won't fit |
| Qwen3 30B A3B | 30.5B | Q4_K_M | 22.5GB | — | ✕ Won't fit |
| Nemotron 3 Nano 30B-A3B | 30B | Q4_K_M | 28.7GB | — | ✕ Won't fit |
| Salience 1.5 Flash | 30B | Q4_K_M | 22.5GB | — | ✕ Won't fit |
| Qwen 3.5 27B | 27B | Q4_K_M | 20.2GB | — | ✕ Won't fit |
| Qwen 3.6 27B | 27B | Q4_K_M | 18.7GB | — | ✕ Won't fit |
| Ternary Bonsai 27B | 27B | Q2_0 | 8.6GB | — | ✕ Won't fit |
| Fara 1.5 27B | 27B | Q4_K_M | 21.1GB | — | ✕ Won't fit |
| Gemma 4 26B-A4B | 26B | Q4_0 | 17.7GB | — | ✕ Won't fit |
| GPT-OSS 20B | 21B | MXFP4 | 15GB | — | ✕ Won't fit |
| Qwen3 14B | 14.8B | Q4_K_M | 11.3GB | — | ✕ Won't fit |
| Phi-4 14B | 14.7B | Q4_K_M | 11.4GB | — | ✕ Won't fit |
| Ministral 3 14B | 14B | Q4_K_M | 10.4GB | — | ✕ Won't fit |
| Gemma 4 12B | 12B | Q4_0 | 9GB | — | ✕ Won't fit |
| Grug 12B | 12B | Q4_K_M | 9.7GB | — | ✕ Won't fit |
| grug 9B (ProCreations) | 9.4B | Q4_K_M | 7.7GB | — | ✕ Won't fit |
| Fara 1.5 9B | 9.4B | Q4_K_M | 7.7GB | — | ✕ Won't fit |
| Qwen 3.5 9B | 9B | Q4_K_M | 7.4GB | — | ✕ Won't fit |
| Ornith 1.0 9B | 9B | Q4_K_M | 7.3GB | — | ✕ Won't fit |
| Qwythos 9B v2 | 9B | Q4_K_M | 7.4GB | — | ✕ Won't fit |
| Qwen3 8B | 8.2B | Q4_K_M | 6.6GB | — | ✕ Won't fit |
| Ministral 3 8B | 8B | Q4_K_M | 6.8GB | — | ✕ Won't fit |
| LFM2.5 8B-A1B | 8B | Q4_K_M | 6.8GB | — | ✕ Won't fit |
| Gemma 4 E4B | 4B | Q4_K_M | 6.3GB | — | ✕ Won't fit |
~ = bandwidth-based estimate · open a row to see the recommended app and exact steps
FAQ
What is the biggest local AI model the MacBook Neo A18 Pro · 8GB can run?
Bonsai 27B (1-bit) is the largest model in our 66-model comparison that fits at Q1_0. It needs about 5GB inside our 5GB working-memory budget, with an estimated 6.6–9.2 tokens/s decode range.
How much of the 8GB unified memory is available to a local LLM?
We use a conservative 5GB working budget, leaving room for macOS, the inference app, and normal background activity. Memory pressure, context length, and other open apps can change the real limit.
Can the MacBook Neo A18 Pro · 8GB run Llama 3.1 8B?
Yes. At Q2_K, our estimate uses 4.7GB and lands around 7.9–10.9 tokens/s.
Can the MacBook Neo A18 Pro · 8GB run Qwen 3.6 27B?
Not at the recommended Q4_K_M quant. It needs about 18.7GB in this 4K-context estimate.
Which app should I use for local AI on the MacBook Neo A18 Pro?
For an exact GGUF repository and quant from this site, start with LM Studio's graphical Discover, download, load, and chat flow. Jan is the open-source GGUF alternative; Ollama is strongest when the exact model already has a trustworthy catalog package or you want coding/API integrations; Msty is useful for mixed GGUF, MLX, and document workflows. For the curated Bonsai 27B build, try Locally AI and first confirm that its catalog shows the model on this Mac.
Can I upgrade the MacBook Neo A18 Pro to more unified memory later?
No. Apple-silicon unified memory is integrated into the chip package and must be chosen at purchase. Storage upgrades or external SSDs do not increase the memory an LLM can use.
Are these MacBook Neo A18 Pro local AI speeds measured?
No. Every speed on this page is a formula range based on Apple’s published memory bandwidth, model size, active parameters, and a cooling-aware efficiency range. App, backend, context length, and thermals can change real results.