Can AREX-2 27B run on Mac Studio M5 Max 128GB?
What this means
AREX-2 27B fits on the Mac Studio M5 Max · 128GB at Q4_K_M. We estimate it uses 20.8GB of the conservative 120GB working budget.
Estimated decode speed is 17.8–25 tokens/s. A roughly 300-word answer may take around 19 seconds.
See how fast it feels
Using the midpoint of our 17.8–25 tokens/s estimate for this demo.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (17.2 GB) + mmap overhead | 18.1 |
| KV cache | 4K context window | 1.9 |
| Runtime | macOS inference app + compute buffers | 0.8 |
| Total needed | at Q4_K_M, 4K context | 20.8 |
| Working budget | 128 GB unified memory − conservative macOS reserve | 120 |
| Headroom | remaining inside the working budget | 99.2 |
Pick your quant
| Quant | Download | Memory | Estimated speed | Verdict |
|---|---|---|---|---|
| Q2_K | 10.6 GB | 13.8 GB | ~29–40.5 tok/s | ✓ Runs great |
| Q3_K_M | 13.2 GB | 16.6 GB | ~23.3–32.6 tok/s | ✓ Runs great |
| IQ4_XS | 15.2 GB | 18.7 GB | ~20.2–28.3 tok/s | ✓ Runs great |
| Q4_0 | 16.1 GB | 19.6 GB | ~19.1–26.7 tok/s | ✓ Runs great |
| Q4_K_S | 16.1 GB | 19.6 GB | ~19.1–26.7 tok/s | ✓ Runs great |
| Q4_K_M ★ | 17.2 GB | 20.8 GB | ~17.8–25 tok/s | ✓ Runs great |
| Q5_K_M | 20.7 GB | 24.4 GB | ~14.8–20.8 tok/s | ✓ Runs great |
| Q6_K | 23.6 GB | 27.5 GB | ~13–18.2 tok/s | ✓ Runs great |
| Q8_0 | 28.7 GB | 32.8 GB | ~10.7–15 tok/s | ✓ Runs great |
Use this model for its intended task
AREX-2 27B is designed to iterate on coding and deep-research tasks using tool results and feedback. This Mac has enough working memory for the tracked Q4_K_M weights, but loading the file in a plain chat window does not supply the inputs, tools, or action loop that make the model useful.
It documents the intended workflow and supported developer runtimes. This is not yet a point-and-click setup for beginners, so do not install LM Studio expecting an ordinary chat tutorial to reproduce the model's task.
Choose Qwen 3.6 27B instead. Its page gives the complete graphical install, download, first prompt, and offline check.
This model is designed to iterate on coding and deep-research tasks using tool results and feedback. Its weights may fit in memory, but it needs a research or coding harness that executes tools and returns feedback between rounds. A normal LM Studio chat does not provide that workflow.
Other models on this Mac
AREX-2 27B on other Mac Studio M5 Max configurations
FAQ
Can the Mac Studio M5 Max · 128GB run AREX-2 27B?
Yes at Q4_K_M. We estimate about 20.8GB of working memory and 17.8–25 tokens/s at 4K context.
Which AREX-2 27B quant should I use on this Mac?
Q4_K_M. It is a 17.2GB download and leaves about 99.2GB inside our conservative working budget.
Which app should I use for AREX-2 27B on this Mac?
AREX-2 27B is a task-specific model, not a normal local chat download. The selected weights fit this Mac, but the model still needs its publisher's intended workflow. This page links that repository and recommends Qwen 3.6 27B if you just want local chat.
Are these speeds measured on a Mac Studio M5 Max?
No. The range is a formula estimate based on memory bandwidth, model size, active parameters, cooling, and a 4K context. The app, backend, thermals, and prompt can change real performance.