Can AREX-2 27B run on MacBook Pro M3 Max 36GB?

YES — Runs great
Q4_K_M · 4K context · formula estimate

What this means

AREX-2 27B fits on the MacBook Pro M3 Max · 36GB at Q4_K_M. We estimate it uses 20.8GB of the conservative 32.4GB working budget.

Estimated decode speed is 8.7–12.2 tokens/s. A roughly 300-word answer may take around 38 seconds.

8.7–12.2tokens/s
estimated · faster than you read
20.8GB
needed at Q4_K_M
Q4_K_M
quant selected
11.6GB
working headroom
needs 20.8 GBworking budget 32.4 GB
Apple M3 Max (30-core GPU)36 GB unified memoryMetalActive cooling

See how fast it feels

Using the midpoint of our 8.7–12.2 tokens/s estimate for this demo.

Live demo · 10.5 tokens/s

▌

Where the memory goes

ComponentDetailGB
Model weightsQ4_K_M GGUF (17.2 GB) + mmap overhead18.1
KV cache4K context window1.9
RuntimemacOS inference app + compute buffers0.8
Total neededat Q4_K_M, 4K context20.8
Working budget36 GB unified memory − conservative macOS reserve32.4
Headroomremaining inside the working budget11.6

Pick your quant

QuantDownloadMemoryEstimated speedVerdict
Q2_K10.6 GB13.8 GB~14.2–19.8 tok/s✓ Runs great
Q3_K_M13.2 GB16.6 GB~11.4–15.9 tok/s✓ Runs great
IQ4_XS15.2 GB18.7 GB~9.9–13.8 tok/s✓ Runs great
Q4_016.1 GB19.6 GB~9.3–13 tok/s✓ Runs great
Q4_K_S16.1 GB19.6 GB~9.3–13 tok/s✓ Runs great
Q4_K_M ★17.2 GB20.8 GB~8.7–12.2 tok/s✓ Runs great
Q5_K_M20.7 GB24.4 GB~7.2–10.1 tok/s✓ Runs great
Q6_K23.6 GB27.5 GB~6.4–8.9 tok/s✓ Runs great
Q8_028.7 GB32.8 GB—✕ Won't fit

Use this model for its intended task

1
This is a task model, not a normal local chat download.

AREX-2 27B is designed to iterate on coding and deep-research tasks using tool results and feedback. This Mac has enough working memory for the tracked Q4_K_M weights, but loading the file in a plain chat window does not supply the inputs, tools, or action loop that make the model useful.

2
Start with the publisher's model page.

It documents the intended workflow and supported developer runtimes. This is not yet a point-and-click setup for beginners, so do not install LM Studio expecting an ordinary chat tutorial to reproduce the model's task.

Open BAAI/AREX-2 ↗

3
Just want to chat locally?

Choose Qwen 3.6 27B instead. Its page gives the complete graphical install, download, first prompt, and offline check.

This model is designed to iterate on coding and deep-research tasks using tool results and feedback. Its weights may fit in memory, but it needs a research or coding harness that executes tools and returns feedback between rounds. A normal LM Studio chat does not provide that workflow.

Other models on this Mac

AREX-2 27B on other MacBook Pro M3 Max configurations

FAQ

Can the MacBook Pro M3 Max · 36GB run AREX-2 27B?

Yes at Q4_K_M. We estimate about 20.8GB of working memory and 8.7–12.2 tokens/s at 4K context.

Which AREX-2 27B quant should I use on this Mac?

Q4_K_M. It is a 17.2GB download and leaves about 11.6GB inside our conservative working budget.

Which app should I use for AREX-2 27B on this Mac?

AREX-2 27B is a task-specific model, not a normal local chat download. The selected weights fit this Mac, but the model still needs its publisher's intended workflow. This page links that repository and recommends Qwen 3.6 27B if you just want local chat.

Are these speeds measured on a MacBook Pro M3 Max?

No. The range is a formula estimate based on memory bandwidth, model size, active parameters, cooling, and a 4K context. The app, backend, thermals, and prompt can change real performance.