Can Ternary Bonsai 1.7B run on MacBook Pro M5 Pro 24GB?
What this means
Ternary Bonsai 1.7B fits on the MacBook Pro M5 Pro · 24GB at Q2_0. We estimate it uses 1.4GB of the conservative 21GB working budget.
Estimated decode speed is 307–429.8 tokens/s. A roughly 300-word answer may take around 1 seconds.
See how fast it feels
Using the midpoint of our 307–429.8 tokens/s estimate for this demo.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q2_0 GGUF (0.5 GB) + mmap overhead | 0.5 |
| KV cache | 4K context window | 0.1 |
| Runtime | macOS inference app + compute buffers | 0.8 |
| Total needed | Q2_0, 4K context | 1.4 |
| Working budget | 24 GB unified memory − conservative macOS reserve | 21 |
| Headroom | remaining inside the working budget | 19.6 |
Pick your quant
| Quant | Download | Memory | Estimated speed | Verdict |
|---|---|---|---|---|
| Q2_0 ★ | 0.5 GB | 1.4 GB | ~307–429.8 tok/s | ✓ Runs great |
| Q2_0_g64 | 0.5 GB | 1.4 GB | ~307–429.8 tok/s | ✓ Runs great |
Get it running on this Mac
This Mac has enough memory, but the tracked Q2_0 ternary file is not supported by the current point-and-click apps. For a first local chat, use Bonsai 27B 1-bit in Locally AI. Do not download PQ2_0; the publisher marks it unsupported.
The advanced path is complete below, but it is not required for ordinary local AI.
Advanced publisher setup (Terminal)
Install Apple's command-line tools if macOS asks, then run:
git clone https://github.com/PrismML-Eng/Bonsai-demo.git
cd Bonsai-demo
BONSAI_OPENWEBUI=0 BONSAI_SKIP_MLX=1 BONSAI_FAMILY=ternary BONSAI_MODEL=1.7B ./setup.sh
BONSAI_FAMILY=ternary BONSAI_MODEL=1.7B BONSAI_CTX=4096 ./scripts/start_llama_server.shWhen the server says it is ready, open http://localhost:8080. Start with 4K context. This route uses the publisher's Prism runtime, not Jan or LM Studio.
Other models on this Mac
Ternary Bonsai 1.7B on other MacBook Pro M5 Pro configurations
FAQ
Can the MacBook Pro M5 Pro · 24GB run Ternary Bonsai 1.7B?
Yes at Q2_0. We estimate about 1.4GB of working memory and 307–429.8 tokens/s at 4K context.
Which Ternary Bonsai 1.7B quant should I use on this Mac?
Q2_0. It is a 0.5GB download and leaves about 19.6GB inside our conservative working budget.
Which app should I use for Ternary Bonsai 1.7B on this Mac?
Current beginner Mac apps do not load this exact file. Use Bonsai 27B 1-bit in Locally AI instead, or open the advanced publisher instructions on this page.
Are these speeds measured on a MacBook Pro M5 Pro?
No. The range is a formula estimate based on memory bandwidth, model size, active parameters, cooling, and a 4K context. The app, backend, thermals, and prompt can change real performance.