Can Holo4 35B-A3B run on Mac mini M5 Pro 64GB?

YES — Runs great
Q4_K_M · 4K context · formula estimate

What this means

Holo4 35B-A3B fits on the Mac mini M5 Pro · 64GB at Q4_K_M. We estimate it uses 25.6GB of the conservative 57.6GB working budget.

Estimated decode speed is 84.1–117.7 tokens/s. A roughly 300-word answer may take around 4 seconds.

84.1–117.7tokens/s
estimated · instant
25.6GB
needed at Q4_K_M
Q4_K_M
quant selected
32GB
working headroom
needs 25.6 GBworking budget 57.6 GB
Apple M5 Pro64 GB unified memoryMetalActive cooling

See how fast it feels

Using the midpoint of our 84.1–117.7 tokens/s estimate for this demo.

Live demo · 100.9 tokens/s

▌

Where the memory goes

ComponentDetailGB
Model weightsQ4_K_M GGUF (21.3 GB) + mmap overhead22.4
KV cache4K context window2.5
RuntimemacOS inference app + compute buffers0.8
Total neededat Q4_K_M, 4K context25.6
Working budget64 GB unified memory − conservative macOS reserve57.6
Headroomremaining inside the working budget32

Pick your quant

QuantDownloadMemoryEstimated speedVerdict
Q4_K_M ★21.3 GB25.6 GB~84.1–117.7 tok/s✓ Runs great

Use this model for its intended task

1
This is a task model, not a normal local chat download.

Holo4 35B-A3B is designed to operate desktop applications using screenshots, clicks, typing and tool results. This Mac has enough working memory for the tracked Q4_K_M weights, but loading the file in a plain chat window does not supply the inputs, tools, or action loop that make the model useful.

2
Start with the publisher's model page.

It documents the intended workflow and supported developer runtimes. This is not yet a point-and-click setup for beginners, so do not install LM Studio expecting an ordinary chat tutorial to reproduce the model's task.

Open Hcompany/Holo4-35B-A3B ↗

3
Just want to chat locally?

Choose Qwen 3.6 35B-A3B instead. Its page gives the complete graphical install, download, first prompt, and offline check.

This model is designed to operate desktop applications using screenshots, clicks, typing and tool results. Its weights may fit in memory, but it needs a screenshot/action harness and the separate vision projector; loading GGUF weights alone does not control a computer. A normal LM Studio chat does not provide that workflow.

Other models on this Mac

Holo4 35B-A3B on other Mac mini M5 Pro configurations

FAQ

Can the Mac mini M5 Pro · 64GB run Holo4 35B-A3B?

Yes at Q4_K_M. We estimate about 25.6GB of working memory and 84.1–117.7 tokens/s at 4K context.

Which Holo4 35B-A3B quant should I use on this Mac?

Q4_K_M. It is a 21.3GB download and leaves about 32GB inside our conservative working budget.

Which app should I use for Holo4 35B-A3B on this Mac?

Holo4 35B-A3B is a task-specific model, not a normal local chat download. The selected weights fit this Mac, but the model still needs its publisher's intended workflow. This page links that repository and recommends Qwen 3.6 35B-A3B if you just want local chat.

Are these speeds measured on a Mac mini M5 Pro?

No. The range is a formula estimate based on memory bandwidth, model size, active parameters, cooling, and a 4K context. The app, backend, thermals, and prompt can change real performance.