Can Holo4 35B-A3B run on MacBook Pro M3 Pro 18GB?

NO — Won’t fit
Q4_K_M · 4K context · formula estimate

What this means

The checked Q4_K_M build needs about 25.6GB, above this Mac's conservative 15GB working budget.

−10.6GB short
above the working budget
25.6GB
needed at Q4_K_M
Q4_K_M
quant selected
0GB
working headroom
needs 25.6 GBworking budget 15 GB · short 10.6 GB
Apple M3 Pro18 GB unified memoryMetalActive cooling

Where the memory goes

ComponentDetailGB
Model weightsQ4_K_M GGUF (21.3 GB) + mmap overhead22.4
KV cache4K context window2.5
RuntimemacOS inference app + compute buffers0.8
Total neededat Q4_K_M, 4K context25.6
Working budget18 GB unified memory − conservative macOS reserve15
Headroommemory shortfall−10.6

Pick your quant

QuantDownloadMemoryEstimated speedVerdict
Q4_K_M ★21.3 GB25.6 GB—✕ Won't fit

Other models on this Mac

Holo4 35B-A3B on other MacBook Pro M3 Pro configurations

FAQ

Can the MacBook Pro M3 Pro · 18GB run Holo4 35B-A3B?

Not with the quants currently tracked. The selected Q4_K_M build needs about 25.6GB, above the 15GB working budget.

Which Holo4 35B-A3B quant should I use on this Mac?

None of the tracked quants fit this configuration safely. Choose a smaller model or a Mac with more unified memory.

Which app should I use for Holo4 35B-A3B on this Mac?

Holo4 35B-A3B is a task-specific model, not a normal local chat download. The selected weights do not fit this Mac, and the model also needs its publisher's intended workflow. This page links that repository and recommends Qwen 3.6 35B-A3B if you just want local chat.

Are these speeds measured on a MacBook Pro M3 Pro?

No. The range is a formula estimate based on memory bandwidth, model size, active parameters, cooling, and a 4K context. The app, backend, thermals, and prompt can change real performance.