Can I run Holo4 35B-A3B locally?

The official Q4_K_M weights are 21.3GB, plus a separate 0.9GB vision projector, context and runtime memory. The model has 35B total parameters: 3B active does not mean a 3B download. A 32GB Mac is a more realistic starting point than a 24GB Mac, not a tested compatibility guarantee.

How to run Holo4 locally

The publisher provides Q4_K_M GGUF weights and mmproj.f16.gguf under Apache 2.0. Use a current multimodal llama.cpp runtime and load both files. The 21.3GB weights alone leave very little headroom on a 24GB machine; images, context, the projector and other applications add memory use.

Holo4 is a computer-use model, not a turnkey desktop agent. A separate harness must send screenshots, execute requested actions and return tool results. The publisher links hai-agents, but its current SDK quickstart uses a hosted API; that is not evidence of a fully local setup. AICanRun has not tested a local end-to-end harness or phone app, and tokens/s estimates do not measure task completion time.

Official GGUF weights and projector ↗ · Publisher model card ↗

Params
35B (3B active)
Family
holo4
Released
2026-09
Tags
vision · coding · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q4_K_M ★21.3 GB~25.4 GB0 memory-fit

Holo4 35B-A3B at Q4_K_M on 151 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Holo4 35B-A3B on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Holo4 35B-A3B on a phone?

At Q4_K_M, Holo4 35B-A3B needs ~25.4GB of usable memory (weights + KV cache + runtime). In practice that means no current Android phone or no current iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Holo4 35B-A3B run on a flagship phone?

It doesn't fit on any of the 151 phones we track at Q4_K_M — it needs ~25.4GB of usable memory.

Can an iPhone run Holo4 35B-A3B?

Not really: the best iPhone we track (iPhone 16 Pro Max) has ~5.2GB usable, but Holo4 35B-A3B needs 25.4GB at Q4_K_M.

What is the best quantization of Holo4 35B-A3B for mobile?

Q4_K_M (21.3GB download) is the size/quality sweet spot. Total memory needed is ~25.4GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Holo4 35B-A3B?

0 of the 151 phones we track have enough estimated usable memory at Q4_K_M; 0 also reach our usable-speed threshold of 3 tokens/s, and 0 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Holo4 35B-A3B?

36 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Holo4 35B-A3B good for on a phone?

It's tagged for vision, coding, reasoning. At 35B parameters (3B active — it's a MoE, so it decodes faster than its size suggests), it prioritizes answer quality over speed — expect slower decoding.

Similar models

Qwen 3.5 35B-A3BQwen 3.6 35B-A3BOrnith 1.0 35B-A3BXYZ-Aquila mini 35B-A3BKAT-Coder V2.5 Dev