G9v3 3B

At Q4_K_M, 134 of 135 phones have enough estimated usable memory; 134 also reach 3+ tokens/s. On Mac, 38 of 38 exact configurations pass the working-memory check.

Params
3B
Family
g9v3
Released
2026-07
Tags
chat · coding · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q2_K1.3 GB~2.2 GB134 memory-fit · 134 usable
Q3_K_M1.6 GB~2.5 GB134 memory-fit · 134 usable
IQ4_XS1.7 GB~2.6 GB134 memory-fit · 134 usable
Q4_01.8 GB~2.7 GB134 memory-fit · 134 usable
Q4_K_S1.8 GB~2.7 GB134 memory-fit · 134 usable
Q4_K_M1.9 GB~2.8 GB134 memory-fit · 134 usable
Q5_K_M2.2 GB~3.1 GB134 memory-fit · 134 usable
Q6_K2.5 GB~3.4 GB134 memory-fit · 134 usable
Q8_03.2 GB~4.2 GB132 memory-fit · 122 usable

G9v3 3B at Q4_K_M on 135 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

G9v3 3B on 38 Mac configurations

Fixed to Q4_K_M at 4K context. Each row links to the same full `/run` report used for phones, with a macOS-specific setup guide.

MacUnified memoryWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run G9v3 3B on a phone?

At Q4_K_M, G9v3 3B needs ~2.8GB of usable memory (weights + KV cache + runtime). In practice that means a 6GB+ Android phone or a 6GB+ iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does G9v3 3B run on a flagship phone?

On the Galaxy S25 Ultra (Snapdragon 8 Elite) it runs at ~18.2 tokens/s at Q4_K_M, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run G9v3 3B?

Yes — the iPhone 17 runs it at ~16.2 tokens/s at Q4_K_M, using 2.8GB of its ~5.2GB usable memory.

What is the best quantization of G9v3 3B for mobile?

Q4_K_M (1.9GB download) is the size/quality sweet spot of the 9 quants available. Total memory needed is ~2.8GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run G9v3 3B?

134 of the 135 phones we track have enough estimated usable memory at Q4_K_M; 134 also reach our usable-speed threshold of 3 tokens/s, and 103 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run G9v3 3B?

38 of the 38 exact Mac configurations we track have enough estimated working memory at Q4_K_M. Each Mac result separates installed unified memory, the macOS reserve, estimated speed, and quant choices.

What is G9v3 3B good for on a phone?

It's tagged for chat, coding, reasoning. At 3B parameters, it's a fast, lightweight pick for quick tasks.

Similar models

Ministral 3 3BSmolLM3 3BLlama 3.2 3BPhi-4 Mini 3.8BQwen3 4B