Ling 3.0 Tiny

At the default Q4_K_M, 105 of 139 phones have enough estimated usable memory; 105 also reach 3+ tokens/s. On Mac, 54 of 54 exact configurations pass using that quant or a smaller tracked fallback.

Params
7.9B (1.4B active)
Family
ling
Released
2026-08
Tags
chat · coding · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q2_K3 GB~4.3 GB136 memory-fit · 136 usable
Q3_K_M3.8 GB~5.1 GB136 memory-fit · 136 usable
IQ4_XS4.4 GB~5.8 GB129 memory-fit · 129 usable
Q4_04.6 GB~6 GB129 memory-fit · 129 usable
Q4_K_S4.8 GB~6.2 GB105 memory-fit · 105 usable
Q4_K_M4.9 GB~6.3 GB105 memory-fit · 105 usable
Q5_K_M5.7 GB~7.1 GB105 memory-fit · 105 usable
Q6_K6.8 GB~8.3 GB102 memory-fit · 102 usable
Q8_08.4 GB~10 GB59 memory-fit · 59 usable

Ling 3.0 Tiny at Q4_K_M on 139 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Ling 3.0 Tiny on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Ling 3.0 Tiny on a phone?

At Q4_K_M, Ling 3.0 Tiny needs ~6.3GB of usable memory (weights + KV cache + runtime). In practice that means a 12GB+ Android phone or a 12GB+ iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Ling 3.0 Tiny run on a flagship phone?

On the Galaxy S25 Ultra (Snapdragon 8 Elite) it runs at ~39.8 tokens/s at Q4_K_M, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run Ling 3.0 Tiny?

Yes — the iPhone Air runs it at ~39.8 tokens/s at Q4_K_M, using 6.3GB of its ~7.8GB usable memory.

What is the best quantization of Ling 3.0 Tiny for mobile?

Q4_K_M (4.9GB download) is the size/quality sweet spot of the 9 quants available. Total memory needed is ~6.3GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Ling 3.0 Tiny?

105 of the 139 phones we track have enough estimated usable memory at Q4_K_M; 105 also reach our usable-speed threshold of 3 tokens/s, and 105 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Ling 3.0 Tiny?

54 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Ling 3.0 Tiny good for on a phone?

It's tagged for chat, coding, reasoning. At 7.9B parameters (1.4B active — it's a MoE, so it decodes faster than its size suggests), it balances quality and speed for daily use.

Full requirements guide
Can I run Ling 3.0 Tiny locally? Real GGUF sizes, 8/12/16GB answers, and setup

Similar models

Ling 3.0 FlashLlama 3.1 8BMinistral 8BTernary Bonsai 8BMinistral 3 8B