Phi-4 Mini 3.8B

At the default Q4_K_M, 139 of 140 phones have enough estimated usable memory; 139 also reach 3+ tokens/s. On Mac, 54 of 54 exact configurations pass using that quant or a smaller tracked fallback.

Params
3.8B
Family
phi
Released
2025-02
Tags
chat · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q2_K1.7 GB~2.7 GB139 memory-fit · 139 usable
Q3_K_M2.1 GB~3.1 GB139 memory-fit · 139 usable
IQ4_XS2.2 GB~3.2 GB139 memory-fit · 139 usable
Q4_02.3 GB~3.3 GB139 memory-fit · 139 usable
Q4_K_S2.3 GB~3.3 GB139 memory-fit · 139 usable
Q4_K_M2.5 GB~3.5 GB139 memory-fit · 139 usable
Q5_K_M2.8 GB~3.8 GB139 memory-fit · 127 usable
Q6_K3.2 GB~4.2 GB137 memory-fit · 125 usable
Q8_04.1 GB~5.2 GB137 memory-fit · 104 usable

Phi-4 Mini 3.8B at Q4_K_M on 140 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Phi-4 Mini 3.8B on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Phi-4 Mini 3.8B on a phone?

At Q4_K_M, Phi-4 Mini 3.8B needs ~3.5GB of usable memory (weights + KV cache + runtime). In practice that means a 6GB+ Android phone or a 6GB+ iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Phi-4 Mini 3.8B run on a flagship phone?

On the OPPO Find X8 Pro (Dimensity 9400) it runs at ~15.4 tokens/s at Q4_K_M, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run Phi-4 Mini 3.8B?

Yes — the iPhone Air runs it at ~13.8 tokens/s at Q4_K_M, using 3.5GB of its ~7.8GB usable memory.

What is the best quantization of Phi-4 Mini 3.8B for mobile?

Q4_K_M (2.5GB download) is the size/quality sweet spot of the 9 quants available. Total memory needed is ~3.5GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Phi-4 Mini 3.8B?

139 of the 140 phones we track have enough estimated usable memory at Q4_K_M; 139 also reach our usable-speed threshold of 3 tokens/s, and 102 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Phi-4 Mini 3.8B?

54 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Phi-4 Mini 3.8B good for on a phone?

It's tagged for chat, reasoning. At 3.8B parameters, it's a fast, lightweight pick for quick tasks.

Similar models

Phi-4 14BQwen3 4BQwen 3.5 4BGemma 4 E4BTernary Bonsai 4B