GPT-OSS 20B

At the default MXFP4, 4 of 140 phones have enough estimated usable memory; 4 also reach 3+ tokens/s. On Mac, 46 of 54 exact configurations pass using that quant or a smaller tracked fallback.

Params
21B (3.6B active)
Family
gpt-oss
Released
2025-08
Tags
chat · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
MXFP412.1 GB~14.8 GB4 memory-fit · 4 usable

GPT-OSS 20B at MXFP4 on 140 phones

This table is fixed to the recommended MXFP4, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

GPT-OSS 20B on 54 Mac configurations

Each Mac uses the recommended MXFP4 at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run GPT-OSS 20B on a phone?

At MXFP4, GPT-OSS 20B needs ~14.8GB of usable memory (weights + KV cache + runtime). In practice that means a 24GB+ Android phone or a 24GB+ iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does GPT-OSS 20B run on a flagship phone?

On the OnePlus 13 (Snapdragon 8 Elite) it runs at ~16.7 tokens/s at MXFP4, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run GPT-OSS 20B?

Not really: the best iPhone we track (iPhone Air) has ~7.8GB usable, but GPT-OSS 20B needs 14.8GB at MXFP4.

What is the best quantization of GPT-OSS 20B for mobile?

MXFP4 (12.1GB download) is the size/quality sweet spot. Total memory needed is ~14.8GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run GPT-OSS 20B?

4 of the 140 phones we track have enough estimated usable memory at MXFP4; 4 also reach our usable-speed threshold of 3 tokens/s, and 4 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run GPT-OSS 20B?

46 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended MXFP4 when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is GPT-OSS 20B good for on a phone?

It's tagged for chat, reasoning. At 21B parameters (3.6B active — it's a MoE, so it decodes faster than its size suggests), it prioritizes answer quality over speed — expect slower decoding.

Similar models

Gemma 4 26B-A4BQwen 3.5 27BQwen 3.6 27BBonsai 27B (1-bit)Ternary Bonsai 27B