Llama 3.3 70B

At the default Q4_K_M, 0 of 140 phones have enough estimated usable memory; 0 also reach 3+ tokens/s. On Mac, 25 of 54 exact configurations pass using that quant or a smaller tracked fallback.

Params
70B
Family
llama
Released
2024-12
Tags
chat · coding

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q2_K26.4 GB~33.2 GB0 memory-fit
Q3_K_M34.3 GB~41.5 GB0 memory-fit
IQ4_XS37.9 GB~45.3 GB0 memory-fit
Q4_040.1 GB~47.6 GB0 memory-fit
Q4_K_S40.3 GB~47.8 GB0 memory-fit
Q4_K_M42.5 GB~50.1 GB0 memory-fit
Q5_K_M49.9 GB~57.9 GB0 memory-fit
Q6_K57.9 GB~66.3 GB0 memory-fit
Q8_075 GB~84.3 GB0 memory-fit

Llama 3.3 70B at Q4_K_M on 140 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Llama 3.3 70B on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Llama 3.3 70B on a phone?

At Q4_K_M, Llama 3.3 70B needs ~50.1GB of usable memory (weights + KV cache + runtime). In practice that means no current Android phone or no current iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Llama 3.3 70B run on a flagship phone?

It doesn't fit on any of the 140 phones we track at Q4_K_M — it needs ~50.1GB of usable memory.

Can an iPhone run Llama 3.3 70B?

Not really: the best iPhone we track (iPhone 16 Pro Max) has ~5.2GB usable, but Llama 3.3 70B needs 50.1GB at Q4_K_M.

What is the best quantization of Llama 3.3 70B for mobile?

Q4_K_M (42.5GB download) is the size/quality sweet spot of the 9 quants available. Total memory needed is ~50.1GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Llama 3.3 70B?

0 of the 140 phones we track have enough estimated usable memory at Q4_K_M; 0 also reach our usable-speed threshold of 3 tokens/s, and 0 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Llama 3.3 70B?

25 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Llama 3.3 70B good for on a phone?

It's tagged for chat, coding. At 70B parameters, it prioritizes answer quality over speed — expect slower decoding.

Similar models

Llama 3.2 1BLlama 3.2 3BLlama 3.1 8BBigBang V1 36B-A3BQwen 3.5 35B-A3B