Can I run Qwen3.8-Flash-Next locally?

At the default IQ4_XS, 0 of 140 phones have enough estimated usable memory; 0 also reach 3+ tokens/s. On Mac, 11 of 54 exact configurations pass using that quant or a smaller tracked fallback.

EXPERIMENTAL RUNTIME · HIGH MEMORY

The model activates about 6B parameters per token, but the full 125B weight set still has to live in memory. The tracked IQ1_S download is roughly 72.5GB before KV cache and runtime overhead, so 96GB is a practical floor and 128GB is the safer target. Qwen labels Next as an experimental architecture preview, not the production Qwen3.8-Flash release.

Runtime support is also early: the GGUF publisher currently points users to Unsloth Desktop or its llama.cpp PR 27742. A reported memory fit does not mean LM Studio, Ollama, or phone apps can load it yet.

Params
125B (6B active)
Family
qwen
Released
2026-08
Tags
chat · coding · vision · reasoning

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
IQ1_S72.5 GB~85.5 GB0 memory-fit
IQ4_XS93.7 GB~107.7 GB0 memory-fit

Qwen3.8-Flash-Next at IQ4_XS on 140 phones

This table is fixed to the recommended IQ4_XS, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Qwen3.8-Flash-Next on 54 Mac configurations

Each Mac uses the recommended IQ4_XS at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Qwen3.8-Flash-Next on a phone?

At IQ4_XS, Qwen3.8-Flash-Next needs ~107.7GB of usable memory (weights + KV cache + runtime). In practice that means no current Android phone or no current iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Qwen3.8-Flash-Next run on a flagship phone?

It doesn't fit on any of the 140 phones we track at IQ4_XS — it needs ~107.7GB of usable memory.

Can an iPhone run Qwen3.8-Flash-Next?

Not really: the best iPhone we track (iPhone 16 Pro Max) has ~5.2GB usable, but Qwen3.8-Flash-Next needs 107.7GB at IQ4_XS.

What is the best quantization of Qwen3.8-Flash-Next for mobile?

IQ4_XS (93.7GB download) is the size/quality sweet spot of the 2 quants available. Total memory needed is ~107.7GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Qwen3.8-Flash-Next?

0 of the 140 phones we track have enough estimated usable memory at IQ4_XS; 0 also reach our usable-speed threshold of 3 tokens/s, and 0 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Qwen3.8-Flash-Next?

11 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended IQ4_XS when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Qwen3.8-Flash-Next good for on a phone?

It's tagged for chat, coding, vision, reasoning. At 125B parameters (6B active — it's a MoE, so it decodes faster than its size suggests), it prioritizes answer quality over speed — expect slower decoding.

Guides for Qwen3.8-Flash-Next

Bonsai 27B vs Qwen3.6 27B on Phones: Is 1-Bit Worth It?

Similar models

Qwen3 0.6BQwen3 1.7BQwen3 4BQwen3 8BQwen3 14B