Can I run Qwen3.8-27B-pi locally?

Yes on a suitable Mac or PC: the publisher's Q4_K_M weights are 16.5GB before runtime and context memory; 32GB unified memory or a 24GB GPU is the practical starting class. At Q4_K_M, the estimated 4K working set is ~19.9GB. This is a Qwen3.8 coding fine-tune for the Pi agent harness, not a new Qwen base model. The fit tables estimate text inference at 4K context without a vision projector or speculative draft. Vision, MTP, DFlash2 and long contexts need extra memory. We have not measured its agent success or inference speed.

Params
27.8B
Family
qwen
Released
2026-09
Tags
chat · coding · vision · reasoning

How to run Qwen3.8-27B-pi locally

This text-only starting command adapts the publisher's llama.cpp recipe to 4K context, omitting the optional MTP draft and vision projector. It is source-checked, not locally benchmarked. Connect Pi separately to the local server with compatible tool calling before expecting repository edits or tool execution. Do not allocate the publisher's full 256K context on a minimum-memory machine.

MAC OR PC · LLAMA.CPP
hf download bytkim/Qwen3.8-27B-pi-GGUF --include Qwen3.8-27B-pi-Q4_K_M.gguf --local-dir models
llama-server --model models/Qwen3.8-27B-pi-Q4_K_M.gguf --host 127.0.0.1 --ctx-size 4096 --parallel 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
Publisher GGUF setup, sampling and optional companion files ↗

Memory figures on this page are planning estimates from GGUF bytes, 5% weight overhead, a 4K KV cache, and runtime allowance—not measured peak allocation or a speed guarantee.

Downloads & memory by quant

The download is only the weights — add KV cache and runtime to get what your phone actually needs. “Memory-fit” only means it can load; “usable” also requires an estimated 3+ tokens/s. Each row uses that row's quant. ★ = recommended.

QuantDownloadMemory neededPhone reality
Q2_K10.7 GB~13.8 GB4 memory-fit · 4 usable
Q3_K_M13.3 GB~16.5 GB3 memory-fit · 0 usable
IQ4_XS15.1 GB~18.4 GB3 memory-fit · 0 usable
Q4_K_S15.6 GB~18.9 GB3 memory-fit · 0 usable
Q4_K_M ★16.5 GB~19.9 GB3 memory-fit · 0 usable
Q5_K_M19.2 GB~22.7 GB0 memory-fit
Q6_K22.1 GB~25.8 GB0 memory-fit
Q8_028.6 GB~32.6 GB0 memory-fit

Qwen3.8-27B-pi at Q4_K_M on 151 phones

This table is fixed to the recommended Q4_K_M, on each phone's largest RAM option. Counts for smaller quants above will differ. All phones are shown, sorted by fit and estimated experience; tap one for the full report.

PhoneChip · RAMNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Qwen3.8-27B-pi on 54 Mac configurations

Each Mac uses the recommended Q4_K_M at 4K context when it fits, or the largest smaller tracked quant that fits. Each row links to the full report with a macOS-specific setup guide.

MacUnified memoryQuantWorking budgetNeedsEstimated speedVerdict

FAQ

How much RAM do I need to run Qwen3.8-27B-pi on a phone?

At Q4_K_M, Qwen3.8-27B-pi needs ~19.9GB of usable memory (weights + KV cache + runtime). In practice that means a 24GB+ Android phone or no current iPhone — iOS lets a single app use less of its RAM than Android does.

How fast does Qwen3.8-27B-pi run on a flagship phone?

On the OnePlus 13 (Snapdragon 8 Elite) it runs at ~2.1 tokens/s at Q4_K_M, estimated from memory bandwidth. Anything above ~8 tokens/s feels smooth for chat.

Can an iPhone run Qwen3.8-27B-pi?

Not really: the best iPhone we track (iPhone 16 Pro Max) has ~5.2GB usable, but Qwen3.8-27B-pi needs 19.9GB at Q4_K_M.

What is the best quantization of Qwen3.8-27B-pi for mobile?

Q4_K_M (16.5GB download) is the size/quality sweet spot of the 8 quants available. Total memory needed is ~19.9GB once the KV cache and runtime are counted — the download size alone understates it.

How many phones can run Qwen3.8-27B-pi?

3 of the 151 phones we track have enough estimated usable memory at Q4_K_M; 0 also reach our usable-speed threshold of 3 tokens/s, and 0 run smoothly. Memory-fit alone does not mean a good experience.

Which Macs can run Qwen3.8-27B-pi?

46 of the 54 exact Mac configurations we track have enough estimated working memory. Each Mac uses the recommended Q4_K_M when it fits, or the largest smaller tracked quant that fits, and separates installed unified memory, the macOS reserve, and estimated speed.

What is Qwen3.8-27B-pi good for on a phone?

It's tagged for chat, coding, vision, reasoning. At 27.8B parameters, it prioritizes answer quality over speed — expect slower decoding.

Guides for Qwen3.8-27B-pi

Bonsai 27B vs Qwen3.6 27B on Phones: Is 1-Bit Worth It?

Similar models

Qwen3 0.6BQwen3 1.7BQwen3 4BQwen3 8BQwen3 14B