Can Grug 12B run on iPhone Duo?

Some specifications for iPhone Duo are not confirmed. Compatibility and speed estimates use provisional chip or RAM information. Announced; available October 23, 2026. RAM capacity is not confirmed; results assume a provisional 12GB configuration.
YES — Runs great
Formula estimate

What this means

Our estimate says Grug 12B should fit comfortably on your iPhone Duo.

Download the Q2_K version, which is about 5.1 GB. We expect it to use about 6.8 GB of the roughly 7.8 GB available to a model on this phone.

At the estimated speed, a roughly 300-word answer may take about 39 seconds to finish.

This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.

10.2tokens/s
estimated · faster than you read
6.8GB
needed at Q2_K
Q2_K
quant checked
needs 6.8 GBusable 7.8 GB
Apple A20 Pro12 GB RAMMetal

See how fast it feels

Estimated at 10.2 tokens/s — faster than you read. A ~300-word reply takes about 39 seconds on this phone.

Live demo · 10.2 tokens/s

▌

Where the memory goes

ComponentDetailGB
Model weightsQ2_K GGUF (5.1 GB) + mmap overhead5.4
KV cache4K context window0.8
Runtimellama.cpp + app overhead0.6
Total neededat Q2_K, 4K context6.8
Working budget12 GB RAM − iOS reserve (jetsam limit)7.8
Headroomremaining inside the working budget1

Pick your quant

QuantDownloadVerdictSpeed
Q2_K BEST HERE5.1 GB✓ Runs great~10.2 tokens/s
Q3_K_M 6.3 GB✕ Won't fitwon't fit
IQ4_XS 6.8 GB✕ Won't fitwon't fit
Q4_0 7.1 GB✕ Won't fitwon't fit
Q4_K_S 7.2 GB✕ Won't fitwon't fit
Q4_K_M 7.7 GB✕ Won't fitwon't fit
Q5_K_M 8.8 GB✕ Won't fitwon't fit
Q6_K 10.2 GB✕ Won't fitwon't fit
Q8_0 12.7 GB✕ Won't fitwon't fit

Get your first offline chat working

Recommended app: PocketPal. Follow the point-and-click steps below. The speed above is an estimate, not a measurement from this exact app and phone.
1
Install or update PocketPal from the App Store. It is free and does not require an account. Use a current version so its loader supports newer model architectures.
2
Open the exact model. In PocketPal, go to Models → + → Add from Hugging Face, then paste bartowski/kai-os_Grug-12B-GGUF.
3
Choose the Q2_K GGUF file. The download is about 5.1 GB, so use Wi-Fi and keep the app open. Choose the main GGUF weights, not a vision projector, mmproj, or other helper file.
4
Tap Download, then Load. Start with a 4K (4096-token) context. Keep PocketPal’s default Metal acceleration; this does not require Xcode.
5
Send a simple first prompt. Try “Explain why the sky is blue in three sentences.” This page estimates about 10.2 tokens/s, but that number is not a PocketPal measurement unless it carries a ✓ Verified label.
6
Confirm it is really offline. After the first reply, turn on airplane mode and ask a second question. If it still answers, the model is running on your phone.
If it does not work
  • Model not listed: update PocketPal and paste the exact repository bartowski/kai-os_Grug-12B-GGUF.
  • App closes while loading: close other apps, restart the phone, and try 2K context. If it still closes, choose a smaller model.
  • No offline reply: confirm that the Q2_K GGUF file is loaded in the chat rather than a remote model.

Related checks

More on iPhone Duo
Qwen3 0.6BQwen3 1.7BQwen3 4BLlama 3.2 1BLlama 3.2 3B
Grug 12B on other phones
Xiaomi 18 FoldOPPO Find X8 Provivo X200 ProGalaxy S26 UltraXiaomi 17