Can Gemma 3 12B run on Nothing Phone (2a)?

YESRuns, barely
Formula estimate

What this means

Our estimate says Gemma 3 12B should fit, but there is very little memory left for the operating system and other apps.

Download the Q4_0 version, which is about 6.9 GB. We expect it to use about 8.7 GB of the roughly 9 GB available to a model on this phone.

At the estimated speed, a roughly 300-word answer may take about 4 minutes to finish.

This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.

1.7tokens/s
estimated · slower than you read
8.7GB
needed at Q4_0
Q4_0
quant checked
needs 8.7 GBusable 9 GB
Dimensity 7200 Pro12 GB RAMCPU / GPUAlso sold with 8 GB

Make it bearable

1
Drop context to 2K — saves ~0.5 GB of KV cache and a little speed.
2
Close every other app before loading — the 0.3 GB headroom is real; Android may kill the app otherwise.
3
Expect throttling after ~10 min of sustained generation on a phone chassis.

See how fast it feels

Estimated at 1.7 tokens/s — slower than you read. A ~300-word reply takes about 235 seconds on this phone.

Live demo · 1.7 tokens/s

Where the memory goes

ComponentDetailGB
Model weightsQ4_0 GGUF (6.9 GB) + mmap overhead7.2
KV cache4K context window0.9
Runtimellama.cpp + app overhead0.6
Total neededat Q4_0, 4K context8.7
Working budget12 GB RAM Android system reserve9
Headroomremaining inside the working budget0.3

Pick your quant

QuantDownloadVerdictSpeed
IQ1_S 3.1 GB! Runs, barely~3.7 tokens/s
Q2_K 4.8 GB! Runs, barely~2.4 tokens/s
Q3_K_M 6 GB! Runs, barely~1.9 tokens/s
IQ4_XS 6.6 GB! Runs, barely~1.7 tokens/s
Q4_0 BEST HERE6.9 GB! Runs, barely~1.7 tokens/s
Q4_K_S 6.9 GB! Runs, barely~1.7 tokens/s
Q4_K_M 7.3 GB Won't fitwon't fit
Q5_K_M 8.4 GB Won't fitwon't fit
Q6_K 9.7 GB Won't fitwon't fit
Q8_0 12.5 GB Won't fitwon't fit

Get your first offline chat working

Recommended app: PocketPal. Follow the point-and-click steps below. The speed above is an estimate, not a measurement from this exact app and phone.
1
Install or update PocketPal from Google Play. It is free and does not require an account. Use a current version so its loader supports newer model architectures.
2
Open the exact model. In PocketPal, go to Models → + → Add from Hugging Face, then paste unsloth/gemma-3-12b-it-GGUF.
3
Choose the Q4_0 GGUF file. The download is about 6.9 GB, so use Wi-Fi and keep the app open. Choose the main GGUF weights, not a vision projector, mmproj, or other helper file.
4
Tap Download, then Load. Start with a 4K (4096-token) context. Keep the app’s default Android backend for the first run.
5
Send a simple first prompt. Try “Explain why the sky is blue in three sentences.” This page estimates about 1.7 tokens/s, but that number is not a PocketPal measurement unless it carries a ✓ Verified label.
6
Confirm it is really offline. After the first reply, turn on airplane mode and ask a second question. If it still answers, the model is running on your phone.
If it does not work
  • Model not listed: update PocketPal and paste the exact repository unsloth/gemma-3-12b-it-GGUF.
  • App closes while loading: close other apps, restart the phone, and try 2K context. If it still closes, choose a smaller model.
  • No offline reply: confirm that the Q4_0 GGUF file is loaded in the chat rather than a remote model.

Related checks

More on Nothing Phone (2a)
Qwen3 0.6BTernary Bonsai 1.7BLFM2.5 8B-A1BOvisOCR2 0.8BTini Cybersec 8B-A1B
Gemma 3 12B on other phones
OPPO Find X8 Provivo X200 ProGalaxy S26 UltraXiaomi 17Xiaomi 17 Pro