AI Models for ROG Phone 9 Pro — What runs on 24GB

34 great · 23 slow · 36 won't fit
Chip
Snapdragon 8 Elite
Memory bandwidth
76.8 GB/s
NPU
45 TOPS
RAM options
16 / 24 GB
Usable for models
~20 GB
Year
2024

Specs checked against manufacturer and public documentation on .

What runs on the ROG Phone 9 Pro

All 93 models at their recommended quant, on the 24GB configuration. Select any row for the full report.

ModelParamsQuantNeedsSpeedVerdict

~ = bandwidth-based estimate · ✓ = measured on real hardware

Best model by use case

Best for Chat

Top everyday assistant & writing pick here — ~57.6 tokens/s at Q8_0, using 1.3 of ~20GB.

Best for Coding

Top code completion & explain-this pick here — ~18.2 tokens/s at Q4_K_M, using 2.8 of ~20GB.

Best for Reasoning

Top math & step-by-step thinking pick here — ~31.4 tokens/s at Q4_K_M, using 1.9 of ~20GB.

FAQ

What is the biggest AI model the ROG Phone 9 Pro can run?

Bonsai 27B (1-bit) (27B parameters) at Q1_0 — it needs 4.8GB of the ~20GB usable on the 24GB ROG Phone 9 Pro and runs at ~9.1 tokens/s.

How much of the ROG Phone 9 Pro's 24GB RAM can AI models actually use?

About 20GB. Android keeps roughly 2–4GB for the system and resident apps, so of the 24GB about 20GB is actually available to a model.

Can the ROG Phone 9 Pro run Llama 3.1 8B?

Yes — at Q4_K_M it needs 6.3GB of the ~20GB usable and runs at ~7.1 tokens/s.

How fast is local AI on the ROG Phone 9 Pro?

The Snapdragon 8 Elite has 76.8GB/s of memory bandwidth, which is what decode speed scales with. Small models like Ternary Bonsai 1.7B reach ~69.1 tokens/s; larger 7–14B models land in the single digits. Anything above ~8 tokens/s feels smooth for chat.

Which quantization should I use on the ROG Phone 9 Pro?

Q4_K_M is the size/quality sweet spot for most models. For example, Qwen3 0.6B at Q8_0 takes 1.3GB of memory here. Only drop to Q3 or IQ4 if a model just misses fitting; Q8 rarely pays off on 24GB of RAM.

Is 24GB of RAM enough for local AI?

57 of the 93 models we track fit on the ROG Phone 9 Pro — 34 run great and 23 run with compromises. 36 models (mostly 12B+) don't fit at their recommended quant.