Can OneJev 27B run on ROG Phone 9 Pro?
Estimated weights fit · custom workflow required
Weight fit is not phone app support
The tracked Q4_K_M weights are 16.5GB; the generic working-memory estimate is 19.8GB against 20GB usable. Image inputs need additional projector memory.
OneJev scores typed decisions in a single forward pass, not chat replies. The decode estimates below are not measured decision latency. The publisher's qev server and System One client are required; AICanRun has not verified a phone app route.
Make it bearable
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (16.5 GB) + mmap overhead | 17.3 |
| KV cache | 4K context window | 1.9 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 19.8 |
| Working budget | 24 GB RAM − Android system reserve | 20 |
| Headroom | remaining inside the working budget | 0.2 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K | 10.7 GB | ! Runs, barely | ~3.2 tokens/s |
| Q3_K_M | 13.3 GB | ! Runs, barely | ~2.6 tokens/s |
| IQ4_XS | 15.2 GB | ! Runs, barely | ~2.3 tokens/s |
| Q4_K_S | 15.6 GB | ! Runs, barely | ~2.2 tokens/s |
| Q4_K_M BEST HERE | 16.5 GB | ! Runs, barely | ~2.1 tokens/s |
| Q5_K_M | 19.2 GB | ✕ Won't fit | won't fit |
| Q6_K | 22.1 GB | ✕ Won't fit | won't fit |
| Q8_0 | 28.6 GB | ✕ Won't fit | won't fit |
Use the right workflow for this model
This model is designed to return calibrated probabilities for typed choices from text or images. PocketPal does not provide the required the publisher's qev server and System One client; image inputs also need a vision projector, so the tracked file is not a beginner phone-chat path.
Read the publisher workflow ↗If you only want a normal private chat, use this simpler model instead:
Check Llama 3.2 3B