Can OneJev 9B run on Motorola Edge 50 Ultra?
Estimated weights fit · custom workflow required
Weight fit is not phone app support
The tracked Q4_K_M weights are 5.6GB; the generic working-memory estimate is 7.1GB against 12GB usable. Image inputs need additional projector memory.
OneJev scores typed decisions in a single forward pass, not chat replies. The decode estimates below are not measured decision latency. The publisher's qev server and System One client are required; AICanRun has not verified a phone app route.
Not measured
decision latency
7.1GB
needed at Q4_K_M
Q4_K_M
quant checked
needs 7.1 GBusable 12 GB
Snapdragon 8s Gen 316 GB RAMCPU / GPUAlso sold with 12 GB
Make it bearable
1
Drop to Q2_K (3.8 GB) — it runs comfortably here at ~8 tokens/s. Recommended over Q4_K_M.
2
Drop context to 2K — saves ~0.3 GB of KV cache and a little speed.
3
Close every other app before loading — the 4.9 GB headroom is real; Android may kill the app otherwise.
4
Expect throttling after ~10 min of sustained generation on a phone chassis.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (5.6 GB) + mmap overhead | 5.9 |
| KV cache | 4K context window | 0.6 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 7.1 |
| Working budget | 16 GB RAM − Android system reserve | 12 |
| Headroom | remaining inside the working budget | 4.9 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K BEST HERE | 3.8 GB | ✓ Runs great | ~8 tokens/s |
| Q3_K_M | 4.6 GB | ! Runs, barely | ~6.6 tokens/s |
| IQ4_XS | 5.2 GB | ! Runs, barely | ~5.8 tokens/s |
| Q4_K_S | 5.4 GB | ! Runs, barely | ~5.6 tokens/s |
| Q4_K_M | 5.6 GB | ! Runs, barely | ~5.4 tokens/s |
| Q5_K_M | 6.5 GB | ! Runs, barely | ~4.7 tokens/s |
| Q6_K | 7.4 GB | ! Runs, barely | ~4.1 tokens/s |
| Q8_0 | 9.5 GB | ! Runs, barely | ~3.2 tokens/s |
Use the right workflow for this model
The model fits in memory, but it is not an ordinary chat model.
This model is designed to return calibrated probabilities for typed choices from text or images. PocketPal does not provide the required the publisher's qev server and System One client; image inputs also need a vision projector, so the tracked file is not a beginner phone-chat path.
Read the publisher workflow ↗If you only want a normal private chat, use this simpler model instead:
Check Llama 3.2 3BRelated checks
More on Motorola Edge 50 Ultra
OneJev 9B on other phones