Can Fara 1.5 9B run on HONOR Robot Phone?
What this means
The checked Q4_K_M version should run, but replies are likely to arrive slowly. The smaller Q2_K version is the safer choice.
Download the Q2_K version, which is about 4.1 GB. We expect it to use about 5.6 GB of the roughly 12 GB available to a model on this phone.
At the estimated speed, a roughly 300-word answer may take about 43 seconds to finish.
This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.
Make it bearable
See how fast it feels
Estimated at 6.5 tokens/s — reading pace. A ~300-word reply takes about 62 seconds on this phone.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (5.9 GB) + mmap overhead | 6.2 |
| KV cache | 4K context window | 0.7 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 7.5 |
| Working budget | 16 GB RAM − Android system reserve | 12 |
| Headroom | remaining inside the working budget | 4.5 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K BEST HERE | 4.1 GB | ✓ Runs great | ~9.3 tokens/s |
| Q3_K_M | 4.9 GB | ! Runs, barely | ~7.8 tokens/s |
| IQ4_XS | 5.2 GB | ! Runs, barely | ~7.3 tokens/s |
| Q4_0 | 5.5 GB | ! Runs, barely | ~6.9 tokens/s |
| Q4_K_S | 5.6 GB | ! Runs, barely | ~6.8 tokens/s |
| Q4_K_M | 5.9 GB | ! Runs, barely | ~6.5 tokens/s |
| Q5_K_M | 6.9 GB | ! Runs, barely | ~5.5 tokens/s |
| Q6_K | 7.7 GB | ! Runs, barely | ~5 tokens/s |
| Q8_0 | 9.5 GB | ! Runs, barely | ~4 tokens/s |
Use the right workflow for this model
This model is designed to operate websites from screenshots as a computer-use agent. PocketPal does not provide the required a browser-control harness that repeatedly supplies screenshots and executes the model's actions, so the tracked file is not a beginner phone-chat path.
Read the publisher workflow ↗If you only want a normal private chat, use this simpler model instead:
Check Llama 3.2 3B