Can Fara 1.5 4B run on iPhone 15 Pro?
What this means
The checked Q4_K_M version should run, but replies are likely to arrive slowly. The smaller Q4_K_S version is the safer choice.
Download the Q4_K_S version, which is about 2.7 GB. We expect it to use about 3.8 GB of the roughly 5.2 GB available to a model on this phone.
At the estimated speed, a roughly 300-word answer may take about 47 seconds to finish.
This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.
Make it bearable
See how fast it feels
Estimated at 7.9 tokens/s — reading pace. A ~300-word reply takes about 51 seconds on this phone.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (2.9 GB) + mmap overhead | 3 |
| KV cache | 4K context window | 0.3 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 4 |
| Working budget | 8 GB RAM − iOS reserve (jetsam limit) | 5.2 |
| Headroom | remaining inside the working budget | 1.2 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K | 2.1 GB | ✓ Runs great | ~11 tokens/s |
| Q3_K_M | 2.4 GB | ✓ Runs great | ~9.6 tokens/s |
| IQ4_XS | 2.5 GB | ✓ Runs great | ~9.2 tokens/s |
| Q4_0 | 2.6 GB | ✓ Runs great | ~8.9 tokens/s |
| Q4_K_S BEST HERE | 2.7 GB | ✓ Runs great | ~8.5 tokens/s |
| Q4_K_M | 2.9 GB | ! Runs, barely | ~7.9 tokens/s |
| Q5_K_M | 3.3 GB | ! Runs, barely | ~7 tokens/s |
| Q6_K | 3.7 GB | ! Runs, barely | ~6.2 tokens/s |
| Q8_0 | 4.5 GB | ✕ Won't fit | won't fit |
Use the right workflow for this model
This model is designed to operate websites from screenshots as a computer-use agent. PocketPal does not provide the required a browser-control harness that repeatedly supplies screenshots and executes the model's actions, so the tracked file is not a beginner phone-chat path.
Read the publisher workflow ↗If you only want a normal private chat, use this simpler model instead:
Check Llama 3.2 3B