Can Agents-A1 4B run on iPhone 17e?
What this means
Our estimate says Agents-A1 4B should fit comfortably on your iPhone 17e.
Download the Q4_K_M version, which is about 2.7 GB. We expect it to use about 3.7 GB of the roughly 5.2 GB available to a model on this phone.
At the estimated speed, a roughly 300-word answer may take about 35 seconds to finish.
This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.
See how fast it feels
Estimated at 11.4 tokens/s — faster than you read. A ~300-word reply takes about 35 seconds on this phone.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (2.7 GB) + mmap overhead | 2.8 |
| KV cache | 4K context window | 0.3 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at Q4_K_M, 4K context | 3.7 |
| Working budget | 8 GB RAM − iOS reserve (jetsam limit) | 5.2 |
| Headroom | remaining inside the working budget | 1.5 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q4_K_M BEST HERE | 2.7 GB | ✓ Runs great | ~11.4 tokens/s |
Use the right workflow for this model
This model is designed to run long-horizon research, tool-calling, and agent workflows. PocketPal does not provide the required a compatible tool harness that can execute calls and return their results to the model, so the tracked file is not a beginner phone-chat path.
Read the publisher workflow ↗If you only want a normal private chat, use this simpler model instead:
Check Llama 3.2 3B