Can OneJev 4B run on iPhone 15?

Estimated weights fit · custom workflow required

Weight fit is not phone app support

The tracked Q3_K_M weights are 2.5GB; the generic working-memory estimate is 3.5GB against 3.9GB usable. Image inputs need additional projector memory.

OneJev scores typed decisions in a single forward pass, not chat replies. The decode estimates below are not measured decision latency. The publisher's qev server and System One client are required; AICanRun has not verified a phone app route.

Not measured
decision latency
3.5GB
needed at Q3_K_M
Q3_K_M
quant checked
needs 3.5 GBusable 3.9 GB
Apple A16 Bionic6 GB RAMMetal

Make it bearable

1
Drop to Q2_K (2.1 GB) — it runs comfortably here at ~11 tokens/s. Recommended over Q3_K_M.
2
Drop context to 2K — saves ~0.2 GB of KV cache and a little speed.
3
Close every other app before loading — the 0.4 GB headroom is real; iOS may evict the model otherwise.
4
Expect throttling after ~10 min of sustained generation on a phone chassis.

Where the memory goes

ComponentDetailGB
Model weightsQ3_K_M GGUF (2.5 GB) + mmap overhead2.6
KV cache4K context window0.3
Runtimellama.cpp + app overhead0.6
Total neededat Q3_K_M, 4K context3.5
Working budget6 GB RAM − iOS reserve (jetsam limit)3.9
Headroomremaining inside the working budget0.4

Pick your quant

QuantDownloadVerdictSpeed
Q2_K BEST HERE2.1 GB✓ Runs great~11 tokens/s
Q3_K_M 2.5 GB! Runs, barely~9.2 tokens/s
IQ4_XS 2.9 GB✕ Won't fitwon't fit
Q4_K_S 2.9 GB✕ Won't fitwon't fit
Q4_K_M 3.1 GB✕ Won't fitwon't fit
Q5_K_M 3.5 GB✕ Won't fitwon't fit
Q6_K 4 GB✕ Won't fitwon't fit
Q8_0 5.2 GB✕ Won't fitwon't fit

Use the right workflow for this model

The model fits in memory, but it is not an ordinary chat model.

This model is designed to return calibrated probabilities for typed choices from text or images. PocketPal does not provide the required the publisher's qev server and System One client; image inputs also need a vision projector, so the tracked file is not a beginner phone-chat path.

Read the publisher workflow ↗

If you only want a normal private chat, use this simpler model instead:

Check Llama 3.2 3B

Related checks

More on iPhone 15
Qwen3 0.6BLlama 3.2 1BGemma 3 1BDeepSeek R1 Distill 1.5BSmolLM2 1.7B
OneJev 4B on other phones
Xiaomi 18 FoldOPPO Find X8 Provivo X200 ProGalaxy S26Galaxy S26+