Can Ministral 8B run on vivo S2?
YES — Runs, barely
Formula estimate
What this means
Our estimate says Ministral 8B should fit, but there is very little memory left for the operating system and other apps.
Download the IQ4_XS version, which is about 4.4 GB. We expect it to use about 5.8 GB of the roughly 6 GB available to a model on this phone.
At the estimated speed, a roughly 300-word answer may take about 4 minutes to finish.
This result is calculated from the phone and model specifications. It has not been measured on this exact phone and setup.
1.7tokens/s
estimated · slower than you read
5.8GB
needed at IQ4_XS
IQ4_XS
quant checked
needs 5.8 GBusable 6 GB
Dimensity 7360-Turbo8 GB RAMCPU / GPU
Make it bearable
1
Drop context to 2K — saves ~0.3 GB of KV cache and a little speed.
2
Close every other app before loading — the 0.2 GB headroom is real; Android may kill the app otherwise.
3
Expect throttling after ~10 min of sustained generation on a phone chassis.
See how fast it feels
Estimated at 1.7 tokens/s — slower than you read. A ~300-word reply takes about 235 seconds on this phone.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | IQ4_XS GGUF (4.4 GB) + mmap overhead | 4.6 |
| KV cache | 4K context window | 0.6 |
| Runtime | llama.cpp + app overhead | 0.6 |
| Total needed | at IQ4_XS, 4K context | 5.8 |
| Working budget | 8 GB RAM − Android system reserve | 6 |
| Headroom | remaining inside the working budget | 0.2 |
Pick your quant
| Quant | Download | Verdict | Speed |
|---|---|---|---|
| Q2_K | 3.2 GB | ! Runs, barely | ~2.4 tokens/s |
| Q3_K_M | 4 GB | ! Runs, barely | ~1.9 tokens/s |
| IQ4_XS BEST HERE | 4.4 GB | ! Runs, barely | ~1.7 tokens/s |
| Q4_0 | 4.7 GB | ✕ Won't fit | won't fit |
| Q4_K_S | 4.7 GB | ✕ Won't fit | won't fit |
| Q4_K_M | 4.9 GB | ✕ Won't fit | won't fit |
| Q5_K_M | 5.7 GB | ✕ Won't fit | won't fit |
| Q6_K | 6.6 GB | ✕ Won't fit | won't fit |
| Q8_0 | 8.5 GB | ✕ Won't fit | won't fit |
Get your first offline chat working
Recommended app: PocketPal. Follow the point-and-click steps below. The speed above is an estimate, not a measurement from this exact app and phone.
1
Install or update PocketPal from Google Play. It is free and does not require an account. Use a current version so its loader supports newer model architectures.
2
Open the exact model. In PocketPal, go to Models → + → Add from Hugging Face, then paste
bartowski/Ministral-8B-Instruct-2410-GGUF.3
Choose the IQ4_XS GGUF file. The download is about 4.4 GB, so use Wi-Fi and keep the app open. Choose the main GGUF weights, not a vision projector, mmproj, or other helper file.
4
Tap Download, then Load. Start with a 4K (4096-token) context. Keep the app’s default Android backend for the first run.
5
Send a simple first prompt. Try “Explain why the sky is blue in three sentences.” This page estimates about 1.7 tokens/s, but that number is not a PocketPal measurement unless it carries a ✓ Verified label.
6
Confirm it is really offline. After the first reply, turn on airplane mode and ask a second question. If it still answers, the model is running on your phone.
If it does not work
- Model not listed: update PocketPal and paste the exact repository
bartowski/Ministral-8B-Instruct-2410-GGUF. - App closes while loading: close other apps, restart the phone, and try 2K context. If it still closes, choose a smaller model.
- No offline reply: confirm that the IQ4_XS GGUF file is loaded in the chat rather than a remote model.
Related checks
More on vivo S2
Ministral 8B on other phones