Can OneJev 4B run on MacBook Pro M3 Max 36GB?
What this means
OneJev 4B fits on the MacBook Pro M3 Max · 36GB at Q4_K_M. We estimate it uses 4.3GB of the conservative 32.4GB working budget.
OneJev scores typed decisions in one forward pass. Generic decode estimates are not decision latency; image inputs need additional projector memory. Use the publisher's qev server and System One client.
Where the memory goes
| Component | Detail | GB |
|---|---|---|
| Model weights | Q4_K_M GGUF (3.1 GB) + mmap overhead | 3.3 |
| KV cache | 4K context window | 0.3 |
| Runtime | macOS inference app + compute buffers | 0.8 |
| Total needed | at Q4_K_M, 4K context | 4.3 |
| Working budget | 36 GB unified memory − conservative macOS reserve | 32.4 |
| Headroom | remaining inside the working budget | 28.1 |
Pick your quant
| Quant | Download | Memory | Estimated speed | Verdict |
|---|---|---|---|---|
| Q2_K | 2.1 GB | 3.3 GB | ~71.4–100 tok/s | ✓ Runs great |
| Q3_K_M | 2.5 GB | 3.7 GB | ~60–84 tok/s | ✓ Runs great |
| IQ4_XS | 2.9 GB | 4.1 GB | ~51.7–72.4 tok/s | ✓ Runs great |
| Q4_K_S | 2.9 GB | 4.1 GB | ~51.7–72.4 tok/s | ✓ Runs great |
| Q4_K_M ★ | 3.1 GB | 4.3 GB | ~48.4–67.7 tok/s | ✓ Runs great |
| Q5_K_M | 3.5 GB | 4.8 GB | ~42.9–60 tok/s | ✓ Runs great |
| Q6_K | 4 GB | 5.3 GB | ~37.5–52.5 tok/s | ✓ Runs great |
| Q8_0 | 5.2 GB | 6.5 GB | ~28.8–40.4 tok/s | ✓ Runs great |
Use this model for its intended task
OneJev 4B is designed to return calibrated probabilities for typed choices from text or images. This Mac has enough working memory for the tracked Q4_K_M weights, but loading the file in a plain chat window does not supply the inputs, tools, or action loop that make the model useful.
It documents the intended workflow and supported developer runtimes. This is not yet a point-and-click setup for beginners, so do not install LM Studio expecting an ordinary chat tutorial to reproduce the model's task.
Choose Llama 3.2 3B instead. Its page gives the complete graphical install, download, first prompt, and offline check.
This model is designed to return calibrated probabilities for typed choices from text or images. Its weights may fit in memory, but it needs the publisher's qev server and System One client; image inputs also need a vision projector. A normal LM Studio chat does not provide that workflow.
Other models on this Mac
OneJev 4B on other MacBook Pro M3 Max configurations
FAQ
Can the MacBook Pro M3 Max · 36GB run OneJev 4B?
The tracked Q4_K_M working memory is estimated at 4.3GB, excluding the vision projector. Weight fit is not verified workflow support; use the publisher's qev server and System One client.
Which OneJev 4B quant should I use on this Mac?
Q4_K_M. It is a 3.1GB download and leaves about 28.1GB inside our conservative working budget.
Which app should I use for OneJev 4B on this Mac?
OneJev 4B is a task-specific model, not a normal local chat download. The selected weights fit this Mac, but the model still needs its publisher's intended workflow. This page links that repository and recommends Llama 3.2 3B if you just want local chat.
Are these speeds measured on a MacBook Pro M3 Max?
Decision latency has not been measured. OneJev scores typed choices in one forward pass; generic decode estimates do not measure this workflow.