Qwen3.8 release date, VRAM requirements & local status
Short answer: Qwen3.8 has no verified official release, downloadable weights, GGUF, Ollama package, or trustworthy VRAM requirement yet. If you need a local Qwen today, run Qwen3.6 instead.
Checked against the official Qwen release site and Qwen Hugging Face organization · no guessed parameter count · no invented quant sizes
Update log
Checked Qwen's official release site and Hugging Face organization; no Qwen3.8 model card or checkpoint is published yet. Next update: the first official announcement, model card or weight repository.
Quick answers
| Question | Verified answer | What is missing |
|---|---|---|
| Official release | Watching | Qwen3.8 model card has not landed yet |
| Open weights | Waiting | Official checkpoint not posted yet |
| Confirmed variants | Waiting | Will populate from the official model family |
| GGUF / Q4 / Q8 | Queued | Actual file sizes will replace these placeholders |
| Ollama | Queued | Verified library package and command will appear here |
| llama.cpp / LM Studio | Pending | Architecture and loader support unknown |
| MLX / vLLM | Pending | Official config and runtime support unknown |
Qwen3.8 memory by quant
These cells stay pending until a real checkpoint and quant files exist. We will report the actual download, loaded weights and total memory separately—a GGUF file size is not the same thing as VRAM required.
| Quant | Actual download | Total @4K | Safer memory |
|---|---|---|---|
| Q2 / IQ2 | Pending | Pending | Pending |
| Q3 / IQ3 | Pending | Pending | Pending |
| Q4_K_M | Pending | Pending | Pending |
| Q5_K_M | Pending | Pending | Pending |
| Q6_K | Pending | Pending | Pending |
| Q8_0 | Pending | Pending | Pending |
| FP8 / BF16 | Pending | Pending | Pending |
Planned total = loaded weights + architecture-aware KV cache + runtime/vision buffers + stated headroom.
Will Qwen3.8 run on 16GB or 24GB VRAM?
No reliable verdict is possible without confirmed variants. A model family can contain a server-scale MoE and much smaller dense checkpoints; using a rumored flagship size to answer for the whole family would be misleading. When files arrive, we will test 8, 12, 16, 24, 32, 48, 64 and 80GB memory tiers at 4K, 32K and longer context.
How to run Qwen3.8 locally
This section is staged for launch day. Once supported files exist, it will provide versioned commands for Ollama, llama.cpp, LM Studio, MLX and vLLM. Until then, verify any similarly named upload against an official Qwen base checkpoint before downloading it.
| Runtime | Status | What will appear here |
|---|---|---|
| Ollama | Unavailable | Verified model tag + pull/run command |
| llama.cpp | Pending | Minimum supported build + GGUF command |
| LM Studio | Pending | Compatible version + recommended load settings |
| MLX | Pending | Official/community conversion + Apple memory guidance |
| vLLM | Pending | Supported version + serve command + context settings |
Qwen3.8 vs Qwen3.6: should you wait?
If you need local inference now, use a model with published weights and runtime support. Qwen3.6 can be evaluated today; Qwen3.8 cannot. A capability or speed comparison before Qwen3.8 has an official model card would be invented rather than measured.
| Decision | Qwen3.8 | Qwen3.6 |
|---|---|---|
| Run locally today | No | Yes — open checkpoints exist |
| Choose an exact quant | No files yet | Available |
| Calculate real memory | Pending | Available |
| Compare measured speed | No verified data | Possible with controlled tests |
| Best choice now | Wait for primary sources | Use if it fits your hardware |
Need numbers for the released generation? See our Qwen3.6 VRAM and RAM requirements.
Can a phone run Qwen3.8?
The answer is pending the actual model family. If Qwen publishes small open-weight variants, we will evaluate each one separately against usable phone RAM, app support and sustained performance. A server-scale flagship and a phone-sized sibling must not share one verdict.
Pick your exact phone and AICanRun will show released Qwen and other models that fit, with the recommended quant.
Check my phone →FAQ
How much VRAM does Qwen3.8 need?
There is no verified Qwen3.8 VRAM figure yet. Qwen has not published an official Qwen3.8 model card, parameter count, checkpoint, or GGUF file. Any precise Q4 or Q8 requirement published before those files exist is speculation.
Has Qwen3.8 been released?
No official Qwen3.8 release can be verified as of July 19, 2026. Qwen's official site currently documents Qwen3.7 as the latest named mainline release, and the official Qwen Hugging Face organization does not list a Qwen3.8 checkpoint.
Can I run Qwen3.8 locally?
Not today. Local inference requires downloadable weights plus support from a runtime such as llama.cpp, Ollama, MLX, or vLLM. No official Qwen3.8 weights or supported local package are currently available.
Is Qwen3.8 available in Ollama or GGUF?
No verified official or community Qwen3.8 GGUF or Ollama package is available yet. A similarly named upload is not enough: its base checkpoint and provenance must match an official Qwen release.
Can Qwen3.8 run on 16GB or 24GB of VRAM?
That depends on which open-weight variants Qwen releases and their real quantized file sizes. We will not infer a fit from an unconfirmed flagship parameter rumor. Once files exist, this page will separate download size, loaded weights, KV cache, runtime overhead, and safer memory.
Can a phone run Qwen3.8?
No phone can run Qwen3.8 locally today because there is no official local checkpoint. A future small open-weight variant could be phone-sized, but Qwen has not confirmed one. Flagship and small-family variants must be evaluated separately.
Should I wait for Qwen3.8 or run Qwen3.6 now?
Use Qwen3.6 if you need a local model now: its open checkpoints and quantized files can be evaluated today. Wait for Qwen3.8 only if you are comfortable waiting for official weights, runtime support, and independent measurements.