LIVE RELEASE TRACKERWEIGHTS NOT POSTED YETLAST VERIFIED · JUL 19, 2026

Qwen3.8 release date, VRAM requirements & local status

Short answer: Qwen3.8 has no verified official release, downloadable weights, GGUF, Ollama package, or trustworthy VRAM requirement yet. If you need a local Qwen today, run Qwen3.6 instead.

VERIFIED ANSWERS · JUL 19, 2026
RELEASED?
No official release
No announcement or model card can be verified.
VRAM?
Unknown
No parameter count or quant files exist to measure.
LOCAL / OLLAMA?
Not available
No official weights, GGUF, MLX, or Ollama package.
RUN NOW
Use Qwen3.6
Open 27B and 35B-A3B weights are available today.

Checked against the official Qwen release site and Qwen Hugging Face organization · no guessed parameter count · no invented quant sizes

NEED A LOCAL QWEN NOW?
Qwen3.6 has downloadable weights and quants today.
Compare 27B vs 35B-A3B, Q2–Q8 memory, 8–48GB hardware and context costs.
See Qwen3.6 requirements →
WHAT THIS TRACKER WILL UPDATE
01 · Model card
Confirmed variants, architecture, context and license.
02 · Weights
Actual checkpoint bytes—never a rumored parameter estimate.
03 · Quants
Real GGUF Q2–Q8 sizes plus 16/24/32GB fit guidance.
04 · Runtimes
Working Ollama, llama.cpp, LM Studio, MLX and vLLM commands.
05 · Measurements
Reproducible speed results labeled ✓ Verified.

Update log

JUL 19, 2026Tracker opened

Checked Qwen's official release site and Hugging Face organization; no Qwen3.8 model card or checkpoint is published yet. Next update: the first official announcement, model card or weight repository.

Quick answers

QuestionVerified answerWhat is missing
Official releaseWatchingQwen3.8 model card has not landed yet
Open weightsWaitingOfficial checkpoint not posted yet
Confirmed variantsWaitingWill populate from the official model family
GGUF / Q4 / Q8QueuedActual file sizes will replace these placeholders
OllamaQueuedVerified library package and command will appear here
llama.cpp / LM StudioPendingArchitecture and loader support unknown
MLX / vLLMPendingOfficial config and runtime support unknown

Qwen3.8 memory by quant

These cells stay pending until a real checkpoint and quant files exist. We will report the actual download, loaded weights and total memory separately—a GGUF file size is not the same thing as VRAM required.

QuantActual downloadTotal @4KSafer memory
Q2 / IQ2PendingPendingPending
Q3 / IQ3PendingPendingPending
Q4_K_MPendingPendingPending
Q5_K_MPendingPendingPending
Q6_KPendingPendingPending
Q8_0PendingPendingPending
FP8 / BF16PendingPendingPending

Planned total = loaded weights + architecture-aware KV cache + runtime/vision buffers + stated headroom.

Will Qwen3.8 run on 16GB or 24GB VRAM?

No reliable verdict is possible without confirmed variants. A model family can contain a server-scale MoE and much smaller dense checkpoints; using a rumored flagship size to answer for the whole family would be misleading. When files arrive, we will test 8, 12, 16, 24, 32, 48, 64 and 80GB memory tiers at 4K, 32K and longer context.

WHY ONE “VRAM” NUMBER IS NOT ENOUGH
Download size — bytes stored on disk.
Loaded weights — memory used by the model parameters.
KV cache — grows with context and depends on the real attention configuration.
Runtime and headroom — buffers and allocation space needed to avoid loading successfully and then crashing.

How to run Qwen3.8 locally

This section is staged for launch day. Once supported files exist, it will provide versioned commands for Ollama, llama.cpp, LM Studio, MLX and vLLM. Until then, verify any similarly named upload against an official Qwen base checkpoint before downloading it.

RuntimeStatusWhat will appear here
OllamaUnavailableVerified model tag + pull/run command
llama.cppPendingMinimum supported build + GGUF command
LM StudioPendingCompatible version + recommended load settings
MLXPendingOfficial/community conversion + Apple memory guidance
vLLMPendingSupported version + serve command + context settings

Qwen3.8 vs Qwen3.6: should you wait?

If you need local inference now, use a model with published weights and runtime support. Qwen3.6 can be evaluated today; Qwen3.8 cannot. A capability or speed comparison before Qwen3.8 has an official model card would be invented rather than measured.

DecisionQwen3.8Qwen3.6
Run locally todayNoYes — open checkpoints exist
Choose an exact quantNo files yetAvailable
Calculate real memoryPendingAvailable
Compare measured speedNo verified dataPossible with controlled tests
Best choice nowWait for primary sourcesUse if it fits your hardware

Need numbers for the released generation? See our Qwen3.6 VRAM and RAM requirements.

Can a phone run Qwen3.8?

The answer is pending the actual model family. If Qwen publishes small open-weight variants, we will evaluate each one separately against usable phone RAM, app support and sustained performance. A server-scale flagship and a phone-sized sibling must not share one verdict.

Want a local model that runs now?

Pick your exact phone and AICanRun will show released Qwen and other models that fit, with the recommended quant.

Check my phone →

FAQ

How much VRAM does Qwen3.8 need?

There is no verified Qwen3.8 VRAM figure yet. Qwen has not published an official Qwen3.8 model card, parameter count, checkpoint, or GGUF file. Any precise Q4 or Q8 requirement published before those files exist is speculation.

Has Qwen3.8 been released?

No official Qwen3.8 release can be verified as of July 19, 2026. Qwen's official site currently documents Qwen3.7 as the latest named mainline release, and the official Qwen Hugging Face organization does not list a Qwen3.8 checkpoint.

Can I run Qwen3.8 locally?

Not today. Local inference requires downloadable weights plus support from a runtime such as llama.cpp, Ollama, MLX, or vLLM. No official Qwen3.8 weights or supported local package are currently available.

Is Qwen3.8 available in Ollama or GGUF?

No verified official or community Qwen3.8 GGUF or Ollama package is available yet. A similarly named upload is not enough: its base checkpoint and provenance must match an official Qwen release.

Can Qwen3.8 run on 16GB or 24GB of VRAM?

That depends on which open-weight variants Qwen releases and their real quantized file sizes. We will not infer a fit from an unconfirmed flagship parameter rumor. Once files exist, this page will separate download size, loaded weights, KV cache, runtime overhead, and safer memory.

Can a phone run Qwen3.8?

No phone can run Qwen3.8 locally today because there is no official local checkpoint. A future small open-weight variant could be phone-sized, but Qwen has not confirmed one. Flagship and small-family variants must be evaluated separately.

Should I wait for Qwen3.8 or run Qwen3.6 now?

Use Qwen3.6 if you need a local model now: its open checkpoints and quantized files can be evaluated today. Wait for Qwen3.8 only if you are comfortable waiting for official weights, runtime support, and independent measurements.