Qwen3.8 VRAM: the 27B That Ties GPT-5.6 Fits 24GB

Qwen3.8 VRAM Requirements: Both Models Now Measured

Qwen3.8-27B weights went public at 15:00 UTC on 14 August 2026, two days after the 2.4T, and on Alibaba’s own comparison table it trades blows with Opus 4.6 Max. Both halves of this page are now measurement rather than projection, and the 27B figures below come from unsloth’s published GGUF files rather than from arithmetic. Unsloth replaced its whole file ladder on 20 August, and its Q4_K_M is now 16.46 GB. On the card that comes to 15.8 GB at 8K, so a 24 GB card holds it at 64K with a full-size cache and the whole 262K window with a 4-bit cache. We have since run it ourselves on a 16 GB Tesla T4 and a 24 GB L4, and the measurements are below. The 2.4T needs about 1,369 GB at 8K and remains datacenter hardware.

This page carried a projection for the 27B for eleven days, labeled as one. The projection said 64 layers, 16 full attention plus 48 linear, 262K context, and 16.4 GB of weights at 4-bit. The released config says 64 layers, 16 full attention plus 48 linear, 262,144 context, and the real file is 17.11 GB. The architecture was exactly right and the size was 4% light. Both numbers stay on the page, because a projection that survives contact with the real file is the only evidence that the next one is worth reading.

Qwen3.8-27B: What Alibaba Actually Shipped

Parameters27,781,427,952
Architecturedense, 64 layers, native multimodal
Attention split16 full attention at every fourth layer, 48 linear attention
Vision27-layer tower, ships as a separate 0.93 GB mmproj file in GGUF builds
KV heads / head dim4 / 256
Native precisionBF16, 55.56 GB
Context262,144 native, extendable to 1M with YaRN
LicenseApache 2.0

Two things here are easy to miss and both change the VRAM answer. The first is that this is a multimodal model, so image and video input need a vision tower that llama.cpp ships as its own file. Quote a text-only quant size for a model you are describing as multimodal and you are about a gigabyte short. The second is that 262,144 is the native context, not an extended one: rope_scaling is null in the shipped config and rope_type is default. The 1M figure Alibaba quotes is YaRN applied by you, after the fact.

The 48 linear-attention layers are what make the rest possible. They hold a fixed recurrent state that does not grow as context grows, so only the 16 full-attention layers pay for a longer prompt. Treating all 64 layers as full attention, which is what a flat layers times heads times dimension formula does, returns 68.7 GB of KV cache at the full 262K context. The real figure is 17.34 GB. That gap is the difference between a 3090 owner trying this model and being told to give up.

Measured VRAM: Qwen3.8-27B by Quantization

Files are Unsloth’s current GGUF builds, read 3 October 2026. “On the card” is what our calculator gives at an 8K context with an F16 KV cache on llama.cpp, text only: llama.cpp keeps the token embedding table in system RAM, which is why the card holds less than the file. Add 0.93 GB to any row for image or video input.

QuantizationFileOn the card at 8K24 GB card
Q8_029.05 GB27.2 GBno
Q6_K21.98 GB20.8 GByes, 87%
Q5_K_M19.77 GB18.8 GByes, 78%
Q4_K_M16.46 GB15.8 GByes, 66%
Q3_K_XL13.15 GB12.8 GByes, 53%
Q2_K_XL9.83 GB9.8 GByes, 41%

A 4090, a 3090 and a 7900 XTX all hold this at 4-bit with working room, and Q6_K now fits a 24 GB card too, at 87% with an 8K context. Earlier versions of this page called Q6_K just over the line; Unsloth’s new Q6_K file is 0.9 GB smaller and our calculator now counts what llama.cpp really puts on the card, measured on real runs. On a card that also drives your desktop, Q5_K_M still leaves the most room.

At the bottom of the ladder, size from the file rather than a formula. Unsloth’s 2-bit build, UD-Q2_K_XL, is 9.83 GB, about 8% more than Q2_K’s nominal 2.625 bits per weight gives for 27.8 billion parameters, because its dynamic quants keep some layers at higher precision.

How Much Context Fits on a 24 GB Card

This is where the architecture pays. At Q4_K_M with an F16 cache, from our calculator:

ContextKV cacheText onlyWith vision24 GB card
8K0.65 GB15.81 GB16.74 GByes
32K2.15 GB17.35 GB18.28 GByes
64K4.15 GB19.40 GB20.33 GByes
128K8.15 GB23.50 GB24.43 GBtight, needs a card with no screen
256K, full16.15 GB31.71 GB32.64 GBno

64K is comfortable on a 24 GB card at default settings, with the vision file loaded too. 128K with an F16 cache clears 24 GB only on a card with nothing else on it; on our 24 GB L4, which reports 22.5 GB, it ran out of memory while loading. A quantized cache fixes that: llama.cpp’s -ctk q8_0 -ctv q8_0 brings 128K down to 20.21 GB (21.14 with vision), and q4_0 holds the full 262,144-token window at 21.09 GB (22.02 with vision). We ran both on the L4: 19.67 GB and 20.55 GB at the peak, with the window filled. Two testers have run the full window on 24 GB cards as well, at 23.00 GB on a 4090 and about 22.2 GB on a 3090; see the measured section below.

Issue 1 of The Local LLM Sizing Guide

The Local LLM Sizing Guide. Issue 1, 30 September

The table above is one model on one card. The guide does it for every model, on every card.

Qwen3.8 27B on a 16 GB card: Q3_K_M, then 16K of context with an F16 cache and 64K with Q4_0. Every model on every tier from 8 GB to 128 GB is laid out like that.

Know what fits before you buy

From $9 a month. Card or crypto. The next issue included.

We Ran It: Qwen3.8 27B on a 16 GB T4 and a 24 GB L4

On 2 October 2026 we ran Unsloth’s files on two cards ourselves, on Linux with NVIDIA driver 580.82.07 and llama.cpp build b11062, one request at a time with flash attention on. Each setup loaded the model, then read one prompt filling 90% of its window and wrote 64 tokens, and we kept the highest memory the card reported during the run. A Tesla T4 is sold as 16 GB and reports 15.0 GB with error correction on; an L4 is sold as 24 GB and reports 22.5 GB.

CardFileContext, cacheCard usedWe sayResult
T4 16 GBUD-Q3_K_XL8K, F1612.31 GB12.82 GBran
T4 16 GBUD-Q3_K_XL16K, F1612.82 GB13.34 GBran
T4 16 GBUD-Q3_K_XL32K, q8_012.99 GB13.54 GBran
T4 16 GBUD-Q3_K_XL32K, q4_012.49 GB13.03 GBran
T4 16 GBUD-Q3_K_XL64K, q4_013.20 GB13.76 GBran
L4 24 GBUD-Q4_K_M32K, F1616.87 GB17.35 GBran
L4 24 GBUD-Q4_K_M64K, F1618.90 GB19.40 GBran
L4 24 GBUD-Q4_K_M128K, q8_019.67 GB20.21 GBran
L4 24 GBUD-Q4_K_M256K, q4_020.55 GB21.09 GBran
L4 24 GBUD-Q4_K_M128K, F16n/a23.50 GBout of memory at load

Two things are settled by this. A 16 GB card runs the 27B: at 3-bit it held a 64K window with a 4-bit cache in 13.20 GB. And a 24 GB card holds the whole 262K window at 4-bit with a 4-bit cache, in 20.55 GB, with room to spare. Our calculator read between 0.48 and 0.56 GB above what the card used on every run, never below, and called the one setup that failed too tight. Every run, on this and other models, is in our VRAM calculator accuracy record.

Filling the window costs time as well as memory. On the L4 the prompt was read at 647 tok/s at 29K tokens and 217 tok/s at 236K, so the full window took 18 minutes before the first word; the speed page has the rest.

A 27B Model at Frontier Level, on a Card You Own

QWEN 3.8 27B Benchmarks opus 4.6

Alibaba put Qwen3.8-27B next to Opus 4.6 Max. Across the nine benchmarks where both carry a score, the 27B leads five. On coding and agent work it is ahead or within a couple of points nearly everywhere. Opus keeps its clearest lead on HLE, broad multidisciplinary reasoning, by 9.2.

That is frontier-class output from a 16.46 GB file that runs on a used RTX 3090. Not long ago that level of work existed only behind an API, metered per token, in a datacenter belonging to someone else.

That reading is no longer only Alibaba’s. On 17 August 2026, Artificial Analysis scored Qwen3.8 27B at 52 on its Intelligence Index. GPT-5.6 Luna (max) scores 52 on the same index. The 27B also ranks first of 135 open-weight models in the 4B to 40B class.

The index is a composite of nine evaluations: GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR.

The distinction matters more than the number. Everything above this paragraph comes from Alibaba’s own comparison table, which is a vendor publishing its own results. This is an independent evaluator running its own suite and landing in the same place. A 27B file you can hold on one consumer card now sits level with a metered frontier API on a third party’s leaderboard.

One honest asterisk, from the same source. Artificial Analysis notes that Qwen3.8 27B produced 160 million output tokens across the evaluation against a 43 million median, and calls it very verbose. GPT-5.6 Luna (max) produced 130 million on the same runs, so against the model it ties with the gap is about 23%, not four times. Locally that verbosity is not a bill, it is decode time, and the speed page has the tok/s to convert it.

One split is worth knowing before you download it. Against Alibaba’s own larger Qwen3.7-Plus, the 27B wins ten of twelve rows and loses only GPQA Diamond and HLE. Procedural skill compresses into 27B. Broad recall does not. For coding, agents and long office tasks this is the top of what you can run at home. For obscure factual recall it is still a 27B, which is what retrieval is for.

What Running It Locally Actually Buys

No metering, no rate limit, no queue, no waiting for capacity, and nothing leaving your machine. Apache 2.0, so the weights are yours to keep, fine-tune or modify, and community builds that remove the built-in refusals already exist. Whatever you make of that, it is a kind of control that only exists while the model sits on your own disk.

The limit on local AI has always been hardware rather than ideas. Confidential work that could not leave a building. Agents that need to run all night without a token bill. Anything where cost per call mattered more than the last few points of quality. All of it was waiting for a model this capable to fit in this much memory, and at 17 GB it now does.

Every figure above is Alibaba’s own, published at launch.

What People Are Measuring on Real Cards

Numbers below were published by the people who ran them, within hours of release. We did not measure any of this. Each figure is attributed to the person who posted it.

HardwareSetupResultSource
RTX 4090, 24 GBUD-Q4_K_XL, q4_0 KV, 260K context23.00 GB, 40.7 tok/s decode@analogalok
RTX 4090, 24 GBsame, with MTP65 tok/s decode@analogalok
RTX 3090, 24 GBQ4_K_M, q4_0 KV, full 262Kabout 22.2 GB, 25.4 tok/s@sudoingX
2x RTX 3090FP8, NVLink, MTP75 to 85 tok/s@Tech2Wild
RTX 5090NVFP4 plus DSpark, SGLang206.1 tok/s decodeSGLang, quoted by Qwen
DGX SparkSGLang38.28 tok/s decodeSGLang

The context ladder on a 24 GB card, measured. @analogalok published the whole thing on a 4090: 100K is the ceiling with an unquantized F16 cache at 23.59 GB, 170K fits with a q8_0 cache at 23.68 GB, and 260K fits with a q4_0 cache at 23.00 GB with no system RAM offload. That matches the shape of the table above and settles the question of what the flags buy you.

Our figures used to run 1 to 2 GB above those measurements. Against five of @analogalok’s configs the gap averaged 1.7 GB, always on the safe side. Since 2 October the calculator is tuned against our own runs of this model, above, and reads about half a gigabyte over what the card used. His files are Unsloth’s UD-Q4_K_XL, 1.1 GB larger than the Q4_K_M in our tables, so compare file to file.

Two other things worth repeating from people who ran it. @sudoingX reports decode staying flat as context fills, 25 tok/s empty and 26 at 90K deep, which is the linear-attention design doing what it is supposed to. And mainline llama.cpp still ignores the model’s MTP tensors (our own runs on build b11062 logged them as unused), so the MTP numbers above come from other stacks. The cafe-llama.cpp fork wires them up; its author reports Qwen 3.8 27B going from 40 to 80 tok/s with them.

Dense model, so it wants a GPU. @TheAhmadOsman makes the point that unified-memory machines, DGX Spark and Apple silicon, suit MoE models better. A 27B dense model will fit on those and run slowly. This one is aimed at 3090s, 4090s, 5090s and similar.

Qwen3.8-2.4T-A95B: The Other Half of the Launch

Total parameters2.4 trillion
Active parameters95 billion
ArchitectureMoE, 512 experts, 10 routed per token
Layers92: 23 full attention + 69 linear attention
Native precisionFP8 (also released BF16 reference)
Context256K config base

Same family, same trick, different scale. At Q4_K_M the weights alone are about 1,355 GB on the card, the 69 linear layers add a fixed 0.60 GB, and only the 23 full-attention layers scale with context: 1,369 GB total at 8K, 1,392 GB at 256K, 1,462 GB at 1M.

What this means in hardware terms: the smallest pools in our catalog that hold it at 8K are 1,728 GB of accelerators: four MI430X, six MI350X or six B300. This is a datacenter-class model at full speed. There is no consumer card that runs it, and there will not be.

But the same sentence was true of Kimi K3, and a 1-bit community build shipped 48 hours after its weights landed. Offloading, streaming weights from disk, and aggressive quantization are how models this size get run on far less hardware than the naive number suggests. When the first such build exists, we will measure it and update this page the way we did for K3.

What the Launch Tells Us

Alibaba shipped the 2.4T first and the 27B two days later, and the 27B is the one almost everyone reading this will actually run. The pattern across Kimi K3, DeepSeek V4 Flash and now Qwen3.8 is that the interesting engineering has moved from parameter count to attention layout. Three quarters of this model’s layers do not cache tokens at all, and that single design choice is worth more to a 24 GB card than any quantization advance of the last year.

To size hardware you already own, the VRAM calculator carries both models with their measured numbers, and the GPU-first tool works from the card side. For the wider picture see our open-weight tier guide, and if a frontier model is out of reach, renting is the cheap way to try it.

FAQ

How much VRAM does Qwen3.8-2.4T-A95B need?
About 1,369 GB at Q4_K_M with an 8K context, sized from Alibaba’s released config.json, with about 1,355 GB of it weights. At 256K it is 1,392 GB, at 1M about 1,462 GB.

Can Qwen3.8-2.4T run at home?
Not at full speed. It needs a datacenter class pool, the smallest in our catalog being 1,728 GB of accelerators: four MI430X, six MI350X or six B300. But community quantization and offloading made Kimi K3 runnable on far less within 48 hours of its release, and the same will happen here. Runnable at a cost is not the same as full speed.

Is Qwen3.8-2.4T open weights?
Yes. It is the first Qwen-Max-class model with open weights, released 12 August 2026 under Alibaba’s open license, in BF16 reference and FP8 native precision.

How much VRAM will Qwen3.8-27B need?
15.8 GB at Q4_K_M with an 8K context, using Unsloth’s current 16.46 GB file, which is 66% of a 24 GB card. Add 0.93 GB for the vision mmproj if you want image or video input. At 64K context the total is 19.40 GB, or 20.33 GB with vision. We measured it ourselves on a 24 GB L4: 16.87 GB at 32K and 18.90 GB at 64K.

Will Qwen3.8-27B run on a 24GB card?
Yes at 6-bit, 5-bit and 4-bit, comfortably at 3-bit, not at 8-bit, which needs about 27 GB. Context matters as much as quantization. 64K fits at 19.40 GB on default settings. Quantize the KV cache and the full 262K context fits too: we measured 20.55 GB on a 24 GB L4 with a q4_0 cache, and @analogalok 23.00 GB on a 4090 with a larger file.

Will Qwen3.8-27B run on a 16 GB card?
Yes, at 3-bit. We ran Unsloth’s UD-Q3_K_XL on a 16 GB Tesla T4: 12.31 GB at an 8K context, and 13.20 GB with a 64K window and a q4_0 cache. A 16 GB gaming card has a little more free memory than the T4, which reports 15.0 GB.

Is Qwen3.8-2.4T better than Claude Opus 4.8?
On Terminal-Bench 2.1 it scored 86.6 to Opus 4.8’s 84.6 in Alibaba’s own launch figures, which are vendor-published claims rather than independent runs. It leads Alibaba’s comparison on PaperBench, IFBench, WideSearch, HealthBench and PRBench-Finance.

Which open-weight model gives the most performance per gigabyte?
DeepSeek V4 Flash 0731, comfortably. It scores 82.7 on Terminal-Bench 2.1 at 284B, against 88.3 for Kimi K3 at 2.8T. Roughly a tenth the size for 94% of the score, and it runs on hardware an individual can buy.

When is Qwen3.8-27B released?
It is out. Weights went public at 15:00 UTC on 14 August 2026, two days after the 2.4T, as BF16 at 55.56 GB under Apache 2.0. GGUF builds appeared within half an hour and the calculator entry went live the same hour.