Strata: Running Qwen 3.8 Flash Next on a 12 GB Card

What Strata is

Strata is an open-source engine that runs Qwen 3.8 Flash Next, a 180-billion-parameter model built for servers, on one NVIDIA gaming card with 12 to 24 GB and 64 GB of system RAM. It comes from Niko Veit (@coldniko), under the MIT licence, installs on Windows or Linux with one script, and serves an OpenAI-compatible API on your own machine. The README's headline: "Run a 125-billion-parameter AI model on a normal gaming PC".

The two parameter counts are both right. Qwen's card lists 125 billion parameters, plus a 51-billion-parameter n-gram embedding table and a 4-billion-parameter multi-token-prediction layer, and the files add up to 180.0 billion. The 125 billion is the network that does the computing, and the 51 billion is a lookup table. Strata's whole design follows from that split.

The first release went up on 24 September. Four days later there were 19, and the repository had passed 800 stars.

Where the model lives

Fully resident, the model does not come close to a gaming card. Our engine, the one behind the calculator, sizes Qwen 3.8 Flash Next at 76.0 GiB at Q2_K and 106.8 GiB at Q4_K_M with 8K of context, which is why our Flash-Next VRAM page points at a DGX Spark. 26.8 GiB of either figure is the n-gram table.

Strata spreads the same model over three kinds of memory, by its own documentation:

Part Fully resident Strata
The n-gram table, 28.8 GB in memory on the SSD, a few rows read per token
All 24,576 routed experts in memory in system RAM, pinned; the most-used also cached on the GPU
Attention and DeltaNet mixers, gated-residual weights, routers, shared experts, output head, the MTP draft layer, the KV cache in memory on the GPU; from 64K, only the most-read part of the KV cache stays there and the rest streams from RAM

The experts are the trick. Each token uses 10 of each layer's 512 experts. The graphics card keeps the ones asked for most, learned from use, and the CPU computes the rest where they sit in RAM, at the same time as the GPU works. Every extra gigabyte of VRAM holds about 700 more experts, which is why the README says a bigger card is faster but does not lower the RAM needed.

This is offloading, and it is the part our tools deliberately do not model: they answer what fits fully on the card. Strata is the case where a careful split beats that answer by a wide margin, the same question our 405B on 8 GB page asks of a larger model.

Which PC it needs

Strata runs ISTA-DASLab's GSQ-RCO quantizations, which give each tensor its own format under a size budget. Each ships as two shards: the network, and the 28.8 GB n-gram table that stays on disk. From ISTA-DASLab's card and Strata's documentation:

Build Bits per weight Shard 1 RAM it uses Fits
Q2_0 2.40 37.6 GB ~34 GB experts + ~6 GB 48 GB of RAM or more
IQ2_XS 2.50 39.2 GB ~36 GB experts + ~6 GB 48 GB of RAM or more
IQ3_XXS 3.00 47.0 GB ~43 GB experts + ~6 GB 64 GB, context to 128K
IQ3_S 3.50 54.8 GB ~50 GB experts 64 GB, with little else open
Coder 1.89 29.6 GB ~23 GB experts 32 GB of RAM

Strata's rule of thumb is RAM of at least shard 1 plus about 10 GB for Windows and everything else. It asks for an RTX 30, 40 or 50 card with 12 GB or more; the documentation adds that 8 GB runs, slowly, which is where the repository description's 8 GB comes from. It also needs a current NVIDIA driver and about 80 GB of disk, plus a one-time copy of about 40 GB if you pick Q2_0 on an AVX-512 processor such as a Ryzen 7000. Its README warns that while the model starts, the PC "can be slow or stop responding for 1-3 minutes", longest the first time, as it loads 35 to 55 GB into RAM.

The Coder is ISTA-DASLab's pruned build: 256 of each layer's 512 experts kept, chosen on code, agent and vision data. Its authors report 91.3% of the full model's SWE-bench Verified and 98.7% of LiveCodeBench v6. Its 1.89 counts bits per parameter of the original model, because half the experts are gone; the experts it keeps are stored at about IQ3_S precision. It is the one that fits a 32 GB PC, and it is weaker outside code.

How fast it writes

Strata's author measured every build on one machine: an RTX 5070 with 12 GB, a Ryzen 5 7600 with six cores, 64 GB of DDR5-5200, Windows, engine 0.1.14, one code-agent prompt per length, 256 generated tokens, with the model's own multi-token prediction drafting. The IQ2_XS row was measured on Swift 1.5's IQ2_XS, a fine-tune the documentation says runs at the original's speed. Output in tok/s:

Build 4K 32K 128K 262K
Q2_0 90.3 73.6 67.2 60.3
IQ2_XS 73.8 71.5 59.8 52.8
IQ3_XXS 62.1 51.4 45.8
IQ3_S 51.6 48.2 40.5
Coder 50.6 53.3 44.0 42.8

IQ3_XXS and IQ3_S were not measured at 262K: with 43 and 50 GB of experts, a full window takes a 64 GB PC to its memory limit, so setup caps them at 128K.

Reading the prompt is the slow part, and the part that moved this week. On a 32K prompt, engine 0.1.13 took Q2_0 from 572 to 1,290 tok/s by reading in chunks of up to 8,192 tokens and streaming the next layer's experts over PCIe during attention. Q2_0 now reads 1,308 tok/s at 32K and 967 at the full 262K window, so a 32K prompt takes about 25 seconds and a full window about 4.5 minutes. Follow-up turns read only what is new.

For other cards the documentation gives estimates, labelled "Not measured" and "±20%": an RTX 3090 at about 140 tok/s on Q2_0 at 4K, an RTX 5060 Ti 16GB at about 87. The estimates' prompt figures predate 0.1.13.

For scale: on one DGX Spark, which holds the whole model, seven tuned runs in our speed records from the second half of September, by five testers including @ViC305 and @yume_arasaki, decode Flash-Next at 43.9 to 79.6 tok/s with a median of 71.7, at 3 to 4 bits (EXL3 at 3.05 bits and NVFP4). That is not like for like (a different quantization, prompt and engine), and it says only this much: a 12 GB card is now in the same range for this model.

What the 2-bit build gives up

ISTA-DASLab publishes the quality of each build against the full BF16 model, 354 GB:

Build Size AIME25 GPQA-Diamond LiveCodeBench v6 Task average
BF16 354 GB 100.00 91.92 87.43 93.12
Q2_0 66.4 GB 96.67 89.39 81.14 89.07
IQ2_XS 68.0 GB 96.67 87.37 83.43 89.16
IQ3_XXS 75.8 GB 100.00 91.41 86.29 92.57
IQ3_S 83.6 GB 100.00 92.93 86.86 93.26

The fastest build costs about four points of task average, most of it in code. ISTA-DASLab reads its own above-100% results as benchmark noise: "recoveries slightly above 100% reflect benchmark variance, not a model that is better than the one it was quantized from." IQ3_S matches the full model on these tests and is the slowest to run. Strata's default suggestion sits between, at IQ2_XS.

The switch named "speed projection"

Strata ships an optional "experimental speed projection", off by default. Strata's documentation quotes the vector's own package on what it is: "a refusal-direction projection", with which the model "declines far fewer requests (it reports 1 of 50 vs 50 of 50 on its test set)". It is not a speed feature; the same page puts its cost at 0.2 to 0.4% per token, and measures it changing the top token at 10% of positions. It is the technique our abliterated models page explains, applied at run time, and the setup asks before turning it on.

Who should try it

  • A 12 to 24 GB NVIDIA card and 64 GB of RAM: the machine it was built for, and a 12 GB card with 64 GB is the one it was measured on. Start with IQ2_XS, the default suggestion.
  • 32 to 48 GB of RAM: the Coder, if your work is code.
  • A Mac: not this engine, because it is NVIDIA only. Flash-Next runs on Macs through MLX engines such as oMLX. One of Strata's credited influences is a Mac engine, Inco Splash, though Splash itself runs Qwen 3.8 27B and Qwen3.6 35B-A3B rather than Flash-Next.
  • To check your card first: the GPU checker says what runs fully on it; Strata is the answer for what does not.

FAQ

What is Strata?

An open-source, MIT-licensed engine by Niko Veit that runs Qwen 3.8 Flash Next on one NVIDIA card with 12 to 24 GB and 64 GB of RAM, by keeping the model's experts in system RAM, the busiest of them cached on the GPU, and its n-gram table on the SSD.

How much RAM does Strata need?

Its rule is shard 1 plus about 10 GB. With 64 GB every build fits; with 48 GB, Q2_0 and IQ2_XS; with 32 GB, the Coder build.

How fast is Strata?

On the author's RTX 5070 12GB with 64 GB of RAM, the Q2_0 build writes 90.3 tok/s at 4K and 67.2 at 128K, and reads a 32K prompt at about 1,300 tok/s. Other cards are estimates in the documentation.

Does the 2-bit model lose quality?

On ISTA-DASLab's tests, Q2_0 scores 89.07 on the task average against 93.12 for the full model; IQ3_S scores 93.26, parity by its authors' reading.

What is the experimental speed projection?

An optional control vector, off by default, that Strata's own documentation describes as a refusal-direction projection: the model declines far fewer requests with it on. It is not a speed optimization.