VRAMCalculator.com
Can Your PC Run Qwen3.8 Flash Next with Strata?
| Build | RAM it uses | On your GPU | Experts the GPU holds | Verdict |
|---|
Closest measured setups
How this is calculated
We read every build's tensor table from its published files and split the bytes three ways: the routed experts, which Strata keeps in system RAM, the dense weights, which go to the graphics card, and the n-gram table, which stays on the SSD. Six of our seven expert sums match the figures in Strata's own installer to within 0.05 GB; the seventh, the Coder, matches once Strata's 23.4 is read as GiB.
RAM it uses: the experts, plus the conversation's KV cache when Strata moves it to RAM (from 64K, and only when the RAM has room for it), plus 10 GB for the system, the figure Strata's installer uses. On Windows, an AMD card of 16 GB or more keeps 2,560 MiB free instead of 700 since 0.1.40.3. For Unsloth's two builds Strata leaves 24 GB for the system and the file cache instead of 10. On the graphics card: the dense weights, the draft layer, the KV cache kept there, and about 2.3 GB for buffers, the 700 MiB Strata leaves free by default, and the model's small fixed state. What is left holds the most-used experts.
Several cards: Strata splits the layers between them. Each card keeps its own copy of the dense weights, its prompt buffers and its reserve, the last card also the draft layer, and each caches the experts of its own layers, so the caches add up while the experts stay in RAM once. NVIDIA cards share a model among themselves, AMD cards among themselves on Linux; Strata's multi-GPU guide lists Intel cards as not supported for splitting, and community reports have split them with engines built by hand, and a card sharing a model needs 8 GB. Strata's setup recommends two cards by default; choosing more is a menu pick. Unsloth's builds have no RAM budget on a split: across cards they need the download plus 24 GB of RAM.
Low-RAM mode: when the experts do not fit the RAM beside the system, Strata keeps on the card the experts it caches and only the rest in RAM, so less RAM is needed and nothing is read from the SSD while it answers. The share of experts on the card is not a hit rate: in a community test published in Strata's docs, a modified 32 GB RTX 4080 SUPER holding about 45% of the experts served 90 to 94% of reads from them. Expert counts are estimates, rounded to the nearest hundred: against the cache sizes in ten startup logs from Strata's community reports with settings like ours (images off, setup's own KV choice), they land between 7% under and 13% over. Images on, extra parallel sessions or a forced KV setting change the real figure by a gigabyte or more.
Measured setups: runs from Strata's community reports and from posts on X, each credited and linked. A run is a claim until it passes our checks: the build, the card model and a speed are stated, the engine version existed on that date when the run names one, our arithmetic says the setup starts, the speed is under what the card's memory can feed even counting draft tokens, and, where there are enough runs to compare, it is not far outside the others. Runs that fail are held for review, not shown. The closest three to your machine are listed first: the same card, then the same vendor and VRAM, then the card count, RAM and OS.
Version 0.2 Alpha: new and experimental. Strata changes every few days, and so will this tool.
Speeds are measured only, each by the person named. Sources: Strata's README, setup.py and docs/DETAILS.md, and the published GGUF files. About Strata itself: our Strata article.
