VRAM Calculator Accuracy: We Ran It Against Real Cards

How accurate is our VRAM calculator?

In this round of tests, every setup our calculator called a fit ran: 27 of 27. We test it the hard way. We load the model on a real card, send one prompt that fills 90% of its context window, and record the most memory the card used while it worked. Then we put our estimate beside that number.

On 33 setups across three cards, our estimate was within 5% of the real figure on 30 of the 31 that loaded, and never below it. The median gap was 2.6%, the largest 6.1%. Where the calculator got a verdict wrong, it said "too tight" for a model that squeezed in. It never said "fits" for one that crashed.

The full VRAM calculator accuracy record is below, every run, including the misses.

How we test it

Each setup is one model file, one card, one context length and one cache type. For each one we:

  1. start llama.cpp with the whole model on the card, one request at a time, flash attention on;
  2. send a short request so the model is fully loaded, then read the card;
  3. send one prompt that fills 90% of the context window and ask for 64 tokens of answer;
  4. read the card's memory every quarter of a second until the answer returns, and keep the highest reading.

That last number is what the card really used. It counts everything llama.cpp put on the card, not only the weights and the cache.

The setups and our estimates were written down and committed on 1 October 2026, before any of these runs, and the table prints them as committed with one exception: Flash Next's four rows include the 0.38 GB these runs found, explained below. The estimates were not blind to the hardware. The engine had already been corrected against an earlier round on the same 33 setups, on 30 September, which read each card right after loading. What this round tested is whether filling the window takes memory a load does not, and whether the corrected engine still holds when it does. The runs took place on 2 October 2026 on Linux, NVIDIA driver 580.82.07, llama.cpp build b11062, on three cards: a Tesla T4 16 GB (it reports 15,360 MiB with error correction on), an NVIDIA L4 24 GB (23,034 MiB, also with error correction on) and an RTX PRO 6000 Blackwell Server Edition 96 GB (97,887 MiB, error correction on). The model files are the public GGUF files from Unsloth, ggml-org and Ornith, named in the table.

The results, every run

"Card used" is the highest reading during the long prompt. "We say" is what the calculator gives for the same model, quant, context and cache, with llama.cpp as the backend. "Calculator" is its verdict for that card: "fits" means our estimate plus 5% fits in the memory the card reports. When the estimate fits but the 5% does not, the calculator still lists the card, flagged "tight, needs a headless card", and five rows below sit in that band; the table calls them too tight. To reproduce a row, pick the quant the file belongs to: Q3_K_M for UD-Q3_K_XL, Q4_K_M for UD-Q4_K_M, and FP16 / BF16 native for the BF16 file and for gpt-oss's MXFP4.

Card Model, file Context, cache Card used We say Calculator Real run
T4 16 GB Qwen 3.8 27B, UD-Q3_K_XL 8K, f16 12.31 GB 12.82 GB fits ran
T4 16 GB Qwen 3.8 27B, UD-Q3_K_XL 16K, f16 12.82 GB 13.34 GB fits ran
T4 16 GB Qwen 3.8 27B, UD-Q3_K_XL 32K, q8_0 12.99 GB 13.54 GB fits ran
T4 16 GB Qwen 3.8 27B, UD-Q3_K_XL 32K, q4_0 12.49 GB 13.03 GB fits ran
T4 16 GB Qwen 3.8 27B, UD-Q3_K_XL 64K, q4_0 13.20 GB 13.76 GB fits ran
T4 16 GB Gemma 4 12B, Q8_0 32K, f16 13.04 GB 13.25 GB fits ran
T4 16 GB Gemma 4 12B, Q8_0 128K, f16 n/a 14.85 GB too tight out of memory at load
T4 16 GB Gemma 4 12B, Q8_0 128K, q4_0 13.04 GB 13.43 GB fits ran
T4 16 GB Qwen3 14B, Q6_K 8K, f16 12.21 GB 12.44 GB fits ran
T4 16 GB Qwen3 14B, Q6_K 16K, f16 13.47 GB 13.71 GB fits ran
T4 16 GB Qwen3 14B, Q6_K 32K, q4_0 12.43 GB 12.75 GB fits ran
T4 16 GB Ornith 1.5 9B, Q8_0 128K, f16 12.25 GB 12.85 GB fits ran
T4 16 GB Ornith 1.5 9B, Q8_0 128K, q4_0 9.85 GB 10.45 GB fits ran
T4 16 GB gpt-oss-20b, MXFP4 64K, f16 12.50 GB 13.01 GB fits ran
T4 16 GB gpt-oss-20b, MXFP4 128K, f16 14.06 GB 14.58 GB too tight ran
L4 24 GB Qwen 3.8 27B, UD-Q4_K_M 32K, f16 16.87 GB 17.35 GB fits ran
L4 24 GB Qwen 3.8 27B, UD-Q4_K_M 64K, f16 18.90 GB 19.40 GB fits ran
L4 24 GB Qwen 3.8 27B, UD-Q4_K_M 128K, q8_0 19.67 GB 20.21 GB fits ran
L4 24 GB Qwen 3.8 27B, UD-Q4_K_M 256K, q4_0 20.55 GB 21.09 GB fits ran
L4 24 GB Qwen 3.8 27B, UD-Q4_K_M 128K, f16 n/a 23.50 GB too tight out of memory at load
L4 24 GB Qwen3.6 35B-A3B, UD-Q4_K_M 32K, f16 21.17 GB 21.27 GB fits ran
L4 24 GB Qwen3.6 35B-A3B, UD-Q4_K_M 128K, q4_0 21.54 GB 21.69 GB too tight ran
L4 24 GB Gemma 4 31B, Q4_K_M 32K, f16 21.17 GB 21.32 GB fits ran
L4 24 GB Gemma 4 31B, Q4_K_M 128K, q4_0 21.64 GB 22.17 GB too tight ran
L4 24 GB Ornith 1.5 9B, BF16 128K, f16 19.30 GB 20.19 GB fits ran
L4 24 GB Ornith 1.5 9B, BF16 256K, q4_0 18.64 GB 19.55 GB fits ran
RTX PRO 6000 96 GB Qwen 3.8 Flash Next, UD-Q3_K_XL 64K, f16 59.72 GB 60.15 GB fits ran
RTX PRO 6000 96 GB Qwen 3.8 Flash Next, UD-Q3_K_XL 128K, f16 61.81 GB 62.24 GB fits ran
RTX PRO 6000 96 GB Qwen 3.8 Flash Next, UD-Q3_K_XL 256K, q4_0 61.17 GB 62.55 GB fits ran
RTX PRO 6000 96 GB Qwen 3.8 Flash Next, UD-Q3_K_XL 256K, f16 66.00 GB 66.40 GB fits ran
RTX PRO 6000 96 GB gpt-oss-120b, MXFP4 128K, f16 63.86 GB 66.15 GB fits ran
RTX PRO 6000 96 GB Llama 3.3 70B, Q6_K 64K, f16 74.02 GB 74.37 GB fits ran
RTX PRO 6000 96 GB Llama 3.3 70B, Q6_K 128K, f16 94.08 GB 94.63 GB too tight ran

Two setups ran out of memory while loading, and the calculator called both too tight. Four setups ran although the calculator called them too tight; they are the misses, and the next section explains them.

Of the five setups the calculator flags "tight, needs a headless card", four ran on these cards, which have no screen. The fifth, Gemma 4 12B at 128K with an f16 cache on the T4, ran out of memory while loading, so that flag was wrong once. The same model and window with a q4_0 cache ran at 13.04 GB.

Why the misses all land on the safe side

All four misses come from the same deliberate margin. The calculator only says "fits" when its estimate plus 5% fits in the memory the card reports. On a gaming PC the same card also draws your desktop, your browser and your monitors, and that memory is not ours to count. Three of the four setups that ran anyway used more than 95% of a card that had no screen attached: Llama 3.3 70B at 128K used 94.08 GB of the 95.6 GB its card reports. The fourth, gpt-oss-20b at 128K on the T4, used 14.06 of 15 GB. There our estimate of 14.58 GB ran 0.52 GB high, and the 5% on top of it crossed the line.

The estimate itself sits just above the real figure on purpose: closest at +0.10 GB, Qwen3.6 35B-A3B at 32K on the L4. A wrong "fits" costs you a 20 GB download and a crash. A wrong "too tight" costs you a model one step smaller than you could have run. We build for the second.

What these runs changed

These numbers changed the engine behind every tool on this site.

  • The allowance came down. The calculator used to add 0.75 GB plus 2% for llama.cpp's own needs. The runs showed the real cost was smaller, so it is now 0.3 GB plus 1%. Ollama and LM Studio keep their older allowances until we measure them.
  • One model needed more. Qwen 3.8 Flash Next takes 0.38 GB more than it loads with once a long prompt runs. No other model in these runs rose by more than 0.06 GB. The calculator now counts it.
  • Datacenter cards are counted at what they report. A Tesla T4 or an L4 keeps its error correction in its own memory and reports about 6% less than its label. The engine knew; the calculator's card list did not use it until these runs showed the gap, and since 2 October it does.
  • A quantized cache has a working layer. With a q8_0 or q4_0 cache, llama.cpp keeps one layer's worth of the cache at full precision while it works. The calculator now adds it, which is why the long 4-bit cache rows above stay on the safe side.

Together these changes turned three "too tight" answers into correct "fits" answers, and added no wrong "fits".

An earlier round of the same test, on 30 September, measured the memory right after loading. These runs filled the window and found that, for every model except Flash Next, the load is already the peak: a long prompt added 0.01 to 0.06 GB.

What these runs cover

The record holds for what we ran: llama.cpp, one request at a time, flash attention on, the default batch sizes, on these three cards. Settings change memory. One published table points the other way: @witcheer's RTX 5090 peaks, sampled during a speed sweep on his own build and settings, sit above our estimate on several of the same models. We have not reproduced that in a controlled run, and it is the first thing we compare when a report comes in. Several parallel slots, a larger batch, or another runtime such as Ollama, LM Studio or vLLM can use more, and they are outside this round. The cards here are datacentre cards without a display; a GeForce card that also runs your screen has less free, which is what the 5% is for.

We add runs as we make them, on more cards and more runtimes, and this page shows every one.

Check your own card, or prove us wrong

Pick your model, quant and context in the calculator, or start from your card in the GPU checker. For speed rather than memory, the tokens-per-second checker uses measured runs too. An earlier in-house run, a 27B on a Tesla T4, is written up in We Ran a 27B on a Tesla T4.

If you have a run where the card used more than we say, send it with the llama.cpp build, the context, the cache type, the number of parallel slots and the peak your card reported. A run we can reproduce goes into this table, whichever way it points.

FAQ

How accurate is the VRAM calculator?

On 33 setups on real cards, every "fits" answer held, 27 of 27, and the estimate was within 5% of the measured peak on 30 of the 31 setups that loaded, never below it. The median gap was 2.6%.

Does the calculator ever say a model fits when it does not?

Not in these 33 runs. Its four wrong verdicts were all the other way: it called a setup too tight and the setup ran. Three of those used more than 95% of the card; on the fourth, gpt-oss-20b at 128K on a T4, our estimate ran 0.52 GB high.

Why does the calculator keep 5% of the card free?

Because the card that runs your model usually also runs your screen. The desktop, the browser and other apps take some of that memory, and the calculator cannot see how much. Keeping 5% free means a "fits" answer still holds on a normal desktop.

Which runtime do these results apply to?

llama.cpp, build b11062, one request at a time with flash attention on. Other runtimes and settings, such as several parallel slots, can use more memory and are outside these runs.

How do I report a result that does not match?

Send the model file, the card, the context length, the cache type, the llama.cpp build, the number of parallel slots and the peak memory your card reported. If we can reproduce it, it goes into the table on this page.