How accurate is our VRAM calculator?
In this round of tests, every setup our calculator called a fit ran: 27 of 27. We test it the hard way. We load the model on a real card, send one prompt that fills 90% of its context window, and record the most memory the card used while it worked. Then we put our estimate beside that number.
On 33 setups across three cards, our estimate was within 5% of the real figure on 30 of the 31 that loaded, and never below it. The median gap was 2.6%, the largest 6.1%. Where the calculator got a verdict wrong, it said "too tight" for a model that squeezed in. It never said "fits" for one that crashed.
The full VRAM calculator accuracy record is below, every run, including the misses.
How we test it
Each setup is one model file, one card, one context length and one cache type. For each one we:
- start llama.cpp with the whole model on the card, one request at a time, flash attention on;
- send a short request so the model is fully loaded, then read the card;
- send one prompt that fills 90% of the context window and ask for 64 tokens of answer;
- read the card's memory every quarter of a second until the answer returns, and keep the highest reading.
That last number is what the card really used. It counts everything llama.cpp put on the card, not only the weights and the cache.
The setups and our estimates were written down and committed on 1 October 2026, before any of these runs, and the table prints them as committed with one exception: Flash Next's four rows include the 0.38 GB these runs found, explained below. The estimates were not blind to the hardware. The engine had already been corrected against an earlier round on the same 33 setups, on 30 September, which read each card right after loading. What this round tested is whether filling the window takes memory a load does not, and whether the corrected engine still holds when it does. The runs took place on 2 October 2026 on Linux, NVIDIA driver 580.82.07, llama.cpp build b11062, on three cards: a Tesla T4 16 GB (it reports 15,360 MiB with error correction on), an NVIDIA L4 24 GB (23,034 MiB, also with error correction on) and an RTX PRO 6000 Blackwell Server Edition 96 GB (97,887 MiB, error correction on). The model files are the public GGUF files from Unsloth, ggml-org and Ornith, named in the table.
The results, every run
"Card used" is the highest reading during the long prompt. "We say" is what the calculator gives for the same model, quant, context and cache, with llama.cpp as the backend. "Calculator" is its verdict for that card: "fits" means our estimate plus 5% fits in the memory the card reports. When the estimate fits but the 5% does not, the calculator still lists the card, flagged "tight, needs a headless card", and five rows below sit in that band; the table calls them too tight. To reproduce a row, pick the quant the file belongs to: Q3_K_M for UD-Q3_K_XL, Q4_K_M for UD-Q4_K_M, and FP16 / BF16 native for the BF16 file and for gpt-oss's MXFP4.
| Card | Model, file | Context, cache | Card used | We say | Calculator | Real run |
|---|---|---|---|---|---|---|
| T4 16 GB | Qwen 3.8 27B, UD-Q3_K_XL | 8K, f16 | 12.31 GB | 12.82 GB | fits | ran |
| T4 16 GB | Qwen 3.8 27B, UD-Q3_K_XL | 16K, f16 | 12.82 GB | 13.34 GB | fits | ran |
| T4 16 GB | Qwen 3.8 27B, UD-Q3_K_XL | 32K, q8_0 | 12.99 GB | 13.54 GB | fits | ran |
| T4 16 GB | Qwen 3.8 27B, UD-Q3_K_XL | 32K, q4_0 | 12.49 GB | 13.03 GB | fits | ran |
| T4 16 GB | Qwen 3.8 27B, UD-Q3_K_XL | 64K, q4_0 | 13.20 GB | 13.76 GB | fits | ran |
| T4 16 GB | Gemma 4 12B, Q8_0 | 32K, f16 | 13.04 GB | 13.25 GB | fits | ran |
| T4 16 GB | Gemma 4 12B, Q8_0 | 128K, f16 | n/a | 14.85 GB | too tight | out of memory at load |
| T4 16 GB | Gemma 4 12B, Q8_0 | 128K, q4_0 | 13.04 GB | 13.43 GB | fits | ran |
| T4 16 GB | Qwen3 14B, Q6_K | 8K, f16 | 12.21 GB | 12.44 GB | fits | ran |
| T4 16 GB | Qwen3 14B, Q6_K | 16K, f16 | 13.47 GB | 13.71 GB | fits | ran |
| T4 16 GB | Qwen3 14B, Q6_K | 32K, q4_0 | 12.43 GB | 12.75 GB | fits | ran |
| T4 16 GB | Ornith 1.5 9B, Q8_0 | 128K, f16 | 12.25 GB | 12.85 GB | fits | ran |
| T4 16 GB | Ornith 1.5 9B, Q8_0 | 128K, q4_0 | 9.85 GB | 10.45 GB | fits | ran |
| T4 16 GB | gpt-oss-20b, MXFP4 | 64K, f16 | 12.50 GB | 13.01 GB | fits | ran |
| T4 16 GB | gpt-oss-20b, MXFP4 | 128K, f16 | 14.06 GB | 14.58 GB | too tight | ran |
| L4 24 GB | Qwen 3.8 27B, UD-Q4_K_M | 32K, f16 | 16.87 GB | 17.35 GB | fits | ran |
| L4 24 GB | Qwen 3.8 27B, UD-Q4_K_M | 64K, f16 | 18.90 GB | 19.40 GB | fits | ran |
| L4 24 GB | Qwen 3.8 27B, UD-Q4_K_M | 128K, q8_0 | 19.67 GB | 20.21 GB | fits | ran |
| L4 24 GB | Qwen 3.8 27B, UD-Q4_K_M | 256K, q4_0 | 20.55 GB | 21.09 GB | fits | ran |
| L4 24 GB | Qwen 3.8 27B, UD-Q4_K_M | 128K, f16 | n/a | 23.50 GB | too tight | out of memory at load |
| L4 24 GB | Qwen3.6 35B-A3B, UD-Q4_K_M | 32K, f16 | 21.17 GB | 21.27 GB | fits | ran |
| L4 24 GB | Qwen3.6 35B-A3B, UD-Q4_K_M | 128K, q4_0 | 21.54 GB | 21.69 GB | too tight | ran |
| L4 24 GB | Gemma 4 31B, Q4_K_M | 32K, f16 | 21.17 GB | 21.32 GB | fits | ran |
| L4 24 GB | Gemma 4 31B, Q4_K_M | 128K, q4_0 | 21.64 GB | 22.17 GB | too tight | ran |
| L4 24 GB | Ornith 1.5 9B, BF16 | 128K, f16 | 19.30 GB | 20.19 GB | fits | ran |
| L4 24 GB | Ornith 1.5 9B, BF16 | 256K, q4_0 | 18.64 GB | 19.55 GB | fits | ran |
| RTX PRO 6000 96 GB | Qwen 3.8 Flash Next, UD-Q3_K_XL | 64K, f16 | 59.72 GB | 60.15 GB | fits | ran |
| RTX PRO 6000 96 GB | Qwen 3.8 Flash Next, UD-Q3_K_XL | 128K, f16 | 61.81 GB | 62.24 GB | fits | ran |
| RTX PRO 6000 96 GB | Qwen 3.8 Flash Next, UD-Q3_K_XL | 256K, q4_0 | 61.17 GB | 62.55 GB | fits | ran |
| RTX PRO 6000 96 GB | Qwen 3.8 Flash Next, UD-Q3_K_XL | 256K, f16 | 66.00 GB | 66.40 GB | fits | ran |
| RTX PRO 6000 96 GB | gpt-oss-120b, MXFP4 | 128K, f16 | 63.86 GB | 66.15 GB | fits | ran |
| RTX PRO 6000 96 GB | Llama 3.3 70B, Q6_K | 64K, f16 | 74.02 GB | 74.37 GB | fits | ran |
| RTX PRO 6000 96 GB | Llama 3.3 70B, Q6_K | 128K, f16 | 94.08 GB | 94.63 GB | too tight | ran |
Two setups ran out of memory while loading, and the calculator called both too tight. Four setups ran although the calculator called them too tight; they are the misses, and the next section explains them.
Of the five setups the calculator flags "tight, needs a headless card", four ran on these cards, which have no screen. The fifth, Gemma 4 12B at 128K with an f16 cache on the T4, ran out of memory while loading, so that flag was wrong once. The same model and window with a q4_0 cache ran at 13.04 GB.
Why the misses all land on the safe side
All four misses come from the same deliberate margin. The calculator only says "fits" when its estimate plus 5% fits in the memory the card reports. On a gaming PC the same card also draws your desktop, your browser and your monitors, and that memory is not ours to count. Three of the four setups that ran anyway used more than 95% of a card that had no screen attached: Llama 3.3 70B at 128K used 94.08 GB of the 95.6 GB its card reports. The fourth, gpt-oss-20b at 128K on the T4, used 14.06 of 15 GB. There our estimate of 14.58 GB ran 0.52 GB high, and the 5% on top of it crossed the line.
The estimate itself sits just above the real figure on purpose: closest at +0.10 GB, Qwen3.6 35B-A3B at 32K on the L4. A wrong "fits" costs you a 20 GB download and a crash. A wrong "too tight" costs you a model one step smaller than you could have run. We build for the second.
What these runs changed
These numbers changed the engine behind every tool on this site.
- The allowance came down. The calculator used to add 0.75 GB plus 2% for llama.cpp's own needs. The runs showed the real cost was smaller, so it is now 0.3 GB plus 1%. Ollama and LM Studio keep their older allowances until we measure them.
- One model needed more. Qwen 3.8 Flash Next takes 0.38 GB more than it loads with once a long prompt runs. No other model in these runs rose by more than 0.06 GB. The calculator now counts it.
- Datacenter cards are counted at what they report. A Tesla T4 or an L4 keeps its error correction in its own memory and reports about 6% less than its label. The engine knew; the calculator's card list did not use it until these runs showed the gap, and since 2 October it does.
- A quantized cache has a working layer. With a q8_0 or q4_0 cache, llama.cpp keeps one layer's worth of the cache at full precision while it works. The calculator now adds it, which is why the long 4-bit cache rows above stay on the safe side.
Together these changes turned three "too tight" answers into correct "fits" answers, and added no wrong "fits".
An earlier round of the same test, on 30 September, measured the memory right after loading. These runs filled the window and found that, for every model except Flash Next, the load is already the peak: a long prompt added 0.01 to 0.06 GB.
What these runs cover
The record holds for what we ran: llama.cpp, one request at a time, flash attention on, the default batch sizes, on these three cards. Settings change memory. One published table points the other way: @witcheer's RTX 5090 peaks, sampled during a speed sweep on his own build and settings, sit above our estimate on several of the same models. We have not reproduced that in a controlled run, and it is the first thing we compare when a report comes in. Several parallel slots, a larger batch, or another runtime such as Ollama, LM Studio or vLLM can use more, and they are outside this round. The cards here are datacentre cards without a display; a GeForce card that also runs your screen has less free, which is what the 5% is for.
We add runs as we make them, on more cards and more runtimes, and this page shows every one.
Check your own card, or prove us wrong
Pick your model, quant and context in the calculator, or start from your card in the GPU checker. For speed rather than memory, the tokens-per-second checker uses measured runs too. An earlier in-house run, a 27B on a Tesla T4, is written up in We Ran a 27B on a Tesla T4.
If you have a run where the card used more than we say, send it with the llama.cpp build, the context, the cache type, the number of parallel slots and the peak your card reported. A run we can reproduce goes into this table, whichever way it points.
FAQ
How accurate is the VRAM calculator?
On 33 setups on real cards, every "fits" answer held, 27 of 27, and the estimate was within 5% of the measured peak on 30 of the 31 setups that loaded, never below it. The median gap was 2.6%.
Does the calculator ever say a model fits when it does not?
Not in these 33 runs. Its four wrong verdicts were all the other way: it called a setup too tight and the setup ran. Three of those used more than 95% of the card; on the fourth, gpt-oss-20b at 128K on a T4, our estimate ran 0.52 GB high.
Why does the calculator keep 5% of the card free?
Because the card that runs your model usually also runs your screen. The desktop, the browser and other apps take some of that memory, and the calculator cannot see how much. Keeping 5% free means a "fits" answer still holds on a normal desktop.
Which runtime do these results apply to?
llama.cpp, build b11062, one request at a time with flash attention on. Other runtimes and settings, such as several parallel slots, can use more memory and are outside these runs.
How do I report a result that does not match?
Send the model file, the card, the context length, the cache type, the llama.cpp build, the number of parallel slots and the peak memory your card reported. If we can reproduce it, it goes into the table on this page.