Know what fits before you buy
A GPU, a DGX Spark or a 128 GB Mac is a purchase you make once. The Local LLM Sizing Guide shows what each one really runs before you spend: every model in the catalogue, sized for every machine from 8 GB to 128 GB, with the best quantization and the context it really holds.
From $9 a month. Card through Whop, or crypto. Every purchase includes the next issue.

Other calculators guess your context. This guide reads it.
of KV cache, guessed from the parameter count
what llama.cpp actually allocates
Qwen3.8 27B at 64K of context: the guess is 7.5 times too high. Enough to send you shopping for memory you do not need, or to rule out a card that runs your model perfectly well.
Everything you need before you buy.
DeepSeek V4.1 Flash: 552B on two DGX Sparks
The month's biggest model fits no single card, yet runs on machines too small to hold it. How, what it costs in memory on every tier, and how fast it went as the recipes caught up.
GPT-6 Luna: 37.3
GPT-6 Luna: 12.6%
#1 of the 9 compared
Benchmarks independently run by Artificial Analysis, Intelligence Index v4.3.2.
DGX Spark against the 128 GB Mac
64 models fit a Spark, 59 a 128 GB Mac by default, and Qwen3.8 Flash-Next drops from Q4_K_M to Q2_K. With measured Spark speeds.
8 GB to 128 GB, every model
The best quantization that fits, and the context it holds with an F16, Q8_0 or Q4_0 cache.
EXL3, analysed
The format everyone is moving to: what it costs in memory, where it runs, what it does for speed.
Up to 2x speed from one setting
What each switch is worth, from real before-and-after runs, credited to the people who ran them.
Qwen3.8 27B on a 16 GB card
Best fit Q3_K_M, 13.9 GB. Then the context it holds with each cache type:
Four times the context from one setting. Every model on every tier is laid out like this.
Issue 1 now, the October issue included.
Card payments through Whop. Your issues arrive by email and in your Whop library.Paid through NOWPayments. After paying, email admin@vramcalculator.com your payment ID so we know where to send your issues.
Subscription
Every issue as it comes out
- Issue 1 on 30 September, then every month
- The month's releases, sized and tested
- Every correction counted, every issue
Single issue
Issue 1, and the next one
- Issue 1, September 2026
- The October issue, included
- No subscription
Before you buy
How is this different from the free calculator?
The calculator answers one model on one card per query. The guide lays out every model on every tier at once, with the context each cache type buys, plus the month's releases and measured speeds, in one document you keep.
When do I get it?
Issue 1 is published on 30 September and emailed to every buyer that day. A new issue comes out on the last day of each month, and whenever you buy, the next issue is included.
Which cards does it cover?
Every tier from 8 GB to 128 GB: one reference card per tier, with every other card of the same memory listed and any difference stated. The DGX Spark and 128 GB Macs get their own chapter.
Will the model really load at the size you give?
Yes. Every fit is the whole model in your card's own memory, weights, cache and runtime all counted, so it loads at the size we give.
Can I pay in crypto?
Yes. Switch the plans above to crypto: the subscription is 3 months upfront for $27, the single issue $24.90. The guide is delivered by email, the same as a card purchase.
Stop guessing what fits.
Issue 1 arrives on 30 September, with the October issue included.