LLM VRAM Guide: Know What Fits Before You Buy

Issue 1, out 30 September

Know what fits before you buy

A GPU, a DGX Spark or a 128 GB Mac is a purchase you make once. The Local LLM Sizing Guide shows what each one really runs before you spend: every model in the catalogue, sized for every machine from 8 GB to 128 GB, with the best quantization and the context it really holds.

From $9 a month. Card through Whop, or crypto. Every purchase includes the next issue.

The Local LLM Sizing Guide, Issue 1, September 2026: cover story DeepSeek V4.1 Flash
Issue 1October issue included
Read, not guessedEvery model checked against its lab's own config file
8 GB to 128 GBEvery tier, with the DGX Spark and 128 GB Macs in their own chapter
Measured, creditedSpeed findings from real before-and-after runs
Every monthThe month's releases, sized and tested
Why it exists

Other calculators guess your context. This guide reads it.

The rule of thumb31.8 GB

of KV cache, guessed from the parameter count

Read from the lab's config4.2 GB

what llama.cpp actually allocates

Qwen3.8 27B at 64K of context: the guess is 7.5 times too high. Enough to send you shopping for memory you do not need, or to rule out a card that runs your model perfectly well.

Inside Issue 1

Everything you need before you buy.

Cover story

DeepSeek V4.1 Flash: 552B on two DGX Sparks

The month's biggest model fits no single card, yet runs on machines too small to hold it. How, what it costs in memory on every tier, and how fast it went as the recipes caught up.

39.5Intelligence Index
GPT-6 Luna: 37.3
26.8%Terminal-Bench 4.0
GPT-6 Luna: 12.6%
68.9%AutomationBench
#1 of the 9 compared

Benchmarks independently run by Artificial Analysis, Intelligence Index v4.3.2.

Same 128 GB?

DGX Spark against the 128 GB Mac

64 models fit a Spark, 59 a 128 GB Mac by default, and Qwen3.8 Flash-Next drops from Q4_K_M to Q2_K. With measured Spark speeds.

Every tier

8 GB to 128 GB, every model

The best quantization that fits, and the context it holds with an F16, Q8_0 or Q4_0 cache.

Format watch

EXL3, analysed

The format everyone is moving to: what it costs in memory, where it runs, what it does for speed.

Tested

Up to 2x speed from one setting

What each switch is worth, from real before-and-after runs, credited to the people who ran them.

One row of the 16 GB chapter

Qwen3.8 27B on a 16 GB card

Best fit Q3_K_M, 13.9 GB. Then the context it holds with each cache type:

F16 cache16K
Q8_0 cache32K
Q4_0 cache64K

Four times the context from one setting. Every model on every tier is laid out like this.

Get the guide

Issue 1 now, the October issue included.

Card payments through Whop. Your issues arrive by email and in your Whop library.

BEST VALUE

Subscription

Every issue as it comes out

$9per month, cancel any time
Subscribe
  • Issue 1 on 30 September, then every month
  • The month's releases, sized and tested
  • Every correction counted, every issue

Single issue

Issue 1, and the next one

$24.90one payment, two issues
Buy Issue 1
  • Issue 1, September 2026
  • The October issue, included
  • No subscription
Questions

Before you buy

How is this different from the free calculator?

The calculator answers one model on one card per query. The guide lays out every model on every tier at once, with the context each cache type buys, plus the month's releases and measured speeds, in one document you keep.

When do I get it?

Issue 1 is published on 30 September and emailed to every buyer that day. A new issue comes out on the last day of each month, and whenever you buy, the next issue is included.

Which cards does it cover?

Every tier from 8 GB to 128 GB: one reference card per tier, with every other card of the same memory listed and any difference stated. The DGX Spark and 128 GB Macs get their own chapter.

Will the model really load at the size you give?

Yes. Every fit is the whole model in your card's own memory, weights, cache and runtime all counted, so it loads at the size we give.

Can I pay in crypto?

Yes. Switch the plans above to crypto: the subscription is 3 months upfront for $27, the single issue $24.90. The guide is delivered by email, the same as a card purchase.

Stop guessing what fits.

Issue 1 arrives on 30 September, with the October issue included.

Sizing GuideAfter paying, email us your payment ID.
Trouble loading? Open the payment page