LLM VRAM Guide: Know What Fits Before You Buy

Issue 1, September 2026

Know what fits before you buy

A GPU, a DGX Spark, a Ryzen AI Max machine or a Mac Studio is a purchase you make once. The Local LLM Sizing Guide shows what each one really runs before you spend: every model we size, on every machine from 8 GB to 512 GB, with the best quantization and the context it really holds. Issue 1 sizes October's new hardware before it ships.

From $9 a month. Card through Whop, or crypto. Every purchase includes the next issue.

The Local LLM Sizing Guide, Issue 1, September 2026: what runs on October's new hardware, and DeepSeek V4.1 Flash against GLM-5.3-Flash and Qwen 3.8 Flash Next
Issue 1October issue included
Read, not guessedEvery model checked against its lab's own config file
8 GB to 512 GBEvery tier, plus eight unified-memory machines side by side
Tested in houseOur own runs, and every outside run credited to the person who ran it
Every monthThe month's releases, sized, with measured runs credited
Tested on real cards Sized by the engine we test on real cards: in 33 runs, every "fits" answer held. Every run published, misses included. See the record
Why it exists

Other calculators guess your context. This guide reads it.

The rule of thumb31.8 GB

of KV cache, guessed from the parameter count

Read from the lab's config4.2 GB

what llama.cpp actually allocates

Qwen 3.8 27B at 64K of context: the guess is 7.5 times too high. Enough to send you shopping for memory you do not need, or to rule out a card that runs your model perfectly well.

Inside Issue 1

Everything you need before you buy.

Cover story

What runs on October's new hardware

RTX Spark, the 192 GB Ryzen AI Max+ PRO 495 and the 512 GB Mac Studio all reach buyers in October. Ten models in demand this month, sized on each before it ships, with the context each holds and a speed ceiling, and a DGX Spark beside them for scale.

63models fit RTX Spark
128 GB, from October
70fit the Ryzen AI Max+
PRO 495, 192 GB
83fit the Mac Studio
M5 Ultra, 512 GB

Every model sized by the engine behind our calculators, each machine at the cautious end of the memory it gives its GPU.

Which Flash model?

DeepSeek V4.1 Flash, GLM-5.3-Flash or Qwen 3.8 Flash Next

Within three points on Artificial Analysis's Intelligence Index, far apart in memory and speed. Which to run on one Spark, two Sparks or a 512 GB Mac Studio, from measured runs.

Every tier

8 GB to 128 GB, every model

The best quantization that fits, and the context it holds with an F16, Q8_0 or Q4_0 cache.

Reference

Every model on eight machines

The whole list side by side, from the 64 GB RTX Spark laptop chip to the 512 GB Mac Studio, with the format and the context each machine holds.

Tested in house

A 16 GB datacenter card gives you 15

Our own runs on a T4 and an A100: the memory ECC takes, what a quantized cache really costs, and our calculator checked against a real card.

One row of the 16 GB chapter

Qwen 3.8 27B on a 16 GB card

Best fit UD-Q3_K_XL, 13.9 GB. Then the context it holds with each cache type:

F16 cache16K
Q8_0 cache32K
Q4_0 cache64K

Four times the context from one setting. Every model on every tier is laid out like this.

Get the guide

Issue 1 now, the October issue included.

Card payments through Whop. Your issues arrive by email and in your Whop library.

BEST VALUE

Subscription

Every issue as it comes out

$9per month, cancel any time
Subscribe
  • Issue 1 now, then every month
  • The month's releases, sized, with measured runs credited
  • Every correction counted, every issue

Single issue

Issue 1, and the next one

$24.90one payment, two issues
Buy Issue 1
  • Issue 1, September 2026
  • The October issue, included
  • No subscription
Questions

Before you buy

How is this different from the free calculator?

The calculator answers one model on one card per query. The guide lays out every model on every tier at once, with the context each cache type buys. It also sizes new hardware before it ships, compares the month's top models on the machines they compete for, and carries our own measurements, in one document you keep.

When do I get it?

Card buyers get Issue 1 in their Whop library the moment they pay; crypto buyers get it by email once they send their payment ID. A new issue comes out on the last day of each month, and whenever you buy, the next issue is included.

Which cards does it cover?

Every tier from 8 GB to 128 GB: one reference card per tier, with every other card of the same memory listed and any difference stated. Beyond that, eight unified-memory machines side by side, from the 64 GB RTX Spark laptop chip to the 512 GB Mac Studio.

Will the model really load at the size you give?

Yes. Every fit is the whole model in your card's own memory, weights, cache and runtime all counted, so it loads at the size we give.

Can I pay in crypto?

Yes. Switch the plans above to crypto: the subscription is 3 months upfront for $27, the single issue $24.90. The guide is delivered by email, the same as a card purchase.

Stop guessing what fits.

Issue 1 is out now, with the October issue included.

Sizing GuideAfter paying, email us your payment ID.
Trouble loading? Open the payment page