How Much VRAM Does Kimi K3 Need?
The honest answer, as of July 16, 2026: no one can run Kimi K3 locally yet, and no one can compute its full memory footprint yet — but the official numbers already frame it. Moonshot’s launch post confirms a 2.8-trillion-parameter Mixture-of-Experts model shipping natively in MXFP4 — which puts the weights alone at roughly 1.5 TB (straight arithmetic: 2.8T parameters at ~4.25 bits each). Moonshot’s own deployment guidance says “supernode configurations with 64 or more accelerators.” Open weights are promised by July 27, 2026; until then there is no config file, so the KV-cache half of the math — the half that depends on layers, KV heads, and head dimensions Moonshot hasn’t disclosed — cannot be computed by anyone. Here is everything that is official, what the leaks got wrong, and what the numbers mean if you were hoping to run it on your own hardware.
Official Specs vs. What the Leaks Said
For weeks the aggregator sites have repeated “2.5 trillion parameters” as fact. The official launch post says otherwise. This table is sourced from Moonshot’s tech blog and announcement only — as of July 2026:
| Spec | The leaks said | Official (Moonshot launch post) |
|---|---|---|
| Total parameters | 2.5T | 2.8T |
| Architecture | “new architecture” (vague) | MoE — 16 of 896 experts activated per token |
| Attention | — | Kimi Delta Attention (KDA) + Attention Residuals; vendor claims up to 6.3× faster decoding at million-token contexts |
| Context window | 1M (uncertain) | 1M tokens, confirmed |
| Multimodality | rumored | Native vision, confirmed |
| Weights format | — | MXFP4 weights / MXFP8 activations, quantization-aware training from the SFT stage onward |
| Open weights | “maybe at launch” | By July 27, 2026 — not at launch; no HuggingFace repo exists yet (org checked today) |
| License | MIT rumored | Not yet specified |
| Active parameters per token | various guesses | Not disclosed as a parameter count — only the 16-of-896 expert ratio is official |
Two of those rows matter more than all the leak-correcting: the weights format and the undisclosed dimensions. We’ll take them in order.
The Benchmarks — Vendor-Published, Plus What Early Testers Report
Moonshot published a 20-row benchmark table comparing K3 against Claude Fable 5, GPT 5.6 Sol, and Claude Opus 4.8, among others. Selected scores, vendor-published (treat accordingly — every lab tunes its own launch table):
| Benchmark | Kimi K3 | Claude Fable 5 | GPT 5.6 Sol |
|---|---|---|---|
| Terminal Bench 2.1 | 88.3 | 84.6 | 88.8 |
| GPQA-Diamond | 93.5 | 92.6 | 94.1 |
| Program Bench | 77.8 | 76.8 | 77.6 |
| MMMU-Pro | 81.6 | 81.2 | 83.0 |
| DeepSWE | 67.5 | 70.0 | 73.0 |
The vendor’s own table places K3 between the Western frontier models — ahead on some agentic and coding rows, behind on others — which is itself notable restraint compared with the usual launch-table sweep.
Early community reports (community-reported, unverified — benchmark videos and API testers, no formal methodology published anywhere yet): testers describe coding performance in the Opus-class range, with the strongest showings in agentic and long-context tasks and mixed results against GPT 5.6 Sol on polish-heavy multimodal work. No independent, methodology-published benchmark of K3 exists as of July 16. When one appears, this section gets updated; until then, the vendor table plus the API price sheet ($0.30/MTok cache-hit input, $3.00 cache-miss, $15.00 output) are the hard public numbers.
The Local Math: What 2.8T Means for Your Hardware
Weights. 2.8T parameters in MXFP4 (~4.25 bits per weight with microscaling) is on the order of 1.5 TB — arithmetic, not a measurement. For scale: the largest model in our calculator today, DeepSeek-R1 at 671B, computes to 407 GB of Q4_K_M weights (the tool’s own reading); K3 is roughly 3.7× that. An H200 holds 141 GB. Moonshot saying “64 or more accelerators” is not hedging — it’s the arithmetic.
The MXFP4 twist — the genuinely interesting part for local AI. K3 was trained quantization-aware from the SFT stage, natively in MXFP4. On every previous frontier model, the 4-bit file you run locally is a lossy afterthought of an FP16 original. K3’s 4-bit tier is the native tier. Whatever the GGUF ecosystem produces after July 27, the usual “how much did Q4 hurt it” question starts from a different place — the vendor already answered it during training. That’s new for a model this size, and it’s the one spec that makes the local conversation more than academic.
Sparsity won’t save your GPU, but it changes the streaming math. 16-of-896 experts means under 2% of expert weights are touched per token. As we covered in the weight-streaming reality check, MoE sparsity is exactly what makes disk-streaming viable-ish — GLM-5.2 at 744B runs from a 370 GB NVMe file at community-reported 0.3–1.2 tok/s. K3 doubles the disk requirement to ~1.5 TB and its resident dense fraction is unknown until the weights land. Streaming K3 will be possible in the way crossing an ocean in a rowboat is possible; the invoice will be measured in minutes per reply.
KV cache at 1M context: cannot be computed yet — and that’s the honest headline. The cache math needs layer count, KV heads, and head dimension. Moonshot has disclosed none of them, and Kimi Delta Attention exists precisely to change this math (that vendor-claimed 6.3× decode speedup at million-token contexts is an attention-mechanism claim). Anyone quoting you a K3 KV-cache figure today is guessing. Our per-model breakdown shows why this line item dominates at long context — and why we won’t publish a K3 number before the config exists.
When Will the Calculator Answer This?
The day the weights land. A provisional Kimi K3 entry is already staged for the calculator with every officially published field — and it ships only when Moonshot’s config.json is public (promised by July 27), because that file, not the launch post and not this article, is the source of truth for layers, heads, and dimensions. That’s the same rule that kept the Gemma 4 numbers right when half the internet had them wrong. Check back after July 27; the entry goes live the day the config does.
FAQ
Can I run Kimi K3 locally right now?
No. As of July 16, 2026, only the API and Kimi’s own products serve K3. Open weights are promised by July 27, and no HuggingFace repository exists yet. Any “run K3 locally” guide you see before the weights drop is describing the API.
How big will the Kimi K3 download be?
Roughly 1.5 TB class — 2.8T parameters shipping natively in MXFP4 (~4.25 bits per weight, arithmetic estimate). That is roughly 3.7 times the Q4_K_M footprint of DeepSeek-R1 671B (407 GB), currently the largest model in our calculator.
Is Kimi K3 really 2.5 trillion parameters?
No — that was the pre-launch leak number, and it survives on many aggregator sites. Moonshot’s official launch post says 2.8 trillion, a Mixture-of-Experts model activating 16 of 896 experts per token.
What is Kimi Delta Attention?
Moonshot’s new attention mechanism in K3, paired with Attention Residuals. The vendor claims up to 6.3× faster decoding at million-token contexts. Its exact dimensions — and therefore K3’s KV-cache footprint — are undisclosed until the open weights and config file ship, promised by July 27, 2026.