Don’t know which model to look up? Start from your graphics card instead: see what LLM your GPU can run. Check my GPU →

VRAMCalculator.com

VRAM Calculator for Local LLMs

Required VRAM 0.0 GB
Model weights 0.0 GB
KV cache 0.0 GB
Runtime overhead 0.0 GB
How this result is calculated

Matching GPU options

smallest viable pools first
Calculator assumptions

Weights: parameters x bits per weight / 8.

KV cache: layers x KV heads x head dimension x context x batch x K/V tensors x bytes per element. Architecture-aware: MLA and sliding-window models cache far less. Quantizing the cache with -ctk/-ctv shrinks it further.

Runtime: backend overhead covers allocator reserve, graph buffers, CUDA/runtime workspace, and practical safety margin.

Hardware: multi-GPU fits include extra headroom for PCIe and tensor-parallel synchronization.

GGUF files: read entirely in your browser, header only. The file is never uploaded. Weight size is measured from the file itself.

Domain for sale: voxllm.com

Got your number? Now see it the other way round: what your GPU can run 

Architectural Infrastructure & Hardware Math

Deploying open-source models requires absolute precision, not guesswork. This tool serves as the definitive gguf vram calculator, allowing engineers to properly architect their local llm hardware stack before provisioning any physical components. By calculating exact parameter weights, KV cache at various context lengths, and 15% system runtime overhead, this gpu memory calculator ensures your deployment fits flawlessly within your available silicon architecture.

Deep multi-GPU orchestration updates inside your inbox.

Get raw hardware configuration profiles, custom layer offloading templates, and performance benchmarks directly from the server room.

Subscription Form (#3)