Effective bits per weight, quantization scales included. Native = exact size of the official checkpoint files.
Requests
Prompt + generated tokens
Requests decoded at the same time; each holds its own KV cache
Set independently of the weight format. FP8 halves the cache.
Hardware
As marketed. The calculation uses what the driver reports when it differs (an L40S sold as 48 GB exposes 44.99 GiB). Unified-memory presets already exclude what the OS keeps.
Auto picks the fewest GPUs that fit: tensor parallel inside a server (up to 8), pipeline parallel across servers.
Serving engine · vLLM
vLLM gpu_memory_utilization, SGLang mem_fraction_static… The rest stays free on purpose.
CUDA context, NCCL, CUDA graphs, activation workspace. An approximation: adjust it from a real nvidia-smi reading.
Reserves gpu_memory_utilization (0.90 by default) of each GPU, then fills what is left after weights and activations with paged KV cache.
Assumes a modern engine with a paged KV cache that keeps only the window on sliding-window layers. Real usage can differ by a few GiB.
2
Your answer
How many GPUs, how fast, and what it costs.
To run it you need
Fits
1 × NVIDIA H100 · 80 GB
Clef, BF16 / FP16 weights, 4 concurrent requests of 32K tokens. 80 GB of GPU memory in total.
From $2.04/h to rent, on Vast.aiOn-demand prices, checked every hour. Prices change: check before you rent.Compare 4 providers ↓
Speed & cost
Per request Generation speed seen by one user, at full context≈ 42 tok/s
Total, 4 requests≈ 168 tok/s
First token, 32K prompt Time to read a prompt as long as the chosen context before the first token appears (prefill), for one request. It depends on the GPU's compute (TFLOPS), not its memory bandwidth.≈ 5.4 s
Your cost per 1M tokens Generated tokens at full load: 1 × NVIDIA H100 on Vast.ai$3.38
No API price for this model on OpenRouter.
Rough estimate (±30–50%) of generation at full context, from memory bandwidth (3.35 TB/s per GPU) and vLLM efficiency; first token from compute (989 TFLOPS). Speculative decoding is not modelled. API prices: OpenRouter, per generated token.
VRAM per GPU
Weights51.0 GiB
KV cache8.00 GiB
Recurrent state0.57 GiB
Runtime overhead3.00 GiB
Needed62.5 GiB
Usable Memory per GPU × usable %71.7 GiB
Headroom9.16 GiB
Capacity
Max requests at 32K8
Max context with 4 requests70,269
KV per token Averaged over the chosen context, whole model. Falls as context grows on sliding-window and compressed layers.64.0 KiB
KV per request At the chosen context, whole model, before tensor-parallel sharding2.00 GiB
Treating every layer as full attention would give 8.00 GiB per request, 4.0× more than this architecture actually caches.
Plus 146.8 MiB of recurrent state per request, fixed whatever the context.
Compare with your real usage
Paste the startup log of vLLM or llama.cpp (llama-server, Ollama, LM Studio). It is read in your browser: nothing is sent.
3
Better quality or a cheaper GPU?
Pick a more faithful weight format (it may need more GPUs) or a cheaper card. Click an option to apply it.
Does it fit? Weights × context on 1 × NVIDIA H100
Weights
4K
8K
32K
128K
256K
BF16 / FP16
FP8
GGUF Q8_0
GGUF Q6_K
GGUF Q5_K_M
GGUF Q4_K_M
GGUF Q3_K_M
Longest context that fits
BF16 / FP16up to 32K
FP8 to GGUF Q3_K_Mup to 128K
✓ fits · ~ tight · ✕ does not fit. Context per request, 4 concurrent requests, vLLM. Click a cell to apply it. Tap a format to use it at its longest context.
Quantization on NVIDIA H100
1 GPU
Fewer bits per weight, fewer GPUs, at some quality cost. The KV cache is set separately.
Cheapest GPUs that fit in one server
Cheapest on-demand price among RunPod, Verda, Vast.ai, Microsoft Azure, checked every hour. Weight format as selected above.
Other models that fit on 1 × NVIDIA H100
Largest first, each with the most faithful weight format that fits 32K tokens × 4 requests.
4
Rent and deploy
Today's NVIDIA H100 prices per provider, then the command to launch it.
Ask Claude, ChatGPT or any MCP client whether a model fits on your GPU: it calls this calculator and answers with the same numbers, for the listed models and any Hugging Face model. No key, no sign-up.
MCP server (Streamable HTTP): https://studiotvai.com/api/mcp
Tools: estimate_vram, models_that_fit, gpu_prices
Claude Code: claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp
Take Llama 3.3 70B. Its weights need 131 GiB in BF16 (2 × H100 80 GB), 67.7 GiB in FP8 and about 39.9 GiB in 4-bit (1 × H100, or 2 × RTX 4090). Then add the KV cache of each request: 2.5 GiB at 8K tokens and 10.0 GiB at 32K.
Can I run an LLM on a 24 GB GPU like the RTX 4090?
Yes, models up to about 35B parameters fit in 4-bit with an 8K-token context: Llama 3.1 8B, Mistral Small 3.2 24B, Qwen3 32B, Qwen3.5 4B, Qwen3.5 9B, Qwen3.8 27B, Qwen3.6 35B-A3B, LFM2.5 8B-A1B, LFM2 24B-A2B, Gemma 4 12B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss 20B, Qwen3-Coder 30B-A3B, Ornith 1.5 35B-A3B, Ornith 1.5 9B, Clef, Clef Flash, D1 3B, Humanizer, JEV 27B VL, Mellum2.1 12B-A2.5B, LightOnOCR-3 4B, Spark-X2.5 4B. Larger models need several GPUs or a card with more memory.
How many GPUs do I need to run an LLM?
Each GPU must hold its share of the weights plus the KV cache of every concurrent request, under the memory the serving engine allows itself (90% by default in vLLM). The calculator tries 1, 2, 4, 8 GPUs and more, splitting the model with tensor parallelism inside a server and pipeline parallelism across servers, and returns the fewest that fit.
What drives LLM VRAM usage?
Four things decide how much memory a model needs once it is serving requests.
Weights are fixed: parameters × bits. “4-bit” formats really cost 4.2 to 4.9 bits once scales are counted.
KV cache grows with context × concurrent sequences; it decides how many users a GPU can serve.
Tensor parallelism splits GQA heads but never below one head per GPU, and replicates MLA latents.
Serving engines leave memory free on purpose: vLLM uses 90% of each GPU by default.
Why is VRAM higher than the model size on disk?
The file holds only the weights. At serving time you also pay for the KV cache, which grows with context × concurrent sequences, runtime buffers such as CUDA graphs and activation workspace, and the share of memory the engine leaves free on purpose.
Why are most LLM VRAM formulas wrong?
The classic estimate, 2 × layers × KV heads × head dim × bytes per token, assumes every layer attends to the whole context. Most models released since 2025 break that rule:
Hybrid linear attention (Qwen 3.5–3.8, Nemotron 3, Kimi K3): only 1 layer in 4, or fewer, keeps a KV cache; the rest hold a fixed-size state.
Sliding windows (Gemma 4, gpt-oss): most layers keep only the last 128 to 1,024 tokens.
MLA (GLM-5.3, Kimi K3): one 576-wide compressed latent per token instead of keys and values.
Compressed attention (DeepSeek V4): a 128-token window plus the rest of the sequence compressed 4× or 128×.
Why does the KV cache differ so much between models of the same size?
2026 architectures cache very different amounts per token. Qwen 3.5+ keeps a KV cache on only 1 layer in 4, Gemma 4 and gpt-oss limit most layers to a short sliding window, MLA models store one compressed latent instead of keys and values, and DeepSeek V4 compresses the sequence itself. At 128K context, Llama 3.3 70B needs 40 GiB of KV per sequence; Qwen3.8 27B needs 8 GiB.
Can quantization make a model fit on fewer GPUs?
Yes. Weights take parameters × bits per weight, so going from BF16 to 8 bits halves them and 4-bit formats cut them by about 3.5×. As a rule of thumb, 8-bit is near lossless, 5 to 6-bit loses very little and 4-bit loses a little more, depending on the model. Set a fixed number of GPUs and the calculator shows the most accurate format that fits.
Does quantizing the weights shrink the KV cache?
No. The weight format (FP8, NVFP4, GGUF Q4_K_M…) and the KV cache precision are set independently. An FP8 KV cache halves the cache whatever the weights are stored in, at a small accuracy cost.
How much memory does fine-tuning take?
Full fine-tuning with AdamW in mixed precision costs about 16 bytes per parameter before activations: BF16 weights and gradients, an FP32 master copy and two FP32 optimizer moments. LoRA keeps the frozen model at 2 bytes per parameter, QLoRA at about 0.52. Activations then scale with sequence length × batch, not with parameter count.