Skip to content

LLM VRAM Calculator

How much VRAM and how many GPUs your LLM needs, with the real KV cache of today's architectures. Pick a model or paste any Hugging Face link.

Trending on Hugging Face

Estimates, not guarantees. Real usage depends on your engine and settings: keep some headroom and confirm with nvidia-smi. Confirm with nvidia-smi.

Estimator
Clef
InferencevLLM1 × 80 GB

Your setup

Pick a model, the context and your GPUs.

Model

Paste a link or org/model: reads the model's config.json and safetensors metadata straight from Hugging Face.

Cloudflare, United States · fine-tune of Qwen3.8 27B

27.4B parameters · 256K context · Apache 2.0

64 layers · 16 full attention (4 KV × 256) · 48 linear attention

Effective bits per weight, quantization scales included. Native = exact size of the official checkpoint files.

Requests

Prompt + generated tokens
Requests decoded at the same time; each holds its own KV cache
Set independently of the weight format. FP8 halves the cache.

Hardware

As marketed. The calculation uses what the driver reports when it differs (an L40S sold as 48 GB exposes 44.99 GiB). Unified-memory presets already exclude what the OS keeps.
Auto picks the fewest GPUs that fit: tensor parallel inside a server (up to 8), pipeline parallel across servers.
Serving engine · vLLM
vLLM gpu_memory_utilization, SGLang mem_fraction_static… The rest stays free on purpose.
CUDA context, NCCL, CUDA graphs, activation workspace. An approximation: adjust it from a real nvidia-smi reading.

Reserves gpu_memory_utilization (0.90 by default) of each GPU, then fills what is left after weights and activations with paged KV cache.

How it is computed

weights = params × bits ÷ 8 ÷ (TP × PP)
KV = Σ layer groups · layers × width × cached tokens × bytes
cached = context · min(context, window) · context ÷ ratio
width = 2 × KV heads × head dim (GQA) · latent + RoPE (MLA)
(KV + state) × sequences + overhead ≤ memory × usable %

Assumes a modern engine with a paged KV cache that keeps only the window on sliding-window layers. Real usage can differ by a few GiB.

Your answer

How many GPUs, how fast, and what it costs.

To run it you need

Fits

1 × NVIDIA H100 · 80 GB

Clef, BF16 / FP16 weights, 4 concurrent requests of 32K tokens. 80 GB of GPU memory in total.

From $2.04/h to rent, on Vast.aiOn-demand prices, checked every hour. Prices change: check before you rent.Compare 4 providers ↓

Speed & cost

Per request Generation speed seen by one user, at full context≈ 42 tok/s
Total, 4 requests≈ 168 tok/s
First token, 32K prompt Time to read a prompt as long as the chosen context before the first token appears (prefill), for one request. It depends on the GPU's compute (TFLOPS), not its memory bandwidth.≈ 5.4 s
Your cost per 1M tokens Generated tokens at full load: 1 × NVIDIA H100 on Vast.ai$3.38

No API price for this model on OpenRouter.

Rough estimate (±30–50%) of generation at full context, from memory bandwidth (3.35 TB/s per GPU) and vLLM efficiency; first token from compute (989 TFLOPS). Speculative decoding is not modelled. API prices: OpenRouter, per generated token.

VRAM per GPU

  • Weights51.0 GiB
  • KV cache8.00 GiB
  • Recurrent state0.57 GiB
  • Runtime overhead3.00 GiB
Needed62.5 GiB
Usable Memory per GPU × usable %71.7 GiB
Headroom9.16 GiB

Capacity

Max requests at 32K8
Max context with 4 requests70,269
KV per token Averaged over the chosen context, whole model. Falls as context grows on sliding-window and compressed layers.64.0 KiB
KV per request At the chosen context, whole model, before tensor-parallel sharding2.00 GiB

Treating every layer as full attention would give 8.00 GiB per request, 4.0× more than this architecture actually caches.

Plus 146.8 MiB of recurrent state per request, fixed whatever the context.

Compare with your real usage

Paste the startup log of vLLM or llama.cpp (llama-server, Ollama, LM Studio). It is read in your browser: nothing is sent.

Better quality or a cheaper GPU?

Pick a more faithful weight format (it may need more GPUs) or a cheaper card. Click an option to apply it.

Does it fit? Weights × context on 1 × NVIDIA H100

Weights4K8K32K128K256K
BF16 / FP16
FP8
GGUF Q8_0
GGUF Q6_K
GGUF Q5_K_M
GGUF Q4_K_M
GGUF Q3_K_M

Longest context that fits

  • BF16 / FP16 up to 32K
  • FP8 to GGUF Q3_K_M up to 128K

✓ fits · ~ tight · ✕ does not fit. Context per request, 4 concurrent requests, vLLM. Click a cell to apply it. Tap a format to use it at its longest context.

Quantization on NVIDIA H100

1 GPU

Fewer bits per weight, fewer GPUs, at some quality cost. The KV cache is set separately.

Cheapest GPUs that fit in one server

Cheapest on-demand price among RunPod, Verda, Vast.ai, Microsoft Azure, checked every hour. Weight format as selected above.

Other models that fit on 1 × NVIDIA H100

Largest first, each with the most faithful weight format that fits 32K tokens × 4 requests.

Rent and deploy

Today's NVIDIA H100 prices per provider, then the command to launch it.

NVIDIA H100 rental prices

ProviderPer GPU-hour1 GPU, per monthRent
Vast.ai$2.04$1,491.39Rent on Vast.ai (opens in a new tab) →
Verda$3.93$2,867.44Rent on Verda (opens in a new tab) →
RunPod$3.99$2,912.70Rent on RunPod (opens in a new tab) →
Microsoft Azure$12.29$8,971.70Rent on Microsoft Azure (opens in a new tab) →

Loading the daily history…

On-demand, per GPU, checked every hour (last check 2026-10-10 14:06 UTC). Monthly = 730 hours. NVIDIA H100 price page

Deploy it

vllm serve Cloudflare/clef \
  --tensor-parallel-size 1 \
  --max-model-len 32768 \
  --max-num-seqs 4 \
  --gpu-memory-utilization 0.90 \
  --dtype bfloat16

Some rental links are affiliate links: StudioTV may earn a commission, at no extra cost to you.

Fits: 1 × NVIDIA H100, $2.04/h. Go to the answer

New & trending open models

Lab countries from the registry of this site, completed with Epoch AI (opens in a new tab) (CC BY 4.0).

Use it from your AI assistant or your code

Ask Claude, ChatGPT or any MCP client whether a model fits on your GPU: it calls this calculator and answers with the same numbers, for the listed models and any Hugging Face model. No key, no sign-up.

MCP server (Streamable HTTP): https://studiotvai.com/api/mcp
Tools: estimate_vram, models_that_fit, gpu_prices

Claude Code: claude mcp add --transport http studiotv-vram https://studiotvai.com/api/mcp

The same data as JSON: /api/models.json (every model, with the link to its file), /api/gpus.json, and one file per model such as /api/models/llama-3-3-70b.json. Each model page also has a VRAM badge to paste in a model card.

VRAM requirements by model

GPU memory for the weights alone, in GiB. Each model page details the KV cache by context length and how many GPUs of each type it takes.

ModelParametersBF16FP84-bit
Dense
Meta, United StatesLlama 3.1 8B VRAM8.0B15.08.54.6
Mistral AI, FranceMistral Small 3.2 24B VRAM24B44.723.613.4
Alibaba, ChinaQwen3 32B VRAM33B61.032.018.6
Meta, United StatesLlama 3.3 70B VRAM71B13167.739.9
Meta, United StatesLlama 3.1 405B VRAM406B756382229
Hybrid linear attention
Alibaba, ChinaQwen3.5 4B VRAM4.7B8.54.82.5
Alibaba, ChinaQwen3.5 9B VRAM9.7B17.510.75.2
Alibaba, ChinaQwen3.8 27B VRAM28B51.027.815.4
Alibaba, ChinaQwen3.6 35B-A3B VRAM36B (3B active)65.433.619.6
Alibaba, ChinaQwen3.5 122B-A10B VRAM125B (10B active)22811668.8
Alibaba, ChinaQwen3.8-Flash-Next VRAM180B (6B active)330166105
Liquid AI, United StatesLFM2.5 8B-A1B VRAM8.5B (1.5B active)15.88.14.8
Liquid AI, United StatesLFM2 24B-A2B VRAM24B (2.3B active)44.422.313.4
DeepReinforce, United StatesOrnith 1.5 35B-A3B VRAM36B (3B active)65.433.619.6
DeepReinforce, United StatesOrnith 1.5 9B VRAM9.7B17.510.75.2
Cloudflare, United StatesClef VRAMNew27B51.027.815.4
Cloudflare, United StatesClef Flash VRAMNew9.4B17.510.75.2
Liquid AI, United StatesD1 3B VRAMNew3.1B5.83.41.6
AutoTrust AI, SingaporeJEV 27B VL VRAMNew28B51.027.815.4
LightOn, FranceLightOnOCR-3 4B VRAMNew4.5B8.54.82.5
Agnes AI, SingaporeAgnes-3.0-Qwen VRAMNew180B (7.8B active)330166105
Sliding-window attention
Google, United StatesGemma 4 12B VRAM12B22.312.16.9
Google, United StatesGemma 4 31B VRAM31B58.230.417.5
Google, United StatesGemma 4 26B-A4B VRAM26B (4B active)48.124.715.4
OpenAI, United Statesgpt-oss 20B VRAM21B (3.6B active)38.920.610.8
OpenAI, United Statesgpt-oss 120B VRAM117B (5.1B active)21811058.3
Meta, United StatesLlama 4 Scout 17B-16E VRAM109B (17B active)20210360.8
Meta, United StatesLlama 4 Maverick 17B-128E VRAM402B (17B active)748376226
Aleph Alpha, GermanyKolibri-1 VRAMNew78B (3.46B active)14573.344.0
Country not known yetHumanizer VRAMNew12B22.313.06.9
JetBrains, CzechiaMellum2.1 12B-A2.5B VRAMNew12B (2.5B active)22.611.77.4
iFlytek, ChinaSpark-X2.5 4B VRAM4.1B7.74.12.4
Hybrid Mamba
NVIDIA, United StatesNemotron 3 Nano 30B-A3B VRAM32B (3B active)58.830.122.0
NVIDIA, United StatesNemotron 3 Super 120B-A12B VRAM124B (12B active)22511375.5
Large MoE
MiniMax, ChinaMiniMax M2.7 VRAM229B (10B active)426214129
DeepSeek, ChinaDeepSeek R1 VRAM684B (37B active)1250627377
Moonshot AI, ChinaKimi K2.6 VRAM1027B (32B active)1913959577
Z.ai, ChinaGLM-5.3 VRAM753B (40.3B active)1385694418
Z.ai, ChinaGLM-5.3-Flash VRAM321B (18B active)585294176
DeepSeek, ChinaDeepSeek V4 Flash VRAM291B (13B active)530266160
DeepSeek, ChinaDeepSeek V4 Pro VRAM1599B (49B active)29301467885
DeepSeek, ChinaDeepSeek V4.1 Flash VRAMNew763B (16B active)1395699421
Xiaomi, ChinaMiMo-V2.6 Pro VRAMNew1024B (42B active)1904954574
Xiaomi, ChinaMiMo-V2.6 Flash VRAMNew311B (15B active)577290174
Moonshot AI, ChinaKimi K3 VRAM2780B (104B active)517825911563
Mistral AI, FranceMistral Large 4 VRAMNew1052B (49B active)1959981591
DeepSeek, ChinaDeepSeek V4 Flash Vision Exp VRAM305B (15B active)530266160
Alibaba, ChinaQwen3.8 2.4T-A95B VRAM2446B (95B active)450722571361
DeepReinforce, United StatesOrnith 1.5 397B VRAM403B (18B active)739371223
Shanghai AI Lab, ChinaAtria Dawn Preview VRAMNew753B (42B active)1385694418
Shanghai AI Lab, ChinaIntern S2 397B VRAMNew403B (18B active)739371223
Coding models
Alibaba, ChinaQwen3-Coder-Next 80B-A3B VRAM80B (3B active)14874.844.9
Alibaba, ChinaQwen3-Coder 30B-A3B VRAM31B (3.3B active)56.929.017.2
Alibaba, ChinaQwen3-Coder 480B-A35B VRAM480B (35B active)894449270

LLM VRAM FAQ

How much VRAM do I need to run a 70B model?

Take Llama 3.3 70B. Its weights need 131 GiB in BF16 (2 × H100 80 GB), 67.7 GiB in FP8 and about 39.9 GiB in 4-bit (1 × H100, or 2 × RTX 4090). Then add the KV cache of each request: 2.5 GiB at 8K tokens and 10.0 GiB at 32K.

Can I run an LLM on a 24 GB GPU like the RTX 4090?

Yes, models up to about 35B parameters fit in 4-bit with an 8K-token context: Llama 3.1 8B, Mistral Small 3.2 24B, Qwen3 32B, Qwen3.5 4B, Qwen3.5 9B, Qwen3.8 27B, Qwen3.6 35B-A3B, LFM2.5 8B-A1B, LFM2 24B-A2B, Gemma 4 12B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss 20B, Qwen3-Coder 30B-A3B, Ornith 1.5 35B-A3B, Ornith 1.5 9B, Clef, Clef Flash, D1 3B, Humanizer, JEV 27B VL, Mellum2.1 12B-A2.5B, LightOnOCR-3 4B, Spark-X2.5 4B. Larger models need several GPUs or a card with more memory.

How many GPUs do I need to run an LLM?

Each GPU must hold its share of the weights plus the KV cache of every concurrent request, under the memory the serving engine allows itself (90% by default in vLLM). The calculator tries 1, 2, 4, 8 GPUs and more, splitting the model with tensor parallelism inside a server and pipeline parallelism across servers, and returns the fewest that fit.

What drives LLM VRAM usage?

Four things decide how much memory a model needs once it is serving requests.

  • Weights are fixed: parameters × bits. “4-bit” formats really cost 4.2 to 4.9 bits once scales are counted.
  • KV cache grows with context × concurrent sequences; it decides how many users a GPU can serve.
  • Tensor parallelism splits GQA heads but never below one head per GPU, and replicates MLA latents.
  • Serving engines leave memory free on purpose: vLLM uses 90% of each GPU by default.

Why is VRAM higher than the model size on disk?

The file holds only the weights. At serving time you also pay for the KV cache, which grows with context × concurrent sequences, runtime buffers such as CUDA graphs and activation workspace, and the share of memory the engine leaves free on purpose.

Why are most LLM VRAM formulas wrong?

The classic estimate, 2 × layers × KV heads × head dim × bytes per token, assumes every layer attends to the whole context. Most models released since 2025 break that rule:

  • Hybrid linear attention (Qwen 3.5–3.8, Nemotron 3, Kimi K3): only 1 layer in 4, or fewer, keeps a KV cache; the rest hold a fixed-size state.
  • Sliding windows (Gemma 4, gpt-oss): most layers keep only the last 128 to 1,024 tokens.
  • MLA (GLM-5.3, Kimi K3): one 576-wide compressed latent per token instead of keys and values.
  • Compressed attention (DeepSeek V4): a 128-token window plus the rest of the sequence compressed 4× or 128×.

Why does the KV cache differ so much between models of the same size?

2026 architectures cache very different amounts per token. Qwen 3.5+ keeps a KV cache on only 1 layer in 4, Gemma 4 and gpt-oss limit most layers to a short sliding window, MLA models store one compressed latent instead of keys and values, and DeepSeek V4 compresses the sequence itself. At 128K context, Llama 3.3 70B needs 40 GiB of KV per sequence; Qwen3.8 27B needs 8 GiB.

Can quantization make a model fit on fewer GPUs?

Yes. Weights take parameters × bits per weight, so going from BF16 to 8 bits halves them and 4-bit formats cut them by about 3.5×. As a rule of thumb, 8-bit is near lossless, 5 to 6-bit loses very little and 4-bit loses a little more, depending on the model. Set a fixed number of GPUs and the calculator shows the most accurate format that fits.

Does quantizing the weights shrink the KV cache?

No. The weight format (FP8, NVFP4, GGUF Q4_K_M…) and the KV cache precision are set independently. An FP8 KV cache halves the cache whatever the weights are stored in, at a small accuracy cost.

How much memory does fine-tuning take?

Full fine-tuning with AdamW in mixed precision costs about 16 bytes per parameter before activations: BF16 weights and gradients, an FP32 master copy and two FP32 optimizer moments. LoRA keeps the frozen model at 2 bytes per parameter, QLoRA at about 0.52. Activations then scale with sequence length × batch, not with parameter count.