Qwen2.5-Coder 32B has 32.8B parameters. In BF16 its weights alone take 61.0 GiB; quantized to 4-bit, about 18.6 GiB. Each 32K-token request adds 8.00 GiB of KV cache. Cheapest way to run it today: 1 × GeForce RTX 3090 (24 GB) at $0.143 per hour.
Weights plus the KV cache of one request (BF16 cache), in GiB. Add 1 to 3 GiB for the inference engine itself.
Weights format
Weights only
+ 8K context
+ 32K context
BF16 / FP16
61.0
63.0
69.0
FP8 / INT8
32.0
34.0
40.0
4-bit (GGUF Q4_K_M)
18.6
20.6
26.6
How many GPUs to run Qwen2.5-Coder 32B
Fewest GPUs that fit the weights plus one 8K-token request. Data-center GPUs use vLLM (90% of memory usable), the others llama.cpp. Several GPUs split the model with tensor or pipeline parallelism.
GPU
BF16 / FP16
8-bit
4-bit
GeForce RTX 4090 24 GB
3
2
1
GeForce RTX 5090 32 GB
3
2
1
RTX PRO 6000 Blackwell 96 GB
1
1
1
Mac M5 Max 128 GB 96 GB usable
1
1
1
NVIDIA H100 80 GB
1
1
1
NVIDIA H200 141 GB
1
1
1
NVIDIA B200 (HGX) 180 GB
1
1
1
8-bit: FP8 with vLLM (INT8 on GPUs without FP8), Q8_0 with llama.cpp. 4-bit: AWQ / GPTQ with vLLM, Q4_K_M with llama.cpp, both counted at the Q4_K_M size.
Cheapest way to run Qwen2.5-Coder 32B today
For each GPU: the fewest cards that fit Qwen2.5-Coder 32B with one 8K-token request, the most faithful weight format at that count, and the cheapest on-demand price (2026-10-11).
On-demand prices from RunPod, Vast.ai, Verda and Azure, checked every hour. Serving many users needs more KV cache, so more memory: size it in the calculator. Some links are affiliate links: StudioTV may earn a commission, at no extra cost to you.
KV cache: how memory grows with context
Architecture: 64 layers · 64 full attention (8 KV × 128). Every layer keeps keys and values for every token (grouped-query attention), so the cache grows linearly with context.
Context per request
BF16 cache
FP8 cache
8K tokens
2.00 GiB
1.00 GiB
32K tokens
8.00 GiB
4.00 GiB
Qwen2.5-Coder 32B VRAM FAQ
How much VRAM does Qwen2.5-Coder 32B need?
In BF16 the weights alone take 61.0 GiB (66 GB). In FP8 that is 32.0 GiB, and about 18.6 GiB with 4-bit quantization (Q4_K_M). Each request then adds KV cache: 2.00 GiB at 8K tokens and 8.00 GiB at 32K (BF16 cache).
Can Qwen2.5-Coder 32B run on a single RTX 4090 (24 GB)?
Yes. In Q4_K_M it fits on one RTX 4090 with an 8K-token context (20.6 GiB including the cache, with llama.cpp).
What is the cheapest way to run Qwen2.5-Coder 32B?
On 2026-10-11, the cheapest on-demand setup is 1 × GeForce RTX 3090 (24 GB) with Q4_K_M weights on Vast.ai, at $0.143 per hour (about $104.39 per month). Next: 1 × GeForce RTX 4090 (24 GB) with Q4_K_M weights on Vast.ai, at $0.356 per hour. Sized for one 8K-token request; prices are checked every hour.
How many H100 GPUs does Qwen2.5-Coder 32B need?
With one request and an 8K-token context: 1 in BF16, 1 in FP8 and 1 in 4-bit (vLLM, 90% of memory usable). Serving many users at once needs more memory for their KV caches: the calculator sizes that for you.
Qwen2.5-Coder 32B VRAM badge
For a model card or a README: the 4-bit size and the smallest GPU it fits on, linked to this page. Data as JSON: /api/models/qwen2-5-coder-32b.json.
Estimates, not guarantees: computed from the official config.json with the same engine as the LLM VRAM Calculator. Real usage depends on your engine version and settings.