Qwen2.5-Coder 7B has 7.62B parameters. In BF16 its weights alone take 14.2 GiB; quantized to 4-bit, about 4.39 GiB. Each 32K-token request adds 1.75 GiB of KV cache. Cheapest way to run it today: 1 × GeForce RTX 3090 (24 GB) at $0.143 per hour.
Read every day from Hugging Face. Split files are added together; vision projectors are left out.
How many GPUs to run Qwen2.5-Coder 7B
Fewest GPUs that fit the weights plus one 8K-token request. Data-center GPUs use vLLM (90% of memory usable), the others llama.cpp. Several GPUs split the model with tensor or pipeline parallelism.
GPU
BF16 / FP16
8-bit
4-bit
GeForce RTX 4090 24 GB
1
1
1
GeForce RTX 5090 32 GB
1
1
1
RTX PRO 6000 Blackwell 96 GB
1
1
1
Mac M5 Max 128 GB 96 GB usable
1
1
1
NVIDIA H100 80 GB
1
1
1
NVIDIA H200 141 GB
1
1
1
NVIDIA B200 (HGX) 180 GB
1
1
1
8-bit: FP8 with vLLM (INT8 on GPUs without FP8), Q8_0 with llama.cpp. 4-bit: AWQ / GPTQ with vLLM, Q4_K_M with llama.cpp, both counted at the Q4_K_M size.
Cheapest way to run Qwen2.5-Coder 7B today
For each GPU: the fewest cards that fit Qwen2.5-Coder 7B with one 8K-token request, the most faithful weight format at that count, and the cheapest on-demand price (2026-10-11).
On-demand prices from RunPod, Vast.ai, Verda and Azure, checked every hour. Serving many users needs more KV cache, so more memory: size it in the calculator. Some links are affiliate links: StudioTV may earn a commission, at no extra cost to you.
KV cache: how memory grows with context
Architecture: 28 layers · 28 full attention (4 KV × 128). Every layer keeps keys and values for every token (grouped-query attention), so the cache grows linearly with context.
Context per request
BF16 cache
FP8 cache
8K tokens
0.44 GiB
0.22 GiB
32K tokens
1.75 GiB
0.88 GiB
Qwen2.5-Coder 7B VRAM FAQ
How much VRAM does Qwen2.5-Coder 7B need?
In BF16 the weights alone take 14.2 GiB (15 GB). In FP8 that is 8.11 GiB, and about 4.39 GiB with 4-bit quantization (Q4_K_M). Each request then adds KV cache: 0.44 GiB at 8K tokens and 1.75 GiB at 32K (BF16 cache).
Can Qwen2.5-Coder 7B run on a single RTX 4090 (24 GB)?
Yes. In BF16 it fits on one RTX 4090 with an 8K-token context (14.6 GiB including the cache, with llama.cpp).
What is the cheapest way to run Qwen2.5-Coder 7B?
On 2026-10-11, the cheapest on-demand setup is 1 × GeForce RTX 3090 (24 GB) with BF16 weights on Vast.ai, at $0.143 per hour (about $104.39 per month). Next: 1 × NVIDIA L4 (24 GB) with BF16 weights on Vast.ai, at $0.268 per hour. Sized for one 8K-token request; prices are checked every hour.
How many H100 GPUs does Qwen2.5-Coder 7B need?
With one request and an 8K-token context: 1 in BF16, 1 in FP8 and 1 in 4-bit (vLLM, 90% of memory usable). Serving many users at once needs more memory for their KV caches: the calculator sizes that for you.
Qwen2.5-Coder 7B VRAM badge
For a model card or a README: the 4-bit size and the smallest GPU it fits on, linked to this page. Data as JSON: /api/models/qwen2-5-coder-7b.json.
Estimates, not guarantees: computed from the official config.json with the same engine as the LLM VRAM Calculator. Real usage depends on your engine version and settings.