BLAST RADIUS

Agents and harnesses · Quantization efficiency

vLLM Documents Prefill-Decode Serving For Qwen

Sep 21, 2026

vLLM documented prefill-decode serving for Qwen3.8-2.4T, including memory measurement for batches exceeding CUDA-graph limits.

Read the original at vllm.aiOpens the publisher's site in a new tab

More in Agents and harnesses