Agents and harnesses · Quantization efficiency
vLLM Documents Prefill-Decode Serving For Qwen
vLLM documented prefill-decode serving for Qwen3.8-2.4T, including memory measurement for batches exceeding CUDA-graph limits.
Read the original at vllm.aiOpens the publisher's site in a new tab