BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Improves CUDA Qwen Throughput With Kernel Fusion

Sep 25, 2026, first seen via Training & inference tooling releases

llama.cpp b11177 fused CUDA RMS normalization and scaling, reporting up to 4.8% higher Qwen3.8-27B throughput under draft-MTP decoding.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack