BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Ships Tiled Quantized Matrix Multiplication

Sep 26, 2026, first seen via Training & inference tooling releases

llama.cpp release b11195 added tiled matrix multiplication for k-quants, reporting 3–6-times faster large matmuls while reducing GEMV performance to 80%.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack