BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp adds Vulkan sparse attention optimizations

Oct 5, 2026, first seen via Training & inference tooling releases

llama.cpp releases b11407 through b11415 added Vulkan sparse flash attention for quantized key-value caches and improved sparse-attention index compaction.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack