BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Improves MUSA FlashAttention And Accelerator Support

Sep 25, 2026, first seen via Training & inference tooling releases

llama.cpp b11178 improved MUSA FlashAttention performance and enabled CUB-based sorting operations on Moore Threads MTT S5000 hardware.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack