BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Enables Additional CUDA DUP Operations

Sep 15, 2026, first seen via Training & inference tooling releases

llama.cpp release b10975 enabled i16 and i32 DUP operations on CUDA.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack