BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Releases Version 0.5.0

Sep 23, 2026, first seen via Training & inference tooling releases

llama.cpp v0.5.0 added CUDA and Metal optimizations, multi-address server binding, function-call image outputs, and additional model architectures.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack