BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp Adds GLM-5.3-Flash Support

Sep 30, 2026, first seen via Training & inference tooling releases. Covered by 2 outlets.

llama.cpp release b11279 added GLM-5.3-Flash support alongside multi-token prediction, caching, quantization, and long-context improvements.

Read the original at github.comOpens the publisher's site in a new tab

Also covering this

More in Open-weight & owned stack