BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama.cpp accelerates vision prompt processing

Sep 28, 2026, first seen via Training & inference tooling releases

llama.cpp release b11227 reduced a reported 24-image prompt from 13,377 milliseconds to 1,759 milliseconds on an RTX 4090.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack