BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

Ollama RC fixes MLX speculative-decoding memory retention

Sep 17, 2026, first seen via Training & inference tooling releases

Ollama’s v0.34.2-rc2 release candidate fixed retained MLX speculative-decoding buffers, keeping a long-context Qwen run near 30 GB instead of above 90 GB.

Read the original at github.comOpens the publisher's site in a new tab

More in Open-weight & owned stack