BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

llama-manager Dynamically Reconfigures Local Inference

Sep 11, 2026

A developer released llama-manager, a llama.cpp wrapper and fork that dynamically adjusts context, speculative decoding, and multimodal settings.

Read the original at reddit.comOpens the publisher's site in a new tab

More in Open-weight & owned stack