Open-weight & owned stack · Quantization efficiency
vLLM 0.29.0 Makes Model Runner V2 Default
vLLM 0.29.0 made Model Runner V2 the default and added model, quantization, speculative-decoding, KV-cache, and GPU-kernel support.
Read the original at github.comOpens the publisher's site in a new tab