Open-weight & owned stack · Quantization efficiency
Developer Adds NVFP4 KV Cache Support For Qwen3.8
A developer reported adding NVFP4 KV-cache support and heterogeneous GPU placement for running Qwen3.8 27B at up to 262k context.
Read the original at reddit.comOpens the publisher's site in a new tab