BLAST RADIUS

Open-weight & owned stack · Quantization efficiency

Strata Reports Fast Qwen Flash Next Inference

Sep 24, 2026

A developer reported running quantized Qwen 3.8 Flash Next at up to 65.1 output tokens per second on a 12GB RTX 5070.

Read the original at reddit.comOpens the publisher's site in a new tab

More in Open-weight & owned stack