Open-weight & owned stack · Quantization efficiency
Qwen3.8 Flash Next Runs Across Three RTX 3090s
A user reported running quantized Qwen3.8 Flash Next across three RTX 3090 GPUs at roughly 38–50 output tokens per second.
Read the original at reddit.comOpens the publisher's site in a new tab