Open-weight & owned stack · Quantization efficiency
Qwen3.8 Next Runs On Six V100 GPUs
A user reported running Qwen3.8 Next on six V100 GPUs at 22.59 tokens per second without speculative decoding and 42.70 with it enabled.
Read the original at reddit.comOpens the publisher's site in a new tab