Open-weight & owned stack · Quantization efficiency
Ninfer user reports 600-token Qwen throughput
A user reported serving Qwen3.6 35B-A3B at approximately 600 tokens per second for one request with Ninfer on an RTX Pro 6000.
Read the original at reddit.comOpens the publisher's site in a new tab