Open-weight & owned stack · Quantization efficiency
Qwen FN shows marginal gains on Strix hardware
A user reported running QFN on Strix hardware at 30–40 tokens per second while finding its usefulness only marginally above a 27B model.
Read the original at reddit.comOpens the publisher's site in a new tab