Open-weight & owned stack · Quantization efficiency
Qwen3.8 Runs On Sixteen-Gigabyte Consumer GPU
A user reported running quantized Qwen3.8 27B at about 30 tokens per second with 100K context on a 16GB Radeon RX 7800 XT.
Read the original at reddit.comOpens the publisher's site in a new tab