Open-weight & owned stack · Quantization efficiency
Fork boosts Qwen decoding on dual Radeon GPUs
A llama.cpp fork reportedly increased Qwen 3.8 Q8 decoding from about 28 to 82 tokens per second on two Radeon RX 7900 XTX GPUs.
Read the original at reddit.comOpens the publisher's site in a new tab