Open-weight & owned stack · Quantization efficiency
Qwen3.8-27B Achieves High Throughput On Two GPUs
A user reported serving an INT4 Qwen3.8-27B model on two RTX 3090s with full 262k context and high decode and prefill throughput.
Read the original at reddit.comOpens the publisher's site in a new tab