Open-weight & owned stack · Quantization efficiency
FreeToken Enables Qwen3.8 Flash Next Caching
A user reported running Qwen3.8-Flash-Next on one RTX 5090 with FreeToken expert caching at up to roughly 60 generation tokens per second.
Read the original at reddit.comOpens the publisher's site in a new tab