Open-weight & owned stack · Quantization efficiency
Community Speeds DeepSeek V4.1 Flash
A GitHub fork reportedly increased DeepSeek V4.1 Flash Q4 decoding on an M3 Ultra from 16.6 to 31.3 tokens per second.
Read the original at reddit.comOpens the publisher's site in a new tab