Open-weight & owned stack · Quantization efficiency
FreeToken fork adds DeepSeek vision and decoding
A community FreeToken fork added DeepSeek-V4.1-Flash, vision, speculative decoding and tensor parallelism with reported speeds up to 48 tokens per second on two RTX 3090 GPUs.
Read the original at reddit.comOpens the publisher's site in a new tab