Open-weight & owned stack · Quantization efficiency
llama.cpp Proposes Maple Ternary MoE Support
A llama.cpp pull request proposed CPU and low-VRAM support for DeepGrove’s preview Maple 20B-A1B ternary MoE architecture.
Read the original at reddit.comOpens the publisher's site in a new tab