Open-weight & owned stack · Quantization efficiency
llama.cpp adds SpacemiT X60 acceleration
llama.cpp b11408 added SpacemiT X60 Q8_0 IME1 acceleration, raising measured throughput on a Milk-V Jupiter from 10.70 to 93.87 tokens per second.
Read the original at github.comOpens the publisher's site in a new tab