Open-weight & owned stack · Quantization efficiency
llama.cpp b10908 Improves Metal Quantized Kernels
llama.cpp b10908 improved Metal quantized-kernel utilization for several IQ formats and corrected row-pointer offsets.
Read the original at github.comOpens the publisher's site in a new tab