Open-weight & owned stack · Quantization efficiency
llama.cpp b10909 Fixes Metal Regression
llama.cpp b10909 reworked Metal fusion tables, fixed a reported token-generation regression, and added cache fusion.
Read the original at github.comOpens the publisher's site in a new tab