Open-weight & owned stack · Quantization efficiency
llama.cpp Adds IQ3S Vulkan Kernels
llama.cpp releases b11035 through b11037 added IQ3_S MMQ Vulkan kernels and related two-byte shared-memory loading changes.
Read the original at github.comOpens the publisher's site in a new tabAlso covering this
b11037github.com, Sep 18b11036github.com, Sep 18b11034github.com, Sep 18b11030github.com, Sep 18b11028github.com, Sep 17b11027github.com, Sep 17