Open-weight & owned stack · Quantization efficiency
llama.cpp Adds AMD CDNA Fattn-MMA Kernel Support
llama.cpp release b11214 enables the fattn-mma kernel on AMD CDNA for dkq above 256 and large batch sizes.
Read the original at github.comOpens the publisher's site in a new tab