Open-weight & owned stack · Quantization efficiency
llama.cpp Adds AMD GCN MMQ Configuration
A llama.cpp pull request adds AMD GCN MMQ support and reports prompt-processing improvements on RDNA2-based MI50 and MI60 accelerators.
Read the original at reddit.comOpens the publisher's site in a new tab