Open-weight & owned stack · Quantization efficiency
LoRDS Unifies LLM Quantization And Adaptation
LoRDS proposed replacing block-wise scaling with continuous low-rank matrices to combine LLM quantization and adaptation.
Read the original at lacuna.tiptreesystems.comOpens the publisher's site in a new tab