Open-weight & owned stack · Quantization efficiency
PagedWeight Dynamically Quantizes Mixture-Of-Experts Weights
PagedWeight proposed dynamically quantizing mixture-of-experts weights during serving to balance GPU memory against growing KV-cache demand.
Read the original at awesomepapers.ioOpens the publisher's site in a new tab