Open-weight & owned stack · Quantization efficiency
Studies Explore Transform Coding For KV-Cache Compression
Two papers examined transform coding and rotated binary quantization as methods for compressing key-value caches during long-context LLM inference.
Read the original at awesomepapers.ioOpens the publisher's site in a new tab