Open-weight & owned stack · Quantization efficiency
VarRate Introduces Variable-Rate KV Cache Compression
VarRate proposed training-free variable-rate KV-cache compression for long-context language models.
Read the original at awesomepapers.ioOpens the publisher's site in a new tab