Open-weight & owned stack · Quantization efficiency
SPARQLe Proposes Sub-Precision Activations For LLM Inference
SPARQLe proposes sub-precision activation representations to reduce compute and memory costs during quantized language-model inference.
Read the original at awesomepapers.ioOpens the publisher's site in a new tab