Open-weight & owned stack · Quantization efficiency
Study examines transforms for four-bit LLM quantization
A study evaluated rotations, scaling, permutations and affine transforms for improving four-bit language-model quantization.
Read the original at awesomepapers.ioOpens the publisher's site in a new tab