Open-weight & owned stack · Quantization efficiency
Nunchux Introduces Low-Bit VC-Attention Kernel
Nunchux introduced a training-free low-bit attention kernel that it says accelerates video diffusion and MiniMax-H3 workloads on B200 GPUs.
Read the original at marktechpost.comOpens the publisher's site in a new tab