Open-weight & owned stack · Quantization efficiency
Splash Demonstrates Chip-Specific Kernel Tuning Gains
Splash reportedly increased decode speed by 20% on an M5 Max by selecting chip-specific matrix-multiplication layouts.
Read the original at reddit.comOpens the publisher's site in a new tab