Open-weight & owned stack · Training tuning
Cross-Tokenizer Scoring Supports Language-Model Distillation
An ICLR 2026 paper introduced algorithms for cross-tokenizer likelihood ratios in language-model distillation and related training methods.
Read the original at mlanthology.orgOpens the publisher's site in a new tab