BLAST RADIUS

Open-weight & owned stack · Training tuning

Pretraining Filtering Builds Open-Weight Model Safeguards

Sep 12, 2026

An ICLR 2026 paper finds that filtering pretraining data can create tamper-resistant safeguards in open-weight language models.

Read the original at mlanthology.orgOpens the publisher's site in a new tab

More in Open-weight & owned stack