Open-weight & owned stack · Training tuning
RM-Distiller Uses Generative LLMs For Reward Distillation
RM-Distiller proposes using generative language models to distill reward models and reduce reliance on human preference data.
Read the original at lacuna.tiptreesystems.comOpens the publisher's site in a new tab