Open-weight & owned stack · Training tuning
Study Evaluates Distillation Models As Diffusion-RL Samplers
A paper evaluates distilled models as samplers for diffusion reinforcement learning and preference alignment.
Read the original at lacuna.tiptreesystems.comOpens the publisher's site in a new tab