
9030 club reading "Reinforcement Learning via Self-Distillation": #63 - ML Paper Reading Group
☕️📝 Paper Link 📝☕️
Join us for an engaging reading session where we explore whether a large language model (LLM) can turn feedback on its failures into its own dense training signal, potentially surpassing standard reinforcement learning with verifiable rewards (RLVR) without the need for a teacher or reward model.
Abstract
Large language models increasingly utilize reinforcement learning in verified domains such as code and mathematics. Existing RLVR methods derive learning from a single scalar reward, which creates a credit-assignment challenge. However, many verified environments offer rich textual feedback explaining failures, enabling a new approach: Self-Distillation Policy Optimization (SDPO). This technique converts feedback into a concentrated learning signal autonomously, leveraging the model's capacity to identify its own mistakes in context.
Key Highlights:
- SDPO demonstrates enhanced sample efficiency and improved accuracy over strong RLVR benchmarks.
- The method successfully utilizes implicit feedback from successful rollouts for failed attempts, outperforming standard RLVR environments.
- At the test phase, applying SDPO to individual questions accelerates discovery in challenging binary-reward tasks, offering greater efficiency with lower attempt counts.
Event Schedule
- 7 PM - 8 PM: Quiet reading time 🍽️📚
- 8 PM - 9 PM: Open discussion about the reading 📝
- 9 PM: Extended time to socialize or network 🤝
Location
Mox SF
1680 Mission St,
San Francisco, CA 94103, USA
RSVP here — please let us know if you can join us!
1680 Mission St, San Francisco, CA 94103, USA
RouteMit der Kamera scannen – die Veranstaltung öffnet sich in der Somo-App.








