9030 club reading "Reinforcement Learning via Self-Distillation": #63 - ML Paper Reading Group

9030 club reading "Reinforcement Learning via Self-Distillation": #63 - ML Paper Reading Group

About the event

β˜•οΈπŸ“ Paper Link πŸ“β˜•οΈ

Join us for an engaging reading session where we explore whether a large language model (LLM) can turn feedback on its failures into its own dense training signal, potentially surpassing standard reinforcement learning with verifiable rewards (RLVR) without the need for a teacher or reward model.

Abstract

Large language models increasingly utilize reinforcement learning in verified domains such as code and mathematics. Existing RLVR methods derive learning from a single scalar reward, which creates a credit-assignment challenge. However, many verified environments offer rich textual feedback explaining failures, enabling a new approach: Self-Distillation Policy Optimization (SDPO). This technique converts feedback into a concentrated learning signal autonomously, leveraging the model's capacity to identify its own mistakes in context.

Key Highlights:

  • SDPO demonstrates enhanced sample efficiency and improved accuracy over strong RLVR benchmarks.
  • The method successfully utilizes implicit feedback from successful rollouts for failed attempts, outperforming standard RLVR environments.
  • At the test phase, applying SDPO to individual questions accelerates discovery in challenging binary-reward tasks, offering greater efficiency with lower attempt counts.

Event Schedule

  • 7 PM - 8 PM: Quiet reading time πŸ½οΈπŸ“š
  • 8 PM - 9 PM: Open discussion about the reading πŸ“
  • 9 PM: Extended time to socialize or network 🀝

Location

Mox SF
1680 Mission St,
San Francisco, CA 94103, USA

RSVP here β€” please let us know if you can join us!

Found by SomoΒ·See original
Location

1680 Mission St, San Francisco, CA 94103, USA

Get directions

This week in San Francisco