9030 club reading "Reinforcement Learning via Self-Distillation": #63 - ML Paper Reading Group

9030 club reading "Reinforcement Learning via Self-Distillation": #63 - ML Paper Reading Group

About the event

โ˜•๏ธ๐Ÿ“ Paper Link ๐Ÿ“โ˜•๏ธ

Join us for an engaging reading session where we explore whether a large language model (LLM) can turn feedback on its failures into its own dense training signal, potentially surpassing standard reinforcement learning with verifiable rewards (RLVR) without the need for a teacher or reward model.

Abstract

Large language models increasingly utilize reinforcement learning in verified domains such as code and mathematics. Existing RLVR methods derive learning from a single scalar reward, which creates a credit-assignment challenge. However, many verified environments offer rich textual feedback explaining failures, enabling a new approach: Self-Distillation Policy Optimization (SDPO). This technique converts feedback into a concentrated learning signal autonomously, leveraging the model's capacity to identify its own mistakes in context.

Key Highlights:

  • SDPO demonstrates enhanced sample efficiency and improved accuracy over strong RLVR benchmarks.
  • The method successfully utilizes implicit feedback from successful rollouts for failed attempts, outperforming standard RLVR environments.
  • At the test phase, applying SDPO to individual questions accelerates discovery in challenging binary-reward tasks, offering greater efficiency with lower attempt counts.

Event Schedule

  • 7 PM - 8 PM: Quiet reading time ๐Ÿฝ๏ธ๐Ÿ“š
  • 8 PM - 9 PM: Open discussion about the reading ๐Ÿ“
  • 9 PM: Extended time to socialize or network ๐Ÿค

Location

Mox SF
1680 Mission St,
San Francisco, CA 94103, USA

RSVP here โ€” please let us know if you can join us!

Found by SomoยทSee original
Location

1680 Mission St, San Francisco, CA 94103, USA

Get directions

This week in San Francisco