Reading Group (+🧋): Train-to-Test (T²) Scaling Laws

Reading Group (+🧋): Train-to-Test (T²) Scaling Laws

Om evenemanget

Join the Snorkel AI Reading Group, a recurring forum to explore the latest frontier developments in AI while building meaningful connections within the community.

In this afternoon's session, Nicholas Roberts will present his recent paper, Train-to-Test (T²) Scaling Laws: Test-Time Scaling Makes Overtraining Compute-Optimal, which will be featured at COLM 2026. You can find the arXiv preprint here.

Agenda:

  • 4 pm - Doors open
  • 4:30 pm - Talk begins

🧋🧋🧋 Boba tea and other refreshments will be provided! 🧋🧋🧋

Key Takeaways:

  • Understanding why pretraining scaling laws like Chinchilla don't account for test-time compute, and the trade-off that arises when inference cost scales with model size and sample count.
  • How Train-to-Test (T²) scaling laws modernize pretraining scaling laws with pass@k modeling, optimizing model size, training tokens, and inference samples under a fixed end-to-end budget.
  • Insights on why forecasts hold up across distinct modeling approaches, including the joint scaling effect on task loss and its impact on task accuracy.
  • Discovering why optimal pre-training decisions shift radically into the overtraining regime across eight downstream tasks, well outside the range of standard pre-training scaling suites.
  • Validation of these findings through pre-training heavily overtrained models in the region T² forecasts, confirming stronger performance even after post-training.

This work will be featured at COLM 2026 and has been covered by VentureBeat: Train-to-test scaling explained.

Hittad av Somo·Se original
Plats

101 Second Street, San Francisco, CA 94105, USA

Vägbeskrivning

Den här veckan i San Francisco