
Reading Group (+🧋): Train-to-Test (T²) Scaling Laws
O události
Join the Snorkel AI Reading Group, a recurring forum to explore the latest frontier developments in AI while building meaningful connections within the community.
In this afternoon's session, Nicholas Roberts will present his recent paper, Train-to-Test (T²) Scaling Laws: Test-Time Scaling Makes Overtraining Compute-Optimal, which will be featured at COLM 2026. You can find the arXiv preprint here.
Agenda:
- 4 pm - Doors open
- 4:30 pm - Talk begins
🧋🧋🧋 Boba tea and other refreshments will be provided! 🧋🧋🧋
Key Takeaways:
- Understanding why pretraining scaling laws like Chinchilla don't account for test-time compute, and the trade-off that arises when inference cost scales with model size and sample count.
- How Train-to-Test (T²) scaling laws modernize pretraining scaling laws with pass@k modeling, optimizing model size, training tokens, and inference samples under a fixed end-to-end budget.
- Insights on why forecasts hold up across distinct modeling approaches, including the joint scaling effect on task loss and its impact on task accuracy.
- Discovering why optimal pre-training decisions shift radically into the overtraining regime across eight downstream tasks, well outside the range of standard pre-training scaling suites.
- Validation of these findings through pre-training heavily overtrained models in the region T² forecasts, confirming stronger performance even after post-training.
This work will be featured at COLM 2026 and has been covered by VentureBeat: Train-to-test scaling explained.
Místo
101 Second Street, San Francisco, CA 94105, USA
TrasaPodrobnosti
Otevřít v mobilu
Naskenujte fotoaparátem – akce se otevře v aplikaci Somo.









