
Daytona AI Researchers - Stanford, July 2026
Daytona AI Researchers - Stanford, July 21, 2026
On Tuesday, July 21, Daytona and FounderCoHo are again co-hosting an exclusive, high-signal evening dedicated to researchers at Stanford University to explore when we take long-horizon, stateful agents seriously — from the infrastructure that makes them possible, to the evaluation frameworks that make them trustworthy.
Agenda
🕒 5:30 pm – 6:00 pm
Welcome Reception and Opening Remarks
🎤 Marijan Cipcic, Principal Events Manager at Daytona
🕒 6:00 pm – 6:15 pm
Talk "Today's Agents Don't Live In Episodes"
🎤 Muhammad Annas Hashmi, DevRel at Daytona
Outline:
The 'episode' (short, stateless, resettable) has been RL's foundational abstraction since ATARI. It underpins the Gym API, GRPO, PPO, and the conventional sandbox lifecycle. Today's agents no longer fit it. Tasks span for days; the env state at hour 18 of an agent session with warm caches, installed deps, live processes, open sockets, dirty git tree, is worth hours of wall clock to reproduce.
Three things are scaling simultaneously. Rollout horizon: seconds -> days. Env state: disposable between episodes -> first-class learning substrate. Branching: absent in modern LLM-RL -> speculative fork trees. Each stresses the inherited toolkit in a different way, and all three have been gated on the same missing primitives: VMs you can fork cheaply, pause without killing processes, snapshot mid-run, and resume hours later.
This talk walks through what opens up when those primitives become available. Live demo of long-horizon sessionful rollouts, mid-trajectory forking, and cross-calendar-time training. The research questions that follow (long-horizon benchmarks, speculative RL algorithms, event-driven training, to name a few) are where the next wave of agent RL gets built.
🕒 6:15 pm – 6:30 pm
Talk "We Scaled Data Wrong"
🎤 Jun Park, CEO at hillclimb
Outline:
For every significant leap in model intelligence, there was a massive dataset that helped us get there. The first chatbots had common crawl, coding agents had github. The next jump in model intelligence requires hyperspecialized training data (finance, health, law, etc) yet we don't have it. What have we, as an industry, done wrong?
🕒 6:30 pm – 6:45 pm
Talk "Open RL Stack for Training Long-Horizon Agents"
🎤 Sijun Tan, PhD researcher at UC Berkeley’s Sky Computing Lab
Outline:
Training long-horizon language agents requires more than a model and a reward signal: it requires a full stack for orchestrating agent sandboxes, running asynchronous rollouts, computing custom rewards and losses, and scaling across different training backends. In this talk, I will introduce rLLM, an open reinforcement learning stack designed to make this workflow practical and reproducible. rLLM lets users train agents on any harness, execute them in flexible sandboxes, and plug in custom evaluation logic for task-specific reward shaping, custom losses, and multi-step trajectories. The system supports fully asynchronous training and multiple backend options, making it easier to move from experimentation to large-scale training without rewriting agent code. I will also share examples of how rLLM is used in long-horizon agent settings and discuss the design tradeoffs behind building an open RL infrastructure for agentic RL training.
🕒 6:45 pm – 7:00 pm
Talk "FrontierCS: Evaluating LLMs on Open-Ended CS problems"
🎤 Hanchen Li, PhD researcher at UC Berkeley’s Sky Computing Lab
Outline:
Auto-Research has been growing in popularity over the months. We will share latest updates of our project Frontier-CS, one of the pioneers in benchmarking Auto-Research before it was hot! We will share the existing benchmark, the design behind the project, and some ongoing work on improving auto-research data pipeline and evaluation.
🕒 7:00 pm - 8:30 pm
Networking
With food and beverages
About event
An engaging meetup designed for AI researchers to connect, share ideas, and explore the latest advancements in artificial intelligence. The event features informal networking, short talks, and discussions on current research trends, fostering collaboration and knowledge exchange within the AI community.









