
ModelOps Night: Evals, Routing & Benchmarks
✨ Join us for a technical deep dive on evaluating AI products, agents, and workflows against real-world tasks! 🤖
As AI models get better, the crucial question shifts from "which model is best?" to whether your product performs effectively on tasks that matter to users. Public benchmarks often miss edge cases, while real production systems contend with messy inputs, multi-step workflows, and various constraints.
This event unites teams constructing the infrastructure for evaluating AI systems in practical scenarios.
Program:
- 5:30 PM: Doors open · 🍕 Food & drinks
- 6:00 PM: Composio: Building evals for real-world workflows (Jayesh Sharma, AI Engineer)
- 6:15 PM: GMI Cloud: Routing models to optimize performance and spend
- 6:30 PM: Handshake: Building benchmarks for economically-valuable agent work (Jonas Mueller, Director of AI Research)
- 6:45 PM: Oqoqo: Evaluating and improving agent experience by building your own benchmark (Haritha Nair, CTO)
- 7:00 PM: Networking 🤝
Who Should Attend:
Founders, product engineers, AI engineers, and teams developing agents or products that must be tested against real user tasks.
📍 Location: Werqwise, 465 California St 7th floor, San Francisco, CA 94104, USA
Werqwise, 465 California St 7th floor, San Francisco, CA 94104, USA
ItinéraireScannez avec l'appareil photo – l'événement s'ouvre dans l'app Somo.









