
VCN #45: Bench
Get ready for an immersive hands-on build night where bringing your laptop is a must! 💻✨
What's in store?
- Ditch the vibe and get ready to measure your coding agent with real numbers! No more guessing—develop a true evaluation and start tracking performance.
- Format:
- Walkthrough: Learn what a coding-agent evaluation really is and discover why public benchmarks often fail when applied to your specific repository.
- Create Your Own Bench: Pull real tasks from your codebase—be it a bug with a confirmed fix, a straightforward refactor, or a feature with a passing test. Five tasks you trust will always outweigh a thousand you can't count on! 🔍
- Build the Harness: Wire up a deterministic oracle for each task, get set for testing with no judgment from LLMs. Run your agent and score it—repeat to measure both pass@k and the vital consecutive-green reliability!
- Leaderboard Analysis: Discover where your agent excels and where it might be flaky—all through real numbers derived from your tasks. 📊
By 10 PM, you’ll have crafted a repeatable evaluation for any coding agent tailored to your tasks. Important: We will utilize this very bench at the next Bake-Off (#46) for head-to-head agent scoring!
Who should come? Builders only! Bring a repository that includes at least one reliable test.
Details:
- Location: Frontier Tower 🧑🚀
- Doors open at: 7 PM
- Walkthrough starts at: 7:30 PM
Hosted by:
- Rayyan Zahid (Immersive Commons)
- Michalis Vasileiadis (Hacker Bob)
- Eric Mockler (AI Geneticist)
- Devinder Sodhi (Learning Layer Labs)
Facilitator: Rayyan Zahid
Guest Speaker: TBD (open call)
Your ticket includes access to z.ai + Claude Code for the session and Nebius Token Factory credits for lab activities.
RSVP if you’ve shipped a coding agent and need to quantify its performance. 📩
Special for Frontier Tower members: Your ticket is complimentary! DM Ray directly for a members-only free RSVP.
Frontier Tower 🧑🚀
Trasa








