
Beyond the Black Box: A Technical Introduction to Mechanistic Interpretability
Join us for a working, technical introduction to mechanistic interpretability — the neuroscience of AI! 🧠✨
We've developed powerful AI systems for high-stakes applications, yet understanding their inner workings remains a challenge. This session will explore:
- Why mechanistic interpretability matters
- The core toolkit, including sparse autoencoders and probes, to uncover the model's tangled internal state.
- Practical applications: finding and moving features, transforming internal knowledge into signals, and identifying hallucinations from within.
- Techniques for reading a model's internal state to preemptively flag risky tool calls. ⚠️
The field is akin to the early days of neuroscience, with open tools and unclaimed problems. This session aims to inspire as much as it instructs.
About the Speaker:
Hariom brings years of experience in AI, machine learning, and finance. He is an O'Reilly author and published researcher, focusing on enhancing the transparency and reliability of large language models in financial and agentic AI contexts. Hariom has spoken at numerous conferences and received the Indian Achiever Award in Machine Learning. He holds an MS from UC Berkeley and a BE from IIT (India). 🎓
To Attend Online:
Looking forward to seeing you! 👋
Scannez avec l'appareil photo – l'événement s'ouvre dans l'app Somo.

![[Women in Claude] First IRL Singapore Meetup!](/_next/image?url=https%3A%2F%2Fimages.gosomo.app%2Fevents%2Fc5c8390d-c547-4089-b0ca-74891502e77b%2Ff5905994-d93a-41ad-b97a-cff927a73a77.webp&w=1200&q=75)







