
How to take apart a neural network (or, Goodfire's AdVersarial Parameter Decomposition Technique)
Join us for an insightful talk on Goodfire's AdVersarial Parameter Decomposition Technique!
Most interpretability work focuses on decomposing a model's activations, but Goodfire's innovative approach breaks down a network's weights into components that serve distinct computational roles, such as predicting emoticons and identifying gender.
This method transforms the model from a simple lookup table into a generalizing algorithm, effectively handling attention computations across multiple heads. It stands as a competitor to SAEs and transcoders, allowing researchers to manually edit model behavior for interpretable and predictable changes.
Speaker: Linda Linsefors, co-author of Goodfire's Interpreting Language Model Parameters paper, will guide us through the workings of AdVersarial Parameter Decomposition (VPD), its findings, and future directions.
π Research here
Schedule:
- π 7:00 PM β Doors open, snacks and refreshments served
- π’ 7:30 PM β Talk begins!
- π£ 8:30 PM β Wrap up Q&A, continued hangouts
This event is part of the Mox Summer Season: talks and socials at the frontier of ideas, running June through August.
Mox, 1680 Mission St, San Francisco, CA 94103, USA
Get directionsScan with your camera β the event opens in the Somo app.









