How to take apart a neural network (or, Goodfire's AdVersarial Parameter Decomposition Technique)

How to take apart a neural network (or, Goodfire's AdVersarial Parameter Decomposition Technique)

About the event

Join us for an insightful talk on Goodfire's AdVersarial Parameter Decomposition Technique!

Most interpretability work focuses on decomposing a model's activations, but Goodfire's innovative approach breaks down a network's weights into components that serve distinct computational roles, such as predicting emoticons and identifying gender.

This method transforms the model from a simple lookup table into a generalizing algorithm, effectively handling attention computations across multiple heads. It stands as a competitor to SAEs and transcoders, allowing researchers to manually edit model behavior for interpretable and predictable changes.

Speaker: Linda Linsefors, co-author of Goodfire's Interpreting Language Model Parameters paper, will guide us through the workings of AdVersarial Parameter Decomposition (VPD), its findings, and future directions.

πŸ”— Research here

Schedule:

  • πŸ•– 7:00 PM β€” Doors open, snacks and refreshments served
  • πŸ•’ 7:30 PM β€” Talk begins!
  • πŸ•£ 8:30 PM β€” Wrap up Q&A, continued hangouts

This event is part of the Mox Summer Season: talks and socials at the frontier of ideas, running June through August.

Found by SomoΒ·See original
Location

Mox, 1680 Mission St, San Francisco, CA 94103, USA

Get directions

This week in San Francisco