Seminar: Mechanistic Interpretability
Neural networks are increasingly integrated into real-world decisions and everyday tasks. Current frontier models are extremely powerful, but also large and complex. We can precisely describe their architecture, control their inputs, and observe their outputs. However, it remains largely unknown what happens as data moves through dozens of layers and billions of weights. This seminar discusses the question Can we look inside a neural network and actually explain, step by step, why it gave a particular output?
For example, an LLM might give the correct answer to math questions, but we may not know whether the process that led to this answer generalizes to other tasks. Similarly, a model might refuse a harmful request, but from the output alone, we cannot tell whether it has genuinely learned to avoid harm or just reacts to certain trigger phrases. To safely apply LLMs, we want to detect which parts in a network are responsible for certain behaviours to intervene and control its behaviour in a desired way.
Mechanistic interpretability (MI) aims to understand what mechanisms such models learn and apply to solve tasks, and to what extent these capabilities generalize. This matters both for safety-critical applications, where predictable behaviour is essential, and for guiding improvements to future models.
This seminar we go over the basic methods and tools to gain insight into a model cognition process and introduce you to a fast-growing subfield of AI, which is increasingly relevant in research and industry.
Some more material for the curious reader:
- Tutorial on Mechanistic Interpretability for Language Models from ICML 2025 by Ziyu Yao, Daking Rai
- Open Problems in Mechanistic Interpretability - a recent paper summarizing the state of research
- The Urgency of Interpretability - a blogpost by Dario Amodei
- Transformer Circuits Thread - a blog by Anthropic
- Neuronpedia - an interactive website to explore AI models
| Course Title | Mechanistic Interpretability |
|---|---|
| Course ID | INF-MSc-102 |
| Registration | drop me an email |
| ECTS | 4 |
| Time | [tentative] Thursday, 10:15-11:45 |
| Language | english |
| #participants | max 10 |
| Location | in-person JvF25; seminar room 3rd floor |
| organized by | Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger |
Requirements
Familiarity with foundations of deep learning, including transformer architectures and in-context learning.
Topics
keywords: circuits, induction heads, probing, activation patching and superposition
| Date | Content |
|---|---|
| tbd | tbd |
How the seminar will look like?
We will meet regularly throughout the semester. In the first few weeks, we will start with introductory lectures on mechanistic interpretability and how to critically review and present research papers. After that, we will have several sessions with presentations, each followed by discussions. In the end we will have one more concluding session.
Other Important information
Grading/Presentations: Grades will be based on your presentation, slides, active participation and a final report/poster. Further details will be discussed in the first session.
Participation/Registration: If more students sign up than there are available topics, we will open a waiting list. Please come to the first lecture even if you are still on the waiting list. If you don’t attend the first session your spot will be freed up.