Seminar: Mechanistic Interpretability

Neural networks are increasingly integrated into real-world decisions and everyday tasks. Current frontier models are extremely powerful, but also large and complex. We can precisely describe their architecture, control their inputs, and observe their outputs. However, it remains largely unknown what happens as data moves through dozens of layers and billions of weights. This seminar discusses the question Can we look inside a neural network and actually explain, step by step, why it gave a particular output?

For example, an LLM might give the correct answer to math questions, but we may not know whether the process that led to this answer generalizes to other tasks. Similarly, a model might refuse a harmful request, but from the output alone, we cannot tell whether it has genuinely learned to avoid harm or just reacts to certain trigger phrases. To safely apply LLMs, we want to detect which parts in a network are responsible for certain behaviours to intervene and control its behaviour in a desired way.

Mechanistic interpretability (MI) aims to understand what mechanisms such models learn and apply to solve tasks, and to what extent these capabilities generalize. This matters both for safety-critical applications, where predictable behaviour is essential, and for guiding improvements to future models.

This seminar we go over the basic methods and tools to gain insight into a model cognition process and introduce you to a fast-growing subfield of AI, which is increasingly relevant in research and industry.

Some more material for the curious reader:

Course TitleMechanistic Interpretability
Course IDINF-MSc-102
Registrationdrop me an email
ECTS4
Time[tentative] Thursday, 10:15-11:45
Languageenglish
#participantsmax 10
Locationin-person JvF25; seminar room 3rd floor
organized byAmir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger

Requirements

Familiarity with foundations of deep learning, including transformer architectures and in-context learning.

Topics

keywords: circuits, induction heads, probing, activation patching and superposition

DateContent
tbdtbd

How the seminar will look like?

We will meet regularly throughout the semester. In the first few weeks, we will start with introductory lectures on mechanistic interpretability and how to critically review and present research papers. After that, we will have several sessions with presentations, each followed by discussions. In the end we will have one more concluding session.

Other Important information

Grading/Presentations: Grades will be based on your presentation, slides, active participation and a final report/poster. Further details will be discussed in the first session.

Participation/Registration: If more students sign up than there are available topics, we will open a waiting list. Please come to the first lecture even if you are still on the waiting list. If you don’t attend the first session your spot will be freed up.