Seminar: Mechanistic Interpretability
Seminar, TU Dortmund, 2026
Neural networks are increasingly integrated into real-world decisions and everyday tasks. Current frontier models are extremely powerful, but also large and complex. We can precisely describe their architecture, control their inputs, and observe their outputs. However, it remains largely unknown what happens as data moves through dozens of layers and billions of weights. This seminar discusses the question Can we look inside a neural network and actually explain, step by step, why it gave a particular output?