Seminar Explainable Machine Learning
(Winter term 2026/27)
Who, when, where
Who: Ulrike von Luxburg together with Maximilian Thiessen and Gunnar König When: Wednesdays 14:15 - 15:45 and two compact days, tentatively 21. and 22.01.27. Where: Maria-von-Linden-Strasse, 1, seminar room A-222 (ground floor) Language: English Credit points: 3 CPDescription
Modern machine learning methods, such as deep learning, random forests, or XGBoost typically produce ``black box models'': they can excel at prediction, but it is completely unclear on which criteria these predictions are being based. While there are many applications where this might not be an issue, there are others where a deeper understanding of the models is necessary: applications in medicine, applications in society (say, credit scoring), or applications in science, where we want to understand the underlying processes. Also, from the legal point of view, explanations are arguably requested in the new AI Act that regulates machine learning applications, and explanations are often considered as a means to establish trust in machine learning applications. The field of explainable machine learning, often abbreviated as XAI, tries to develop methods and algorithms that supposedly produce explanations for machine learning models. In this seminar, we are going to discuss many of the standard approaches in this field. We will also discuss critically whether and under which conditions the suggested methods might achieve their goal or not.Prerequesits
This seminar is intended for master students in machine learning, computer science or related fields. Basic knowledge on machine learning, for example at least one of the standard classes in the ML master program, is required.Registration and time line
- Oct 14: first seminar meeting, Intro and orga
- Oct 21: First lecture and paper bidding deadline: If you want to participate, you need to bid for papers and register on Ilias by Oct 21 (links will follow after the first seminar on Oct 14 has taken place).
- Oct 28: Second lecture; assignment to papers
- Dec 1rst: you need to have contacted your supervisor and agreed on a meeting time before christmas.
- Before christmas, you need to upload a first version of your slides and then meet your supervisor to get feedback.
- Final presentations: Jan 21 + 22, all day.
Organization
- Phase 1 (Mid Oct - Mid Nov): During the first 2-3 weeks, the seminar proceeds as a lecture: Ulrike Luxburg and two of her postdocs give an overview on the basic methods and questions in the field.
- Phase 2 (Mid Nov - End Dec): The seminar participants work on their individual presentations.
- Phase 3 (Jan): Presentations take place on two whole days (dates tba). Presentations need to be in english.
To pass the seminar, each participant has to give a presentation, needs to act as sparring partner for another paper, and needs to be present at the compact seminar days. Details will be explained during the seminar.
Schedule
Tentative and subject to change:- 12.10. Lecture 1 (Ulrike von Luxburg): Introduction to XAI, setup of this seminar.
- 19.10. Lecture 2 (Gunnar König): Conflicting XAI goals, contrastive explanations, global methods.
- 26.10. Lecture 3 (Maximilian Thiessen): Overview on explanation methods.
- 21.01. All day: presentations by seminar participants.
- 22.02. All day: presentations by seminar participants.
List of all papers
- Paper 1: T. Han, S. Srinivas, and H. Lakkaraju. "Which explanation should I choose? A function approximation perspective to characterizing post hoc explanations." Neural Information Processing Systems (NeurIPS), 2022.
- Paper 2: S. Krishna, T. Han, A. Gu, S. Wu, S. Jabbari, and H. Lakkaraju. "The disagreement problem in explainable machine learning: A practitioner's perspective." Transactions on Machine Learning Research (TMLR), 2024.
- Paper 3: M. Taimeskhanov, and D. Garreau. "Feature attribution from first principles." arXiv preprint arXiv:2505.24729 (2025).
- Paper 4: D. Slack, S. Hilgard, E. Jia, S. Singh, and H. Lakkaraju. "Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods." Conference on AI, Ethics, and Society (AIES), 2020.
- Paper 5: H. Hwang, A. Bell, J. Fonseca, V. Pliatsika, J. Stoyanovich, and S. E. Whang. "SHAP-based explanations are sensitive to feature representation." Conference on Fairness, Accountability, and Transparency (FAccT), 2025.
- Paper 6: J. Skirzynski, D. Danks, and B. Ustun. "Discrimination exposed? On the reliability of explanations for discrimination detection." Conference on Fairness, Accountability, and Transparency (FAccT), 2025.
- Paper 7: S. Tekkesinoglu. "When explanations deceive: Understanding unintentional and intentional deception in XAI." Conference on Fairness, Accountability, and Transparency (FAccT), 2026.
- Paper 8: S. Jain, and B. C. Wallace. "Attention is not explanation." Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2019.
- Paper 9: S. Serrano, and N. A. Smith. "Is attention interpretable?" Annual Meeting of the Association for Computational Linguistics (ACL), 2019.
- Paper 10: S. Wiegreffe, and Y. Pinter. "Attention is not not explanation." Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019.
- Paper 11: B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, and R. Sayres. "Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)." International Conference on Machine Learning (ICML), 2018.
- Paper 12: M. Böhle, M. Fritz, and B. Schiele. "B-cos networks: Alignment is all we need for interpretability." Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
- Paper 13: H. Fokkema, T. van Erven, and S. Magliacane. "Sample-efficient learning of concepts with theoretical guarantees: from data to concepts without interventions." Neural Information Processing Systems (NeurIPS), 2025.
- Paper 14: Y. Elisha, O. Barkan, Z. Haddad, and N. Koenigstein. "ConEx: Human-interpretable saliency maps via concept-aware attribution." International Conference on Machine Learning (ICML), 2026.
- Paper 15: P. Atanasova, O. Camburu, C. Lioma, T. Lukasiewicz, J. G. Simonsen, and I. Augenstein. "Faithfulness tests for natural language explanations." Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
- Paper 16: M. Turpin, J. Michael, E. Perez, and S. R. Bowman. "Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting." Neural Information Processing Systems (NeurIPS), 2023.
- Paper 17: Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, V. Mikulik, S. R. Bowman, J. Leike, J. Kaplan, and E. Perez. "Reasoning models don't always say what they think." arXiv preprint arXiv:2505.05410 (2025).
- Paper 18: A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, A. Tamkin, E. Durmus, T. Hume, F. Mosconi, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan. "Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet." Transformer Circuits Thread (2024), arXiv preprint arXiv:2605.29358.
- Paper 19: E. Ameisen, J. Lindsey, A. Pearce, W. Gurnee, N. L. Turner, B. Chen, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. Marcus, M. Sklar, A. Templeton, T. Bricken, C. McDougall, H. Cunningham, T. Henighan, A. Jermyn, A. Jones, A. Persic, Z. Qi, T. B. Thompson, S. Zimmerman, K. Rivoire, T. Conerly, C. Olah, and J. Batson. "Circuit tracing: Revealing computational graphs in language models." Transformer Circuits Thread, 2025.
- Paper 20: D. Braun, L. Bushnaq, S. Heimersheim, J. Mendel, and L. Sharkey. "Interpretability in parameter space: Minimizing mechanistic description length with attribution-based parameter decomposition." arXiv preprint arXiv:2501.14926 (2025).
- Paper 21: D. Chanin, J. Wilken-Smith, T. Dulka, H. Bhatnagar, S. Golechha, and J. Bloom. "A is for absorption: Studying feature splitting and absorption in sparse autoencoders." Neural Information Processing Systems (NeurIPS), 2025.
- Paper 22: T. Heap, T. Lawson, L. Farnik, and L. Aitchison. "Automated interpretability metrics do not distinguish trained and random transformers." International Conference on Learning Representations (ICLR), 2026.
- Paper 23: C. Panigutti, R. Hamon, I. Hupont, D. Fernandez Llorca, D. Fano Yela, H. Junklewitz, S. Scalzo, G. Mazzini, I. Sanchez, J. Soler Garrido, and E. Gomez. "The role of explainable AI in the context of the AI Act." Conference on Fairness, Accountability, and Transparency (FAccT), 2023.
- Paper 24: T. Schmude, L. Koesten, T. Möller, and S. Tschiatschek. "On the impact of explanations on understanding of algorithmic decision-making." Conference on Fairness, Accountability, and Transparency (FAccT), 2023.
- Paper 25: C. Yadav, M. Moshkovitz, and K. Chaudhuri. "XAudit: A learning-theoretic look at auditing with explanations." Transactions on Machine Learning Research (TMLR), 2024.
- Paper 26: M. Stewart. "Beyond explanation: Evidentiary rights for algorithmic accountability." Conference on Fairness, Accountability, and Transparency (FAccT), 2026.