Home

Track

Interpretability

Goal: Learn how to explain model predictions and inspect internals, from permutation importance, PDP/ICE, LIME, and SHAP through gradient attribution, Grad-CAM, concept tests, faithfulness checks, and circuits.

Prereqs: ML Basics. Deep Learning helps for attribution, Grad-CAM, and circuits.

Status: done

Work through the steps in order. Bold links open YouTube.

Step Concept YouTube Read
1 What interpretability is for What is Interpretable Machine Learning Molnar — Interpretability
2 Permutation feature importance Permutation Feature Importance Molnar — Permutation Feature Importance
3 Partial dependence and ICE Partial Dependence Plots Molnar — Partial Dependence Plot
4 LIME Ribeiro — Why Should I Trust You? Ribeiro et al. 2016 — “Why Should I Trust You?” Explaining the Predictions of Any Classifier
5 Shapley values and SHAP The Fair Way to Attribute Feature Importance: Shapley Values Lundberg & Lee 2017 — A Unified Approach to Interpreting Model Predictions
6 Integrated Gradients   Sundararajan et al. 2017 — Axiomatic Attribution for Deep Networks
7 Grad-CAM Grad-CAM Selvaraju et al. 2017 — Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
8 Feature visualization   Distill — Feature Visualization
9 Concept activation vectors (TCAV) Interpretability Beyond Feature Attribution Kim et al. 2018 — Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
10 Sanity checks for saliency maps   Adebayo et al. 2018 — Sanity Checks for Saliency Maps
11 Attention is not explanation   Jain & Wallace 2019 — Attention is not Explanation
12 Counterfactual explanations   Wachter et al. 2017 — Counterfactual Explanations without Opening the Black Box
13 Circuits A Walkthrough of A Mathematical Framework for Transformer Circuits Distill — Zoom In: An Introduction to Circuits

Found a broken link or an unclear step? Report a problem with this track.