You've marked 0 of 4 pages in Interpretability understood. View your progress →
Interpretability — Roadmap
1. Classical Interpretability: Feature Importance, SHAP & LIME
- Built-in feature importance (tree-based models)
- Permutation importance
- SHAP (SHapley Additive exPlanations)
- LIME (Local Interpretable Model-agnostic Explanations)
2. Deep Learning & LLM Interpretability
- Saliency maps and gradient-based attribution
- Activation visualization
- Attention analysis
- Probing classifiers
- Sparse autoencoders
- Circuits
- Mechanistic interpretability
Next: Reinforcement Learning — a separate track with its own theory (MDPs, policy/value functions) underlying RLHF, agents, and much more.