Neural Mastery
You've marked 0 of 4 pages in Interpretability understood. View your progress →
0%

Interpretability — Roadmap

1. Classical Interpretability: Feature Importance, SHAP & LIME

  • Built-in feature importance (tree-based models)
  • Permutation importance
  • SHAP (SHapley Additive exPlanations)
  • LIME (Local Interpretable Model-agnostic Explanations)

2. Deep Learning & LLM Interpretability

  • Saliency maps and gradient-based attribution
  • Activation visualization
  • Attention analysis
  • Probing classifiers
  • Sparse autoencoders
  • Circuits
  • Mechanistic interpretability

Next: Reinforcement Learning — a separate track with its own theory (MDPs, policy/value functions) underlying RLHF, agents, and much more.

Last updated Sep 5, 2026Edit this pageReport an issue
← Previous
Interpretability — Overview
Next →
Classical Interpretability: Feature Importance, SHAP & LIME