Neural Mastery
You've marked 0 of 4 pages in AI Safety & Alignment understood. View your progress →
0%

AI Safety & Alignment — Roadmap

1. Alignment & RLHF

  • What "alignment" means, precisely
  • RLHF and Constitutional AI as alignment techniques
  • Reward hacking
  • Specification gaming
  • Deception

2. Scalable Oversight & Frontier Safety

  • Scalable oversight: evaluating a system that may exceed the evaluator's own capability
  • Interpretability's role in safety (see Interpretability)
  • Model behavior evaluation for safety specifically
  • Frontier model safety practices

Next: Interpretability — the tools that make a model's internal behavior inspectable, which both safety and security work depend on.

Last updated Sep 5, 2026Edit this pageReport an issue
← Previous
AI Safety & Alignment — Overview
Next →
Alignment & RLHF