You've marked 0 of 4 pages in AI Safety & Alignment understood. View your progress →
AI Safety & Alignment — Roadmap
1. Alignment & RLHF
- What "alignment" means, precisely
- RLHF and Constitutional AI as alignment techniques
- Reward hacking
- Specification gaming
- Deception
2. Scalable Oversight & Frontier Safety
- Scalable oversight: evaluating a system that may exceed the evaluator's own capability
- Interpretability's role in safety (see Interpretability)
- Model behavior evaluation for safety specifically
- Frontier model safety practices
Next: Interpretability — the tools that make a model's internal behavior inspectable, which both safety and security work depend on.