You've marked 0 of 5 pages in Reinforcement Learning understood. View your progress →
Reinforcement Learning — Roadmap
1. RL Fundamentals
- Markov Decision Processes: states, actions, rewards, transitions
- Policy and value functions
- The Bellman equations
- Q-learning
- SARSA
- On-policy vs off-policy learning
2. Advanced RL
- DQN, policy gradients, actor-critic, PPO (see Advanced Architectures)
- SAC (Soft Actor-Critic) and TD3 for continuous control
- Offline RL
- Imitation learning
- Inverse RL
- RLHF / RLAIF / GRPO in the RL framework (see Training Pipeline)
3. Multi-Armed & Contextual Bandits
- Regret, and why it's the right way to score an explore/exploit strategy
- ε-greedy, UCB1, and Thompson Sampling — real regret tradeoffs between them
- Contextual bandits and LinUCB
- When a bandit is the right model vs. when you need full RL
Next: Graph ML — another specialized track, for data with explicit relational structure.