Home

Track

Reinforcement Learning

Goal: Learn sequential decision making from MDPs and value methods through deep RL, policy gradients, and model-based planning. RLHF appears only as a short pointer at the end.

Prereqs: Deep Learning. Probability from Math for ML helps.

Status: done

Work through the steps in order. Bold links open YouTube.

Step Concept YouTube Read
1 RL problem and the agent loop David Silver — Introduction to Reinforcement Learning Lilian Weng — A (Long) Peek into Reinforcement Learning
2 Markov decision processes David Silver — Markov Decision Process Sutton & Barto — Finite Markov Decision Processes (ch. 3)
3 Temporal-difference learning and Q-learning David Silver — Model Free Control OpenAI Spinning Up — Key Concepts in RL
4 Deep Q-Networks Yannic Kilcher — Playing Atari with Deep Reinforcement Learning Mnih et al. 2013 — Playing Atari with Deep Reinforcement Learning
5 Policy gradients David Silver — Policy Gradient Methods Lilian Weng — Policy Gradient Algorithms
6 Asynchronous advantage actor-critic   Mnih et al. 2016 — Asynchronous Methods for Deep Reinforcement Learning
7 Proximal policy optimization Costa Huang — PPO Implementation Details Schulman et al. 2017 — Proximal Policy Optimization Algorithms
8 Soft actor-critic Phil Tabor — Soft Actor Critic in PyTorch Haarnoja et al. 2018 — Soft Actor-Critic
9 Twin delayed DDPG   Fujimoto et al. 2018 — Addressing Function Approximation Error in Actor-Critic Methods
10 MuZero model-based planning Yannic Kilcher — MuZero Schrittwieser et al. 2019 — Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
11 RLHF pointer   Ouyang et al. 2022 — Training language models to follow instructions with human feedback

Found a broken link or an unclear step? Report a problem with this track.