Track
Reinforcement Learning
Goal: Learn sequential decision making from MDPs and value methods through deep RL, policy gradients, and model-based planning. RLHF appears only as a short pointer at the end.
Prereqs: Deep Learning. Probability from Math for ML helps.
Status: done
Work through the steps in order. Bold links open YouTube.
Found a broken link or an unclear step? Report a problem with this track.