Track
Evals & Safety
Goal: Learn how to frame, attack, and measure model safety, from the alignment problem, specification gaming, and goal misgeneralization through red teaming, jailbreaks, Constitutional AI, and toxicity, truthfulness, and refusal evals.
Prereqs: LLMs. Eval Harnesses helps for harness tooling. Fine-Tuning helps for RLHF and preference training context.
Status: done
Work through the steps in order. Bold links open YouTube.
Found a broken link or an unclear step? Report a problem with this track.