Home

Track

Evals & Safety

Goal: Learn how to frame, attack, and measure model safety, from the alignment problem, specification gaming, and goal misgeneralization through red teaming, jailbreaks, Constitutional AI, and toxicity, truthfulness, and refusal evals.

Prereqs: LLMs. Eval Harnesses helps for harness tooling. Fine-Tuning helps for RLHF and preference training context.

Status: done

Work through the steps in order. Bold links open YouTube.

Step Concept YouTube Read
1 The alignment problem Stuart Russell — AI: What If We Succeed?  
2 Concrete problems in AI safety   Amodei et al. 2016 — Concrete Problems in AI Safety
3 Specification gaming Specification Gaming: How AI Can Turn Your Wishes Against You DeepMind — Specification gaming: the flip side of AI ingenuity
4 Goal misgeneralization Goal Misgeneralization: How a Tiny Change Could End Everything Shah et al. 2022 — Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals
5 Scalable oversight via debate   Irving, Christiano & Amodei 2018 — AI safety via debate
6 Red teaming language models   Ganguli et al. 2022 — Red Teaming Language Models to Reduce Harms
7 Why jailbreaks defeat safety training Jailbroken: How Does LLM Safety Training Fail? — Paper Explained Wei, Haghtalab & Steinhardt 2023 — Jailbroken: How Does LLM Safety Training Fail?
8 Automated adversarial suffixes   Zou et al. 2023 — Universal and Transferable Adversarial Attacks on Aligned Language Models
9 Constitutional AI and RLAIF Constitutional AI: Harmlessness from AI Feedback Bai et al. 2022 — Constitutional AI: Harmlessness from AI Feedback
10 Toxicity evaluation   Gehman et al. 2020 — RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
11 Truthfulness evaluation   Lin, Hilton & Evans 2021 — TruthfulQA: Measuring How Models Mimic Human Falsehoods
12 Standardized red-team and refusal evals   Mazeika et al. 2024 — HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Found a broken link or an unclear step? Report a problem with this track.