Home

Track

LLMs

Goal: Large language models from the transformer stack through pretraining, scaling, and instruction alignment.

Prereqs: Deep Learning. NLP through transformers helps.

Status: done

Work through the steps in order. Bold links open YouTube.

Step Concept YouTube Read
1 Transformer overview 3Blue1Brown — Transformers, the tech behind LLMs SLP3 ch. 7 Transformers and Pretraining
2 Attention 3Blue1Brown — Attention in transformers, step-by-step The Illustrated Transformer
3 The Transformer paper StatQuest — Transformer Neural Networks Vaswani et al. 2017 — Attention Is All You Need
4 Decoder-only transformers StatQuest — Decoder-Only Transformers  
5 Tokens and BPE Karpathy — Let’s build the GPT Tokenizer  
6 Build a GPT Karpathy — Let’s build GPT  
7 GPT-2 Karpathy — Let’s reproduce GPT-2 (124M) Radford et al. 2019 — Language Models are Unsupervised Multitask Learners
8 GPT-3 and in-context learning Yannic Kilcher — GPT-3: Language Models are Few-Shot Learners Brown et al. 2020 — Language Models are Few-Shot Learners
9 Scaling laws   Hoffmann et al. 2022 — Training Compute-Optimal Large Language Models
10 Instruction following and RLHF AssemblyAI — How ChatGPT actually works Ouyang et al. 2022 — InstructGPT
11 Post-training overview   SLP3 ch. 8 Post-training
12 Open foundation models   Touvron et al. 2023 — LLaMA

Found a broken link or an unclear step? Report a problem with this track.