Track
Speech & Audio
Goal: Represent audio for models, then cover ASR (CTC, wav2vec, Whisper) and synthesis (WaveNet, Tacotron 2), plus controllable music generation.
Prereqs: Deep Learning.
Status: done
Work through the steps in order. Bold links open YouTube.
Found a broken link or an unclear step? Report a problem with this track.