Home

Track

Speech & Audio

Goal: Represent audio for models, then cover ASR (CTC, wav2vec, Whisper) and synthesis (WaveNet, Tacotron 2), plus controllable music generation.

Prereqs: Deep Learning.

Status: done

Work through the steps in order. Bold links open YouTube.

Step Concept YouTube Read
1 Sound and waveforms Valerio Velardo — Sound and Waveforms  
2 Short-time Fourier transform Valerio Velardo — Short-Time Fourier Transform Explained Easily  
3 Mel spectrograms Valerio Velardo — Mel Spectrograms Explained Easily Ketan Doshi — Audio Deep Learning Made Simple: Why Mel Spectrograms perform better
4 Loading and inspecting audio   librosa — Getting started with audio data
5 Mel-frequency cepstral coefficients Valerio Velardo — Mel-Frequency Cepstral Coefficients Explained Easily  
6 Connectionist temporal classification Connectionist Temporal Classification (CTC) Explained Graves et al. 2006 — Connectionist Temporal Classification
7 Self-supervised speech representations Wav2Vec 2.0: Paper Overview Baevski et al. 2020 — wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
8 Large-scale robust ASR OpenAI’s Whisper Model Explained Radford et al. 2022 — Robust Speech Recognition via Large-Scale Weak Supervision
9 Neural waveform synthesis   van den Oord et al. 2016 — WaveNet: A Generative Model for Raw Audio
10 Neural text-to-speech   Shen et al. 2018 — Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
11 Controllable music generation   Copet et al. 2023 — Simple and Controllable Music Generation

Found a broken link or an unclear step? Report a problem with this track.