Track
Multimodal
Goal: Connect vision and language with contrastive encoders, then generative VLMs (Flamingo, BLIP-2, LLaVA), and finally joint spaces across many modalities.
Prereqs: Deep Learning. Computer Vision and LLMs help.
Status: done
Work through the steps in order. Bold links open YouTube.
Found a broken link or an unclear step? Report a problem with this track.