Instructor:
Shulei Wang (shuleiw at illinois dot edu)
TA:
Arghya Chakraborty (arghyac2 at illinois dot edu)
Course Website: Canvas
Office Hours:
TBD by Arghya Chakraborty
Tuesday 9-10:00am (CST) by Shulei Wang
This course provides a rigorous introduction to self-supervised representation learning (SSL), focusing on bridging the gap between empirical deep learning and formal statistical theory. We will explore how models learn representations from unlabeled data by solving pretext tasks. The curriculum is structured around two central themes: first, a comprehensive survey of the algorithmic landscape, examining state-of-the-art frameworks such as Contrastive Methods, Masked Modeling, Video Representation Learning, and Multi-Modal Alignment; and second, an exploration of statistical foundations and recent progress. In the latter half, we will discuss the evolving theoretical underpinnings of these methods, focusing on the role of self-supervised signals, the geometry of latent spaces, and the statistical properties that enable these representations to transfer effectively to downstream tasks. Throughout the course, students are expected to connect algorithms to statistical questions: what self-supervised signal is used, what representation is identifiable, what information is preserved or discarded, how collapse is avoided, and why the representation transfers to downstream tasks.
No textbook is required.
Foundation of Representation Learning
Contrastive Learning
Non-contrastive Siamese Methods
Geometry and Theory of Contrastive Representations
Self-Supervised Language Modeling
Masked Image Modeling and Vision Foundation Model
Geometry and Theory of Masked Prediction
Predictive Latent-Space Modeling
Speech and Video Self-Supervised Learning
Multimodal Alignment
Three Mini-Projects (60%)
Final Project (40%)