Course description: Modern statistical applications often involve data sets in which the number of variables is comparable to, or substantially larger than, the number of observations. This course develops the mathematical and statistical foundations needed to analyze such high-dimensional data. Topics include concentration inequalities, random vectors and random matrices, high-dimensional geometry, sparse linear models, regularization methods, the Lasso, model selection, low-rank matrix estimation, principal component analysis, graphical models, minimax lower bounds, and computational considerations.
Particular emphasis is placed on understanding how structural assumptions—such as sparsity, low rank, and graphical structure—make statistically meaningful estimation possible in high dimensions. Students will study nonasymptotic techniques for deriving finite-sample guarantees and will learn to evaluate estimators in terms of consistency, prediction error, estimation error, and variable-selection performance.
Instructor: Yun Wei
Semester: Fall 2026
Textbook: Martin J. Wainwright’s High-Dimensional Statistics: A Non-Asymptotic Viewpoint
(optional) Roman Vershynin's High Dimensional Probability
Lecture notes (it is evolving while I am teaching it this semester) 👉
Schedule
8/24: Subgaussian distributions
8/26: Subgaussian distributions and Sub-exponential distributions