Academic employment: Associate Professor of Economics (with tenure), UCSD, since July 2025; Assistant Professor of Economics, UCSD, 2019-2025; Assistant Professor of Statistics and Computer Science, Purdue University, West Lafayette, 2018-2019
Past research areas: Mathematical Foundations of Machine Learning, Nonasymptotic Statistics, High-Dimensional Estimation and Inference, Semiparametric and Nonparametric Methods, Algorithmic Economics
To prospective PhD students: I do not operate a traditional, hierarchical lab. Instead, I view research as a direct partnership. When I work with a student, we act as true coauthors—co-developing an idea we are both genuinely excited about and tackling the execution side-by-side. If you are looking for a hands-on, highly collaborative environment where we build the project together from the ground up, our working styles will align well.
Below, ** indicates papers that were either single authored or co-authored with a (former or current) student.
New Papers
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs** by Xingwei Gan, Ying Zhu (alphabetical ordering)
(We are currently exploring industry collaborations to test and scale the following findings to larger foundation models.)
Core Innovation: Replaces traditional Kullback-Leibler (KL) regularization and critic architectures in Group Relative Policy Optimization (GRPO) with a novel logit-averaging mechanism.
Post-Training Compute Savings: Eliminates the computational overhead associated with computing gradients for the explicit KL penalty term during the post-training phase, and maintains comparable memory requirements relative to GRPO.
Complementary Skill Composition: Couples a trainable reasoning policy with a frozen Supervised Fine-Tuning (SFT) anchor directly in the logit space, allowing the model to retain SFT formatting rigor while leveraging unconstrained mathematical reasoning.
Pathway to 1x Inference Deployment: Leveraging the mixed policy as a teacher to train a single student network collapses the capabilities back into a standalone model to restore standard inference efficiency.
Approximating invariant functions with the sorting trick is theoretically justified** by Wee Chaimanowong, Ying Zhu (alphabetical ordering), R&R at SIAM Journal on Mathematics of Data Science (SIMODS)
The Gap in AI Practice: In modern AI, models often need to process unordered data—such as 3D point clouds in autonomous driving (e.g., PointNet), molecular structures in Graph Neural Networks, or set-based Transformers. To handle this, developers widely use a practical "sorting trick" (a form of canonicalization) because it is computationally much faster and more efficient than calculating all possible data arrangements (group averaging).
The Mathematical Skepticism: Despite its widespread practical use, the mathematical foundation of this sorting trick has been heavily doubted because it breaks traditional rules; sorting data makes the underlying mathematical functions "rough" or jumpy (non-differentiable and discontinuous).
The Certification: This paper explores the approximation theory behind this practice and provides the formal mathematical certification that justifies using the sorting trick in machine learning and AI models.
Error Reduction: This paper proves that sorting reorganizes the data points efficiently, which fundamentally reduces the model's overall approximation error by a significant factor.
Publications and Accepted Papers
Permutation Invariant Functions: Statistical Testing, Density Estimation and Metric Entropy** by Wee Chaimanowong, Ying Zhu (alphabetical ordering) - AISTATS 2025
On the limitations of data-based price discrimination** by Haitian Xie, Ying Zhu, Denis Shishkin (Xie and Zhu share the first authorship and are listed alphabetically).
Theoretical Economics, 2025
An extended abstract in ACM Economics and Computation 2024 (full paper peer refereed)
Phase transitions in nonparametric regressions** by Ying Zhu - Accepted at Journal of Econometrics, 2023
Omitted Variable Bias of Lasso-Based Inference Methods: A Finite Sample Analysis by Kaspar Wuthrich, Ying Zhu (alphabetical ordering) - The Review of Economics and Statistics, 2023
Inference in Approximately Sparse Correlated Random Effects Probit Models with Panel Data by Jeffrey Wooldridge, Ying Zhu (alphabetical ordering) - The Journal of Business & Economic Statistics, 2020
High Dimensional Inference in Partially Linear Models by Ying Zhu, Zhuqing Yu, Guang Cheng, AISTATS 2019
Sparse Linear Models and l1-Regularized 2SLS with High-Dimensional Endogenous Regressors and Instruments** {Additional materials from 2013} by Ying Zhu - The Journal of Econometrics, 2018
Nonasymptotic Analysis of Semiparametric Regression Models with High-Dimensional Parametric Coefficients** by Ying Zhu - The Annals of Statistics, 2017
Nonparametric Density Estimation Based on Truncated Mean** by Ying Zhu - Statistics and Probability Letters, 2013
Unpublished Manuscripts