A public ELLIS reading group exploring the interplay between the mathematical foundations of deep learning and the practical challenge of making ML efficient — from optimization theory to hardware-aware training. Learn more about our topics and scope →
Not everything we find interesting makes it into a session. For the rest papers, talks, and ideas worth sharing see Writeups →
12. October 2026 @ 5pm CEST / 11am EST / 8am PST [timezone converter]
Cracking the Hessian: Closed-Form Hessian Spectra for some Fundamental Neural Networks
Sidak Pal Singh, Google DeepMind, USA
Abstract: The Hessian and its spectrum hold significant theoretical and practical relevance for building op- timizers, measuring generalization, compressing models, and more. Prior works have characterized the Hessian through its spectral density, rank, and the outlier–bulk structure of its spectrum, often relying on approximations. However, the precise behavior of Hessian eigenvalues and eigenvectors remains unclear, owing both to the absence of closed-form results for non-trivial neural networks and the computational expense of empirical estimation. As a first in the literature, in this work, we derive closed-form expressions for all Hessian eigenvalues and eigenvectors in two-layer linear and ReLU networks with scalar input, arbitrary hidden width, and where the loss is aggregated over any number of samples. We further provide closed-form eigenvalues for the core component of Transformer architectures — a single self-attention layer with arbitrary sequence length. Our results reveal a previously undiscovered "paired" structure of outlier eigenvalues, a cell-wise decomposition of the Hessian spectrum with ReLU, and the sensitivity of the Hessian condition number to the query and key matrix norms, as well as the presence of attention sinks. We complement these findings with experiments beyond the assumed model setting, showing strong correlation between the largest eigenvalue and the spectral norm of weight matrices, and empirical evidence that the paired eigenvalue structure persists more generally. Overall, this advances our understanding of the Hessian via neatly exhibiting its exact and complete spectrum for a core set of networks.
OpenReview: https://openreview.net/pdf?id=gW30Rx4eZP
26. October 2026 @ 5pm CET / 12pm EST / 9am PST [timezone converter]
Exploring and Exploiting Stability in Latent Flow Matching
Rania Briq, Jülich Supercomputing Centre and Technical University Dortmund, Germany
Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capacity shrinkage. We characterize this stability by these models' tendency to generate similar outputs under identical noise seeds. We provide a perspective relating this phenomenon to flow matching theory, which indicates that this stability is inherent to the FM objective. We further exploit this stability to derive practical algorithms for more efficient training and inference. Concretely, first, we show that by training LFM models on significantly reduced datasets, performance is preserved, and in compute-constrained regimes, the model converges faster while maintaining quality. This yields multiple advantages, including savings in the training time due to faster convergence, and alleviating annotation effort when training conditional models. Second, LFM stability under architectural shrinkage gives rise to a two-model coarse-to-fine approach, one using a light-weight architecture for the first phase of the FM trajectory, and one with higher capacity for the second, thereby reducing the inference cost substantially. To determine which samples are informative, we introduce three sample-scoring criteria and evaluate them under standard metrics for generative models. Our results are thoroughly evaluated on multiple datasets, demonstrating the practical advantage of this stability, including data savings and a more than two-fold inference speedup while generating comparable outputs.
28. September 2026 @ 5pm CEST — ▶️ YouTube
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
Ali Hojjat, Kiel University and Hamburg University of Technology (TUHH), Germany
arXiv: https://arxiv.org/abs/2507.10800
Web: https://ds-kiel.github.io/ThinkingViT-project-page/
14. September 2026 @ 5pm CEST — ▶️ YouTube
Temperature Paths Reveal Generalization in the Interpolation Regime
Erfan Mirzaei, University of Genova, Italy and Institut Polytechnique de Paris, France
arXiv: https://arxiv.org/abs/2510.06028
20. July 2026 @ 5pm CEST — ▶️ YouTube
Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction
Yumao Liu, The Hong Kong University of Science and Technology, China
arXiv: https://arxiv.org/abs/2605.10388
13. July 2026 @ 5pm CEST — ▶️ YouTube
Geometry-aware similarity metrics for neural representations on Riemannian and statistical manifolds
N Alex Cayco Gajic, École Normale Supérieure Paris, France
Arthur Pellegrino, University College London, UK and Ecole Normale Supérieure Paris, France
arXiv: https://arxiv.org/pdf/2603.28764
29. June 2026 @ 5pm CEST — ▶️ YouTube
Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
Shikang Zheng, Shanghai Jiao Tong University and South China University of Technology, China
arXiv: https://arxiv.org/abs/2510.04188
22. June 2026 @ 5pm CEST — ▶️ YouTube
WK, WV is (Linearly) All You Need: On the Necessity of the QKV Weight Triplet in Self-Attention Transformers
Marko Karbevski, In Simplicity Technologies, Skopje, Macedonia
Antonij Mijoski, Institut de Recherche Mathématique Avancée (IRMA), Université de Strasbourg, France
arXiv: https://arxiv.org/pdf/2510.23912
15. June 2026 @ 5pm CEST — ▶️ YouTube
Generalization at the Edge of Stability
Mario Tuci, INRIA, CNRS, PSL, France and Imperial College London, UK
arXiv: https://arxiv.org/abs/2604.19740
8. June 2026 @ 5pm CEST — ▶️ YouTube
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
Mariia Seleznova, Ludwig Maximilian University of Munich, Germany
arXiv: https://arxiv.org/pdf/2505.19827
1. June 2026 @ 5pm CEST — ▶️ YouTube
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
Eugenia Iofinova, Institute of Science and Technology Austria
Andrej Jovanovic, University of Cambridge, UK
arXiv: https://arxiv.org/abs/2407.10994
11. May 2026 @ 5pm CEST — ▶️ YouTube
Finite-Time Lyapunov Exponents of Deep Neural Networks
Bernhard Mehlig, Department of Physics, University of Gothenburg, Sweden
DOI: 10.1103/PhysRevLett.132.057301
27. April 2026 @ 5pm CEST — ▶️ YouTube
It's not a Lottery, it's a Race: Understanding How Gradient Descent Adapts the Network's Capacity to the Task
Hannah Pinson, Eindhoven University of Technology, Netherlands
arXiv: https://arxiv.org/abs/2602.04832
13. April 2026 @ 5pm CEST — ▶️ YouTube
Sustainable Development and Energy Efficiency in Deep Learning
Raphael Fischer, TU Dortmund and Lamarr Institute, Germany
arXiv: https://arxiv.org/abs/2509.22092
30. March 2026 @ 5pm CEST — ▶️ YouTube
s1: Simple test-time scaling
Niklas Muennighoff, Stanford University, Allen Institute for AI, Contextual AI, USA
arXiv: https://arxiv.org/abs/2501.19393
16. March 2026 @ 5pm CET — ▶️ YouTube
Procedural Pretraining: Warming Up Language Models with Abstract Data
Liangze Jiang, EPFL and Idiap Research Institute, Switzerland
Zachary Shinnick, Australian Institute for Machine Learning (AIML), Adelaide University, Australia
arXiv: https://arxiv.org/pdf/2601.21725
9. March 2026 @ 5pm CET — ▶️ YouTube
How Does Sharpness-Aware Minimization Minimize Sharpness?
Kaiyue Wen, Stanford University, USA
arXiv: https://arxiv.org/abs/2211.05729
2. March 2026 @ 5pm CET — ▶️ YouTube
When Flatness Does (Not) Guarantee Adversarial Robustness
Nils Philipp Walter, CISPA Helmholtz Center for Information Security, Germany
arXiv: https://arxiv.org/pdf/2510.14231
9. February 2026 @ 5pm CET — ▶️ YouTube
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
Yedi Zhang, Gatsby Computational Neuroscience Unit, University College London, UK
arXiv: https://arxiv.org/pdf/2512.20607
The paper on Muon Yedi mentioned in the talk is now on arXiv: https://arxiv.org/abs/2603.00742
19. January 2026 @ 5pm CET — ▶️ YouTube
Fast Video Generation (multiple papers)
Rahim Entezari, Wayve.ai
12. January 2026 @ 5pm CET — ▶️ YouTube
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
Ting Han, Lamarr Institute, TU Dortmund, Germany and Institute for AI in Medicine, UK Essen, Germany
OpenReview: https://openreview.net/pdf?id=lbtOctHDQ3
Contact us for questions or suggestions via efficientml@gmail.com.
Self-nominations to present your published work in the reading group are welcome.
Olga Saukh
(primary contact)
Linara Adilova
(primary contact)