"This conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields. The New Trends in Statistical Learning VI conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields.
The program will feature two courses and six one-hour invited talks, designed to strike a balance between theoretical and applied perspectives. Topics will cover a wide range of areas, from mathematical foundations and algorithmic developments to applications in data science, machine learning, neuroscience, and the social sciences.
Each speaker is expected to introduce their domain, highlight the state of the art, key challenges, and open questions of broad interest. Participants are encouraged to engage actively with the speakers and with each other throughout the sessions.
To foster discussion and collaboration, all presentations will take place in the morning, leaving afternoons free for informal exchanges, small working groups, and scientific discussions on the day’s topics.
To preserve the interactive and friendly atmosphere of the meeting, the number of participants is intentionally limited.
We hope this conference will once again create new opportunities for dialogue and collaboration across disciplines and strengthen connections among researchers with complementary perspectives."
The organizing commitee.
Conference Registration & Accommodation Fee
Conference Fee: €1 248,17: Includes conference registration, accommodation, and meals.
Organizing Commitee
Katia Meziani (Ceremade,University Dauphine-PSL)
Karim Lounici (CMAP, Ecole Polytechnique)
Scientific Committee
Karim Lounici (CMAP, Ecole Polytechnique)
Katia Meziani (Ceremade,University Dauphine-PSL)
Madalina Olteanu (Ceremade,University Dauphine-PSL)
Participants
Akhavan Arya (Oxford University)
Brunel Victor-Emmanuel (CREST, ENSAE)
Braun Baptiste (Ceremade, Dauphine-PSL University)
Dalalyan Arnak (CREST, ENSAE)
Danjou Arthur (CMAP, École Polytechnique de Paris)
Denis Christophe (Panthéon Sorbonne University)
El Mamdhi El Mahdi (CMAP, École Polytechnique de Paris)
Gaucher Solène (CMAP, École Polytechnique de Paris)
Ganassali Luca (CMAP, École Polytechnique de Paris)
Germain Thibaut (CMAP, École Polytechnique de Paris)
Ghariani Yassine (Ceremade, Dauphine-PSL University)
Frohlich Alek (Instituto Italiano di Tecnologia-Genova)
Hebiri Mohamed (Lama, Gustave Eiffel University)
Lelievre Tony (CMAP, École Polytechnique de Paris)
Liu Yating (Ceremade, Dauphine-PSL University)
Killick Rebecca (Lancaster University)
Kostic Vladimir (Instituto Italiano di Tecnologia-Genova)
Lounici Karim (CMAP, École Polytechnique de Paris)
Meric Axelle (Ceremade,University Dauphine-PSL)
Meziani Katia (Ceremade,University Dauphine-PSL)
Ndiaye Aminata (Ceremade, Dauphine-PSL University)
Olteanu Madalina (Ceremade, Dauphine-PSL University)
Pontil Massimiliano (Instituto Italiano di Tecnologia-Genova & University College London)
Reynaud-Bouret, Patricia (Côte d'Azur University )
Ribordy Mathis (Ceremade, Dauphine-PSL University)
Rivoirard Vincent (Ceremade, Dauphine-PSL University)
Saci Léo (Ceremade, Dauphine-PSL University)
Salmon Joseph (Montpellier University)
Schreuder Nicolas (CNRS, Gustave Eiffel University)
Focused Sessions
Arnak Dalalyan (CREST, ENSAE, GENES) Lecture
Title: Flow matching and denoising diffusions: a selective overview
Abstract : Flow Matching and Denoising diffusions are the most popular methods of generative modeling
used in Machine Learning and Artificial Intelligence. The goal of these lectures will be to introduce
these methods using rigorous mathematical language and present some recent results
explaining their good performance in practice. The first lecture will be devoted to a general overview
of generative modeling and recent developments in this field of Statistics and Machine Learning.
The formal presentation of the Flow Matching methodology will also be introduced and discussed.
The second lecture will focus on Denoising Diffusion probabilistic models and will provide more
details on theoretical results quantifying the performance of these algorithms.
El Mahdi El Mhamdi (Ecole polytechnique)
Title: Could Adversarial Machine Learning ever be Defensive?
Abstract: In this talk, we will start by reviewing the last decade of advances in defensive machine learning, with a focus on train-phase defensive measures and distributed machine learning. In particular, we argue that the most credible results point to an impossibility of full-safety guarantees, even in very mild threat models. We then present new approaches, based on offensive measures, to better capture the surface attack of machine learning, and prepare for a new era of defensive ML that is not relying solely on agnostic measures from robust statistics and robust mean estimation.
Patricia Reynaud-Bouret (CNRS, Université Côte d'Azur)
Title: Hawkes Processes for Modeling Learning in Neural Networks
Abstract: Thanks to the neurons in our brain, we are capable of learning and memorizing. Unlike artificial neural networks, these biological networks locally adjust their synaptic weights to achieve global learning. We will discuss a toy model based on Hawkes processes, where we can mathematically demonstrate that local rules can lead to global learning. Additionally, these networks can help recover measurable macroscopic quantities, such as the evolution of reaction times during a learning task. This presentation will be based on several studies conducted during Sophie Jaffard's PhD thesis, in collaboration with Samuel Vaiter, Etienne Tanré (LJAD, Nice), and Giulia Mezzadri (Columbia, USA)
Luca Ganassali (Universite Paris-Saclay)
Tilte : What is a good matching of probability measures? A counterfactual lens on transport maps
Abstract : Coupling probability measures is central to statistics and machine learning, yet transport maps are generally non-unique. The common recourse to optimal transport, motivated by cost minimization and cyclical monotonicity, obscures the fact that several distinct notions of multivariate monotone matchings coexist. In this talk, we will first compare three constructions of transport maps—cyclically monotone, quantile-preserving, and triangular maps—characterizing when they coincide and highlighting their structural properties. We will then connect this analysis to causal inference, showing how counterfactual reasoning can be framed as selecting a transport map and when causal assumptions align with classical statistical transports. Taken together, these results aim to enrich the theoretical understanding of families of transport maps and to clarify their possible causal interpretations.
This talk is based on joint work with Lucas De Lara.
Germain Thibaut (CMAP, Ecole polytechnique)
Tilte : Geometric Dictionary Learning of Dynamical Systems with Optimal Transport
Abstract: Learning dynamical systems through operator-theoretic representations provides a powerful framework for analyzing complex dynamics, as spectral quantities such as eigenvalues and invariant structures encode characteristic time scales and long-term behavior. However, dynamical operators are typically estimated independently for each system, preventing the discovery of shared structure across related dynamics. To address this limitation, we posit that related dynamical systems lie near a low-dimensional manifold in spectral operator space. Based on this hypothesis, we introduce DOODL (Dynamical OperatOr Dictionary Learning), a framework that learns a dictionary of characteristic spectral dynamics whose combinations approximate this manifold and yield compact, interpretable embeddings of individual systems.
Beyond representation learning, DOODL enables fast and interpretable operator estimation from short and partially observed trajectories by constraining the estimation to the learned operator manifold. Experiments on metastable Langevin dynamics and turbulent plasma simulations demonstrate that DOODL scales to highly complex multiscale regimes while capturing characteristic spectral structure governing the dynamics rather than merely fitting trajectories, achieving errors one to two orders of magnitude lower than independent operator estimation methods in challenging low-data regimes.
Killick Rebecca (Lancaster University)
Tilte : An Introduction to the fundamentals of changepoint analysis
Abstract: Traditional statistical model building assumes that the same model (and fitted parameters) can describe the data at any point in an observed process. To tackle early violations to this, statisticians introduced regressors to describe time-varying features such as trend and seasonality. These have enabled model building to become ingrained in everyday applications across all fields. But what happens when this neat assumption of a static model is no longer appropriate?
The simplest departure from this static model assumption is to piece together static models, which we understand well. Questions then surface around which bits of data follow which static model, and how many of different static models should we use? This field of statistical modelling is called changepoint detection. I will introduce the fundamental challenges and components required for changepoint modelling and highlight open issues in the field that will hopefully spark discussion.
Lelievre Tony (Ecole des Ponts et Chaussées)
Tilte : Challenges related to sampling in molecular dynamics: Gradient Flows and Adaptive Biasing Techniques
Abstract: I will then focus on free energy adaptive biasing techniques, and discuss convergence results for these methods, in particular a recent work in collaboration with Xuyang Lin and Pierre Monmarché. Free-energy-based adaptive biasing methods, such as Metadynamics, the Adaptive Biasing Force (ABF) and their variants, are enhanced sampling algorithms widely used in molecular simulations. Although their efficiency has been empirically acknowledged for decades, providing theoretical insights via a quantitative convergence analysis is a difficult problem, in particular for the kinetic Langevin diffusion, which is non-reversible and hypocoercive. We obtain the first exponential convergence result for such a process, in an idealized setting where the dynamics can be associated with a mean-field non-linear flow on the space of probability measures. A key of the analysis is the interpretation of the (idealized) algorithm as the gradient descent of a suitable functional over the space of probability distributions.
Liu Yating (Ceremade, University Dauphine)
Tilte : Learning drift functions in diffusion processes: from estimation to supervised classification via neural networks
Abstract: We study learning problems for time-homogeneous diffusion processes observed at discrete times, focusing on both drift estimation and supervised classification. We propose a neural network–based nonparametric estimator for the drift function using high-frequency observations from independent trajectories, and derive non-asymptotic convergence rates. For compositional drift structures, the rates show weak dependence on the dimension, and numerical experiments show clear advantages over classical B-spline methods, especially in higher dimensions. Building on this estimation framework, we introduce a neural network–based plug-in classifier for multiclass diffusion models, where each class is defined by a distinct drift function. We establish convergence rates for the excess misclassification risk and demonstrate that exploiting the diffusion structure leads to improved classification performance compared to direct end-to-end neural network classifiers. This presentation is based on the results developed in [Zhao, Liu, Hoffmann (2025); Zhao, Fan, Liu (2026)].
Salmon Joseph (Inria)
Tilte : Conformal Prediction for Long-Tailed Classification
Abstract: Many real-world classification problems, such as plant identification, have extremely long-tailed class distributions.
In order for prediction sets to be useful in such settings, they should
(i) provide good class-conditional coverage, ensuring that rare classes are not systematically omitted from the prediction sets,
(ii) be a reasonable size, allowing users to easily verify candidate labels.
Unfortunately, existing conformal prediction methods, when applied to the long-tailed setting, force practitioners to make a binary choice between small sets with poor class-conditional coverage or sets with very good class-conditional coverage but that are extremely large.
We propose methods with guaranteed marginal coverage that smoothly trade off between set size and class-conditional coverage. First, we introduce a new conformal score function called prevalence-adjusted softmax that targets macro-coverage, a relaxed notion of class-conditional coverage.
Second, we propose a new procedure that interpolates between marginal and class-conditional conformal prediction by linearly interpolating their conformal score thresholds.
We demonstrate our methods on Pl@ntNet-300K and iNaturalist-2018, two long-tailed image datasets with 1,081 and 8,142 classes, respectively.
https://arxiv.org/abs/2507.06867
Learning dynamical systems through operator-theoretic representations provides a powerful framework for analyzing complex dynamics, as spectral quantities such as eigenvalues and invariant structures encode characteristic time scales and long-term behavior. However, dynamical operators are typically estimated independently for each system, preventing the discovery of shared structure across related dynamics. To address this limitation, we posit that related dynamical systems lie near a low-dimensional manifold in spectral operator space. Based on this hypothesis, we introduce DOODL (Dynamical OperatOr Dictionary Learning), a framework that learns a dictionary of characteristic spectral dynamics whose combinations approximate this manifold and yield compact, interpretable embeddings of individual systems.
Beyond representation learning, DOODL enables fast and interpretable operator estimation from short and partially observed trajectories by constraining the estimation to the learned operator manifold. Experiments on metastable Langevin dynamics and turbulent plasma simulations demonstrate that DOODL scales to highly complex multiscale regimes while capturing characteristic spectral structure governing the dynamics rather than merely fitting trajectories, achieving errors one to two orders of magnitude lower than independent operator estimation methods in challenging low-data regimes.
🚆 Arrived: Hyères Train Station
🚍 Option 1: Bus line 67 to La Tour Fondue (approx. 30 min – €2)
🚖 Taxi to La Tour Fondue (approx. 20–25 min – around €30–€40)
🛳️ Board the ferry (TLV‑TVM) to Porquerolles at La Tour Fondue (Giens Peninsula) (approx. 20 min)
The conférence will be held at the center https://www.igesa.fr/