"This conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields. The New Trends in Statistical Learning VI conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields.
The program will feature two courses and six one-hour invited talks, designed to strike a balance between theoretical and applied perspectives. Topics will cover a wide range of areas, from mathematical foundations and algorithmic developments to applications in data science, machine learning, neuroscience, and the social sciences.
Each speaker is expected to introduce their domain, highlight the state of the art, key challenges, and open questions of broad interest. Participants are encouraged to engage actively with the speakers and with each other throughout the sessions.
To foster discussion and collaboration, all presentations will take place in the morning, leaving afternoons free for informal exchanges, small working groups, and scientific discussions on the day’s topics.
To preserve the interactive and friendly atmosphere of the meeting, the number of participants is intentionally limited.
We hope this conference will once again create new opportunities for dialogue and collaboration across disciplines and strengthen connections among researchers with complementary perspectives."
The organizing commitee.
This event has benefited from support by the French National Research Agency (ANR), through the CAMELOT project.
Scientific Committee
Mohamed Hebiri (Lama, Gustave Eiffel University)
Karim Lounici (CMAP, Ecole Polytechnique)
Katia Meziani (Ceremade,University Dauphine-PSL)
Madalina Olteanu (Ceremade,University Dauphine-PSL)
Organizing Commitee
Katia Meziani (Ceremade,University Dauphine-PSL)
Karim Lounici (CMAP, Ecole Polytechnique)
Participants
Akhavan Arya (CMAP, École Polytechnique de Paris)
Brunel. Victor-Emmanuel (CREST, ENSAE)
Chzhen Evgenii (CNRS, Paris-Saclay University)
Denis Christophe (Panthéon Sorbonne University)
Chérief-Abdellatif Badr-Eddine (CNRS, Sorbonne University)
Fermanian Jean-Baptiste (Montpellier University - Inria)
Hebiri Mohamed (Lama, Gustave Eiffel University)
Kostic Vladimir (Instituto Italiano di Tecnologia-Genova)
Lounici Karim (CMAP, École Polytechnique de Paris)
Novelli Pietro (Instituto Italiano di Tecnologia-Genova)
Mirzaei Erfan (Instituto Italiano di Tecnologia-Genova)
Frohlich Alek (Instituto Italiano di Tecnologia-Genova)
Halconruy Helene ( Telecom-Sud Paris)
Meziani Katia (Ceremade,University Dauphine-PSL)
Minasyan Arshak (CentraleSupélec - Paris-Saclay University)
Mourtada Jaouad (CREST, ENSAE)
Ndiaye Aminata (Ceremade, Dauphine-PSL University)
Olteanu Madalina (Ceremade, Dauphine-PSL University)
Pontil Massimiliano (Instituto Italiano di Tecnologia-Genova & University College London)
Rossi Fabrice (Ceremade, Dauphine-PSL University)
Salmon Joseph (Montpellier University)
Schreuder Nicolas (CNRS, Gustave Eiffel University)
Tiapkin Daniil (CMAP, École Polytechnique de Paris)
Tsybakov Alexander (CREST, ENSAE)
Saulpic David (CNRS, Paris Cité University)
Focused Sessions
Akhavan Arya
Title: "A dimension-free Bernstein-type inequality for self-normalised martingales and tight bounds on worst-case information gain"
Abstract: In the first part of my talk, I will present my recent work in which we introduce a dimension-free Bernstein-type tail inequality for self-normalised martingales normalised by their predictable quadratic variation. This allows direct application to infinite-dimensional Hilbert spaces, significantly broadening its applicability to non-parametric statistical settings. As an important application, it resolves the recent open problems posed by Mussi et al. (2024), providing computationally efficient confidence sequences for logistic regression with adaptively chosen RKHS-valued covariates, and establishing instance-adaptive regret bounds in the corresponding kernelised bandit setting. All of the results presented in the first part of the talk are expressed in terms of a quantity known as information gain.
The second part of the talk is dedicated to establishing tight bounds on the worst-case information gain under the Matérn and squared exponential kernels. I present a proof that relies on elementary Fourier-analytic techniques inspired by Widom (1963), avoiding the strong and unverified assumption of uniformly bounded Mercer eigenfunctions used by Vakili et al. (2021). This yields a tight characterisation of the information gain for these kernels and directly improves a range of bounds for kernel-based sequential decision-making algorithms- such as kernelised bandits and reinforcement learning algorithms- discussed in the first part of the talk.
Schreuder Nicolas
Title: "An Efficient Permutation-Based Kernel Two-Sample Test"
Abstract: Two-sample hypothesis testing---determining whether two sets of data are drawn from the same distribution---is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing, maximum mean discrepancy (MMD) has gained popularity as a test statistic due to its flexibility and strong theoretical foundations. However, its use in large-scale scenarios is plagued by high computational costs. I will show how a Nyström approximation of the MMD can be used to design a computationally efficient and practical testing algorithm while preserving statistical guarantees. I will present some techniques we used to obtain a finite-sample bound on the power of our permutation-based test and discuss its optimality.
Based on a joint work with Antoine Chatalic, Marco Letizia, and Lorenzo Rosasco (https://arxiv.org/abs/2502.13570v2).
Novelli Pietro
Title: "Operator World Models for Reinforcement Learning"
Abstract: In this Focused Session, we'll take a deep dive into Reinforcement Learning (RL), exploring it through the lens of stochastic processes and the tool of transfer operators. At the cost of a little abstraction, this "operator way" offers new insights and a fresh perspective on tackling RL problems. Crucially, it highlights how RL can be understood as two intertwined challenges: Learning the environment and optimizing behavior. These theoretical underpinnings motivate a novel algorithm that combines policy mirror descent with conditional mean embeddings, for which theoretical convergence rates can be rigorously established.
Halconruy Helene
Title: "LDP drift parameter estimation for i.i.d. paths of diffusion processes"
Abstract: As large-scale sensitive data becomes more common, balancing privacy and utility is increasingly
important. Consider a clinical trial involving N patients, where drug diffusion is modelled by an
stochastic differential equation (SDE). How can we estimate a drift parameter while preserving each
patient's privacy? This question falls under the framework of (local) differential privacy (LDP). Most work in
statistical inference under LDP has focused on independent random variables without temporal
structure, raising challenges in hypothesis testing and estimation. In joint work with Chiara Amorino and Arnaud Gloter, we tackle drift parameter estimation from N i.i.d. diffusion paths under LDP. We introduce a pseudo-likelihood contrast function and add carefully scaled Laplace noise to preserve privacy. We establish conditions ensuring privacy, consistency, and asymptotic normality of the estimator. In this talk, I’ll introduce LDP, illustrate private randomized algorithms, and present our LDP mechanism for SDEs and its statistical guarantees.
Meziani Katia
Title: "Clustering Beyond Limits: Scalable Algorithms and Federated Accessibility"
Abstract: In this talk, we introduce CoHiRF (Consensus Hierarchical Random Features), a novel meta-clustering algorithm designed to enhance the scalability of existing clustering methods such as k-means, DBSCAN, and others. CoHiRF acts as a wrapper, enabling these methods to scale efficiently both in the number of observations (n) and in dimensionality (p), while preserving performance and interpretability. A key strength of CoHiRF lies in its interpretability. As a hierarchical method, it does not require specifying the number of clusters in advance and provides a clear view of how clusters progressively merge at each iteration, offering insights into the data structure at multiple levels of granularity.
We further extend CoHiRF to a federated setting with CoCoHiRF (Collaborative CoHiRF in Vertical Federated Learning), allowing multiple agents - each holding partial feature sets for the same individuals-to collaborate without sharing raw data. Our approach ensures privacy by limiting communication to simple exchanges of integer vectors (e.g., cluster labels or identifiers), with no central server required. We present promising empirical results for CoHiRF in centralized settings, and demonstrate significant performance gains in federated scenarios, particularly under strong feature partitioning.
This work is the result of a collaboration with Vladimir Kostić, Bruno Belluci-Teixeira, Karim Lounici and Erfan Mirzaei .
Mourtada Jaouad
Title: "Finite-sample performance of the maximum likelihood estimator in logistic regression"
Abstract: Logistic regression is a classical model for describing the probabilistic dependence of binary responses to multivariate covariates. We consider the predictive performance of the maximum likelihood estimator (MLE) for logistic regression, assessed in terms of logistic risk. We consider two questions: first, that of the existence of the MLE (which occurs when the dataset is not linearly separated), and second that of its accuracy when it exists. These properties depend on both the dimension of covariates and on the signal strength. In the case of Gaussian covariates and a well-specified logistic model, we obtain sharp non-asymptotic guarantees for the existence and excess logistic risk of the MLE. We then generalize these results in two ways: first, to non-Gaussian covariates satisfying a certain two-dimensional margin condition, and second to the general case of statistical learning with a possibly misspecified logistic model. Finally, we consider the case of a Bernoulli design, where the behavior of the MLE is highly sensitive to the parameter direction.
Tiapkin Daniil T
Title: "Game Theoretical Approaches for LLM Alignment"
Abstract: Aligning Large Language Models (LLMs) with nuanced human preferences, values, and desired behaviors is a paramount challenge in contemporary AI research. While Reinforcement Learning from Human Feedback (RLHF) has achieved notable success by training LLMs to optimize a learned scalar reward model, this paradigm faces inherent limitations in expressiveness, often struggling with complex, non-transitive, or diverse human preferences, and can be susceptible to reward hacking.
This talk explores the emerging field of game-theoretical approaches as a more robust and expressive framework for LLM alignment. We will dive into how recasting the alignment problem as a "game" – where policies compete based on human preferences – can overcome the limitations of scalar rewards. Key concepts such as Nash Equilibria, von Neumann winners, and preference modeling (e.g., P(y ≻ y'|x)) will be introduced as alternative solution concepts that naturally handle intricate preference structures and aim for policies that are demonstrably preferred against alternatives. Next, we will focus on existing algorithms and consider novel ones that allow us to examine the problem of LLM alignment from a different perspective.
Tsybakov Alexander
Title: "Conversion theorem and minimax optimality for continuum contextual bandits"
Abstract: We study the continuum contextual bandit problem, where the learner sequentially receives a side information vector (a context) and has to choose an action in a convex set, minimizing a function depending on the context. The goal is to minimize the dynamic contextual regret, which provides a stronger guarantee than the standard static regret. Considering a meta-algorithm that to any input non-contextual bandit algorithm associates an output contextual bandit algorithm, we prove a conversion theorem, which allows one to derive a bound on the contextual regret from the static regret of the input algorithm. We apply this strategy to obtain upper bounds on the contextual regret in several settings (losses that are Lipschitz, convex and Lipschitz, strongly convex and smooth with respect to the action variable). Inspired by the interior point method and employing self-concordant barriers, we propose an algorithm achieving a sub-linear contextual regret for strongly convex and smooth functions in noisy setting. We show that it achieves, up to a logarithmic factor, the minimax optimal rate of the contextual regret as a function of the number of queries. Joint work with Arya Akhavan, Karim Lounici and Massi Pontil.
Saulpic David
Title: "A review of clustering, from practice to statistics"
Abstract: The goal of clustering is to group similar points together, while at the same time separating dissimilar points. This can be formalized by different objective functions or algorithms: for points in a metric space, one may want to minimize the quantization error (and solve the k-means problem) or use algorithms such as DBSCAN ; for points in a graph, one may resort to minimizing the modularity (with Louvain algorithm) or to spectral clustering. In this talk, I will review those different algorithms and their comparative advantages. In particular, I will present situations in which those algorithms are known to succeed : in other words, what properties of the input are enough to make the algorithms recover "good" clusters. I will then compare with the landscape for subspace clustering, where each cluster is a linear subspace, highlighting few open questions.
Small Talks
Small Talks are short sessions (around 30 minutes) led by PhD students, during which they present their ongoing research. These talks encourage the sharing of ideas, highlight current scientific progress, and promote interaction between early-career researchers and the wider academic community.
Frohlich Alek
Title: "PersonalizedUS: Interpretable Breast Cancer Risk Assessment with Local Coverage Uncertainty Quantification"
Abstract: Correctly assessing the malignancy of breast lesions identified during ultrasound examinations is crucial for effective clinical decision-making. However, the current" gold standard" relies on manual BI-RADS scoring by clinicians, often leading to unnecessary biopsies and a significant mental health burden on patients and their families. In this paper, we introduce PersonalizedUS, an interpretable machine learning system that leverages recent advances in conformal prediction to provide precise and personalized risk estimates with local coverage guarantees and sensitivity, specificity, and predictive values above 0.9 across various threshold levels. In particular, we identify meaningful lesion subgroups where distribution-free, model-agnostic conditional coverage holds, with approximately 90% of our prediction sets containing only the ground truth in most lesion subgroups, thus explicitly characterizing for which patients the model is most suitably applied. Moreover, we make available a curated tabular dataset of 1936 biopsied breast lesions from a recent observational multicenter study and benchmark the performance of several state-of-the-art learning algorithms. We also report a successful case study of the deployed system in the same multicenter context. Concrete clinical benefits include up to a 65% reduction in requested biopsies among BI-RADS 4a and 4b lesions, with minimal to no missed cancer cases.
Ndiaye Aminata
Title: "MCD: Marginal Contrastive Discrimination for conditional density estimation"
Abstract: Conditional density estimation is a central problem in statistics and machine learning, particularly challenging when the dimensionality of the conditioning set is high. To address this issue, we propose an innovative method inspired by noise-contrastive methods. This approach reformulates the conditional density estimation problem into two simpler sub-problems: density estimation and binary classification. Our method, called Marginal Contrastive Discrimination (MCD), demonstrates performance that is comparable to, and sometimes surpasses, state-of-the-art conditional density estimation approaches, especially in scenarios where the dimensionality of the conditioning set is high
🚆 Arrived: Hyères Train Station
🚍 Option 1: Bus line 67 to La Tour Fondue (approx. 30 min – €2)
🚖 Taxi to La Tour Fondue (approx. 20–25 min – around €30–€40)
🛳️ Board the ferry (TLV‑TVM) to Porquerolles at La Tour Fondue (Giens Peninsula) (approx. 20 min)
The conférence will be held at the center https://www.igesa.fr/