"This conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields. The New Trends in Statistical Learning VI conference aims to bring together researchers from diverse backgrounds to present and discuss current and emerging directions in statistical learning and related fields.
The program will feature two courses and six one-hour invited talks, designed to strike a balance between theoretical and applied perspectives. Topics will cover a wide range of areas, from mathematical foundations and algorithmic developments to applications in data science, machine learning, neuroscience, and the social sciences.
Each speaker is expected to introduce their domain, highlight the state of the art, key challenges, and open questions of broad interest. Participants are encouraged to engage actively with the speakers and with each other throughout the sessions.
To foster discussion and collaboration, all presentations will take place in the morning, leaving afternoons free for informal exchanges, small working groups, and scientific discussions on the day’s topics.
To preserve the interactive and friendly atmosphere of the meeting, the number of participants is intentionally limited.
We hope this conference will once again create new opportunities for dialogue and collaboration across disciplines and strengthen connections among researchers with complementary perspectives."
The organizing commitee.
Organizing commitee
Mohamed Hebiri (Lama, Gustave Eiffel University) and Katia Meziani (Ceremade,University Dauphine-PSL)
Other participants
Akhavan Arya (CMAP,École Polytechnique de Paris)
Bouret Patricia (CNRS,Côte d'Azur University)
Celisse Alain (Panthéon Sorbonne University)
Chérief-Abdellatif Badr-Eddine (CNRS, Sorbonne University)
Cheysson Félix (Gustave Eiffel University)
Chzhen Evgenii (CNRS, Paris-Saclay University)
Denis Christophe (Gustave Eiffel University)
Fermanian Jean-Baptiste (Paris-Saclay University)
Khaleghi Azadeh (CREST, ENSAE)
Kostic Vladimir (Istituto Italiano di Tecnologia-Genova)
Lounici Karim (CMAP,École Polytechnique de Paris)
Mourtada Jaouad (CREST, ENSAE)
Ndiaye Aminata (Ceremade, Dauphine-PSL University)
Olteanu Madalina (Ceremade, Dauphine-PSL University)
Pacreau Gregoire (CMAP,École Polytechnique de Paris)
Pontil Massimiliano (Istituto Italiano di Tecnologia-Genova & University College London)
Rivoirard Vincent (Ceremade, Dauphine-PSL University)
Rossi Fabrice (Ceremade, Dauphine-PSL University)
Salmon Joseph (Montpellier University)
Schreuder Nicolas (CNRS, Gustave Eiffel University)
Taturyan Gayane (Gustave Eiffel University)
Valade Florian (Gustave Eiffel University)
Program
The objective of this "focused session" is to introduce a particular field by presenting its vocabulary, context, issues, and the state of the art, from both theoretical and practical perspectives. This format not only allows for the discovery of a topic but also deepens one's knowledge and keeps one informed of the latest research developments at a given point in time.
Hebiri Mohamed :
" Conformal Predictor Insights"
Conformal prediction is a popular approach for constructing prediction sets with a prescribed coverage. Importantly, the theoretical validity of the resulting prediction sets is almost distribution-free -- they only require exchangeability of the data. A key feature of this framework is that it can be applied to any machine learning algorithm. In this talk, we provide a comprehensive overview of conformal prediction through a step-by-step presentation of various conformal prediction methods. We highlight the main theoretical tools needed to prove the validity of conformal predictors and illustrate both the advantages of this framework and its limitations (particularly when dealing with the width of prediction sets).
Rossi Fabrice :
"Causality Insights"
Bouret, Patricia:
"Statistic for learning data"
When a human or animal learns a rule, a categorization etc, he/she/it can do it only once, whereas the way this learning has been done is essentially individual. So inferring parameters or doing model selection for these learning data means in fact that we are doing statistics on non stationary data with no repetitions. The purpose of this talk is to show what can be said about these statistical problems.
Chzhen, Evgenii:
"Online Adversarial reinforcement learning."
we consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environment during $T$ episodes, each of which consists of $H$ stages, and each episode is evaluated with respect to a reward function that will be revealed only at the end of the episode. In talk I will present some recent results on learning in such an environment. Leveraging generic black-box online learning algorithm, I will describe an approach that gives \sqrt{S} improvement over previous methods that are based on occupancy measures. The talk does not assume any prior knowledge of RL theory and a short introduction is provided.The talk is based on joint work with D. Tiapkin and G. Stoltz
Khaleghi, Azadeh:
"Nonparametric methods for stationary ergodic time-series"
In this talk I will focus on some results for time-series analysis in the case where long-range dependencies are present. The relevant literature on this topic typically involves such parametric structural assumptions as autoregressive or Markovian models. However, the theoretical guarantees obtained under standard modelling assumptions do not hold in the presence of long-range dependencies. I will recall a paradigm based on ergodic theory which allows to view the observations as sample-paths of ergodic measure-preserving transformations, reducing the inference problem to that on a metric space of probability measures on \mathbb{R}^{\mathbb N }. Ergodicity ensures the point-wise convergence of empirical measures to the true probability measures without the need to impose assumptions on the memory of the processes. I will discuss some opportunities and limitations of this approach in the context of change-point estimation, time-series clustering, and restless bandits.
Kostic, Vladimir:
"Statistical perspective on dynamical representation learning"
We address novel formulation of representation learning for dynamical systems that is based on statistical learning theory of operator regression. Coupled, learned representation space and the estimated operator on that space, allow efficient forecasting of state distributions for diverse stochastic processes, leading to scientific discoveries across disciplines. In this talk we present some theoretical and empirical results that bridge statistical learning theory and deep learning, we pose diverse open questions and point out exciting avenues for future work on machine learning for dynamical systems.
Mourtada, Jaouad:
"Estimation of discrete distributions in relative entropy, and deviations of the missing mass"
We consider the problem of estimating a distribution on a finite set (alphabet) using an i.i.d. sample. More precisely, we are interested in estimators achieving a small relative entropy (Kullback-Leibler divergence) with respect to the true distribution, with high probability over the sampling of the dataset. While the empirical distribution (or MLE) is arguably the most natural estimator, it turns out to be inadequate for this problem, since it may underestimate the frequency of certain classes, or even miss them altogether. A simple correction due to Laplace consists in smoothing the empirical distribution by adding 1 to the count of each class. The resulting estimator is known to be minimax-optimal in expectation, yet its tail behavior is not fully understood.
In this talk, we will describe the best high-probability guarantees of the Laplace estimator. We will also discuss the best achievable high-probability guarantees (by any estimator) in a minimax sense; these rates are slightly larger than one might expect, and exhibit a separation between confidence-independent and confidence-dependent estimators.
Finally, we will present a modified estimator, which unlike the Laplace estimator adapts to the "effective support size" or "effective sparsity" of the true distribution. If time permits, we will also discuss the question of bounding the "missing mass" from the sample, which plays an essential role in the analysis.
Pacreau, Gregoire:
"Introduction to asset management"
In this talk, we present the theoretical bases of portfolio selection and risk management, going from the historical foundations to the most recent advances.
We introduce metrics used to analyse the performance of a portfolio, as well as the field known as "technical analysis", and explore their limitations.
We also detail the Basel 3 framework for risk management, created as an answer to the 2008 crisis, and present recent works taking into account market microstructure, systematic risk and market impact.
Salmon, Joseph:
"Crowdsourcing and supervised learning"
In supervised learning - for instance in image classification - modern massive datasets are commonly labeled by a crowd of workers. The obtained labels in this crowdsourcing setting are then aggregated for training. The aggregation step generally leverages a per-worker trust score. Yet, such worker-centric approaches discard each task's ambiguity. Some intrinsically ambiguous tasks might even fool expert workers, which could eventually be harmful to the learning step. In a standard supervised learning setting - with one label per task - the Area Under the Margin (AUM) is tailored to identify mislabeled data. We adapt the AUM to identify ambiguous tasks in crowdsourced learning scenarios, introducing the Weighted AUM (WAUM). The WAUM is an average of AUMs weighted by task-dependent scores. We show that the WAUM can help discard ambiguous tasks from the training set, leading to better generalization or calibration performance. We report improvements over existing strategies for learning a crowd, both for simulated settings and for the CIFAR-10H, LabelMe and Music crowdsourced datasets.
Practical informations
How to get there : https://www.bateaux-taxi.com/
The conférence will be held at the center https://www.igesa.fr/