Machine learning is increasingly transforming actuarial modelling by offering flexible tools to capture complex, nonlinear relationships in insurance data. Yet, for actuarial applications, predictive accuracy alone is not sufficient: models must also remain interpretable, robust, and compatible with the statistical and regulatory requirements of insurance practice.
My research in this area focuses on the interaction between machine learning and actuarial modelling, with particular emphasis on interpretable learning methods, extreme-risk modelling, and individual claims reserving. A recurring objective is to combine the predictive power of modern machine-learning techniques with the transparency and statistical structure of traditional actuarial models. This includes extracting simple decision rules from complex models, translating tree-based predictions into interpretable regression frameworks, and developing data-driven approaches for heterogeneous claims and tail risks.
More broadly, this research aims to develop machine-learning methods that actuaries can not only use for prediction, but also understand, validate, and integrate into insurance decision-making.
Maillart, A. and Robert, C. (2024) — “Distill knowledge of additive tree models into generalized linear models: a new learning approach for non-smooth generalized additive models”
Annals of Actuarial Science, 18(3), 692–711.
Abstract
Generalized additive models (GAMs) are a leading model class for interpretable machine learning. GAMs were originally defined with smooth shape functions of the predictor variables and trained using smoothing splines. Recently, tree-based GAMs where shape functions are gradient-boosted ensembles of bagged trees were proposed, leaving the door open for the estimation of a broader class of shape functions (e.g. Explainable Boosting Machine (EBM)). In this paper, we introduce a competing three-step GAM learning approach where we combine (i) the knowledge of the way to split the covariates space brought by an additive tree model (ATM), (ii) an ensemble of predictive linear scores derived from generalized linear models (GLMs) using a binning strategy based on the ATM, and (iii) a final GLM to have a prediction model that ensures auto-calibration. Numerical experiments illustrate the competitive performances of our approach on several datasets compared to GAM with splines, EBM, or GLM with binarsity penalization. A case study in trade credit insurance is also provided.
Maillart, A. and Robert, C. (2023) — “Tail index partition-based rules extraction with application to tornado damage insurance”
ASTIN Bulletin, 53(2), 258–284.
DOI : 10.1017/asb.2023.1
Abstract
The tail index is an important parameter that measures how extreme events occur. In many practical cases, this tail index depends on covariates. In this paper,we assume that it takes a finite number of values over a partition of the covariate space. This article proposes a tail index partition-based rules extraction method that is able to construct estimates of the partition subsets and estimates of the tail index values. The method combines two steps: first an additive tree ensemble based on the Gamma deviance is fitted, and second a hierarchical clustering with spatial constraints is used to estimate the subsets of the partition. We also propose a global tree surrogate model to approximate the partition-based rules while providing an explainable model from the initial covariates. Our procedure is illustrated on simulated data. A real case study on wind property damages caused by tornadoes is finally presented
Baudry, M. and Robert, C. (2019) — “A machine learning approach for individual claims reserving in insurance”
Applied Stochastic Models in Business and Industry, 35(5), 1127–1155.
DOI : 10.1002/asmb.2455
Abstract
Accurate loss reserves are an important item in the financial statement of an insurance company and are mostly evaluated by macrolevel models with aggregate data in run-off triangles. In recent years, a new set of literature has considered individual claims data and proposed parametric reserving models based on claim history profiles. In this paper, we present a nonparametric and flexible approach for estimating outstanding liabilities using all the covariates associated to the policy, its policyholder, and all the information received by the insurance company on the individual claims since its reporting date. We develop a machine learning–based method and explain how to build specific subsets of data for the machine learning algorithms to be trained and assessed on. The choice for a nonparametric model leads to new issues since the target variables (claim occurrence and claim severity) are right-censored most of the time. The performance of our approach is evaluated by comparing the predictive values of the reserve estimates with their true values on simulated data. We compare our individual approach with the most used aggregate data method, namely, chain ladder, with respect to the bias and the variance of the estimates. We also provide a short real case study based on a Dutch loan insurance portfolio.