The cluster welcomes interest from anyone who is looking pursuing a PhD in our areas of interest, which broadly encompass statistics, probability, and mathematical biology. We recommend that you start with an informal approach to potential supervisors who work in areas of interest to you; they will be able to discuss potential projects and give you advice about applications. You can then apply here.
EPSRC funded studentships: Typically, there are a small number of fully funded places available each year and decisions on these are made between January and March for an October start. Funding information can be found here. Other start dates during the year are possible, but PhD funding is only allocated once a year.
Below are some current projects on offer within the research cluster.
Diseases and pests of trees lead to large economic costs to the forestry industry, limit natural spaces for recreation and impact climate change by reducing trees’ ability to capture carbon. A key element of infection specific to plants is spatial structure: since plants are usually fixed in location, infection tends to be localised. In this project you would develop advanced, spatially-explicit mathematical models using nonlinear ordinary differential equations and computational individual-based models. You will fit your models to real data of disease spread and then test different management and mitigation strategies.
Antagonistic coevolution between parasites and the hosts they infect is a major driver of disease dynamics. Mathematical models have been used for some time to understand these co-evolutionary dynamics, however, they tend to assume constant environments and no external interactions with other species. In this project you will develop coevolutionary models of hosts and parasites within the adaptive dynamics framework, primarily using nonlinear ordinary differential equations and some stochastic numerical models. You will incorporate complex environments either through fluctuating parameters (for example representing seasonality) or external interactions (such as competiiton with other species).
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Network archaeology studies the problem of reconstructing the temporal history of a network when only its final state is observed. Most mathematical work on network archaeology for random trees focuses on root finding. Here, given a tree T and accuracy \epsilon, the goal is to output a (minimal) subset S of vertices in T such that the probability that the root is within the set S is at least 1-\epsilon. The project will investigate root finding and other problems in preferential attachment trees with vertex removal. This random tree model generalises popular models such as affine preferential attachment and uniform attachment by including vertex removal. In the PAVER model, starting from a single root vertex, a random tree is constructed recursively by, at every step, either connecting a new vertex to a randomly selected vertex in the tree or by removing the selected vertex from the tree. The project will (1) extend known results for root finding in preferential attachment and uniform attachment models to the more general setting of PAVER models, and (2) develop and investigate algorithms to find the oldest unremoved vertex in the tree, a generalised root finding problem that is novel and unique to PAVER models (compared to models without vertex removal).
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
Random graph models are probabilistic models for the development of networks. They are often inspired by networks found in areas such as computing, social networks and biology, but the emphasis is on studying mathematical properties of the graphs. One family of random graph models which has been studied extensively for the last 25 years or so is the preferential attachment model and its variations. Some of these variations feature an interesting phenomenon often referred to as condensation, where a small proportion of vertices become unusually important in the network in that a large proportion of edges connect to them. The aim of this project will be to study this phenomenon further and its occurrence in new graph models.
The location of individual animals and species emerge from a complex combination of their interactions with one another and with the underlying environment. This combination makes it difficult to predict how species distributions may change as the landscape is altered (e.g. by human infrastructure or the effects of climate change). Mathematical models can help by providing a rigorous framework for understanding how the mechanisms of between-species interactions and animal-landscape interactions lead to the emergent patterns of location distributions observed in nature. This project seeks to further that understanding by analysing pattern formation in a class of partial differential equation (PDE) systems in heterogeneous environments. The technical difficulty arises from the fact that most PDE pattern formation studies assume a homogeneous environment, but this study will build on recent techniques of the supervisor for dealing with heterogeneous environments. If you are interested in this project, please contact the supervisor on the above link for more information.
This project studies a recently-proposed class of PDE systems for studying the spatial arrangement of interacting populations of animals (and, in some situations, cells). Numerical investigations have revealed rich pattern formation properties in these systems, whereby continuous changes in parameters can cause qualitative changes in spatial patterns (bifurcations), which in some circumstances are discontinuous (tipping points). This project will aim to understand these bifurcations and tipping points analytically. We will use traditional tools, such as Crandall-Rabinowitz bifurcation theory and weakly non-linear analysis, to analyse these phenomena, but also explore the extent to which we can ease the complicated calculations that arise, through careful use of large language models. Such usage is quite novel and has only been made potentially possible through the very recent rapid developments in LLM technology. Please contact the supervisors on the above links for more information.
Index trackers (ETFs, index funds, index mutual funds) aim to replicate the return of a reference equity index while keeping implementation frictions---notably turnover and trading costs---under control. Index tracking is a passive investment approach that frequently beats their actively managed counterparts over the long term. In practice, full index-replication can be costly or operationally infeasible; consequently, pure tracking with a sparse basket is common for segregated mandates, often subject to constraints such as a full-investment budget and, in some mandates, long-only weights. A large body of work formulates index tracking as an optimisation problem with convex losses and penalties, or as mixed-integer programs that directly control support size and transactions. These approaches typically deliver a single allocation weight vector and tune regularisation hyperparameters by heuristics or cross-validation. Such tuning requires manual interventions and is often subject to uncertainty.
This project develops novel Bayesian methodology with the aim to return a posterior over weights, tuning parameters, and the residual variance that drives tracking error. The aim is decision-grade uncertainty quantification (UQ) and principled data-driven parameter selection. Hence, the project will develop automatic portfolio rebalancing driven by UQ analysis, based on the posterior distribution of the allocation weights.
We will model the tracking residual returns—index minus portfolio—as a noisy outcome with unknown variance, regressing the index returns on the returns of the individual assets. Portfolio weights live in a constrained space reflecting business rules: non-negativity or signed positions, leverage or budget limits, optional sector exposures, and explicit cardinality. Sparsity comes from imposing hierarchical priors that prefer few active names: either spike-and-slab families (selection indicators plus a continuous “slab” for active weights) or continuous shrinkage alternatives such as horseshoe or Dirichlet-Laplace. Tuning parameters are not fixed numbers in this approach: global and local shrinkage scales, inclusion probabilities, tracking-error and turnover multipliers, and the residual variance all carry their own priors. The project is expected to result in a coherent hierarchy that links portfolio structure, transaction costs, prior beliefs, and a data-driven narrative for model selection.
Inference combines two ingredients. First, state-of-the-art MCMC samplers, which are designed to handle non-smooth penalties and hard constraints. These samplers perform gradient-informed random walks while enforcing feasibility at every step. These take advantage of recent advances and contribute to convex nonsmooth optimisation and numerical analysis of stochastic differential equations. Second, Gibbs or Metropolis updates for the hyperparameters: residual variance, global shrinkage, local scales, and selection indicators.
Alongside expert-set priors, the project proposes to also use machine learning techniques to learn parts of the prior and penalty structure from data in an interpretable way. Examples include mapping liquidity and volatility features to a prior preference for sparsity, learning sector-level shrinkage from historical co-movement, or learning a turnover multiplier that depends on spread and market depth. The plan is to start with simple parametric maps and progress to neural networks with monotonicity or convexity constraints so that learned priors remain stable and defensible. Expert ranges can be encoded as hyperpriors, allowing the data to refine but not overrule domain knowledge.
Please email the supervisor for more information
Please email the supervisor for more information
Please email the supervisor for more information
The cost and duration of traditional clinical trials imply a high societal and economic cost, hampering the release of new treatments to public. One possible avenue to improve on this situation is the use of mathematical models to produce synthetic data, which can then be shared with contemporary clinical trials to reduce the number of patients and the duration of the trial. The approach can take advantage of well established mathematical models of diseases, or use current information to generate statistical forecasting models for less well understood systems. A Bayesian approach allows combining disparate sources of information, along with coherent uncertainty propagation, fundamental to feed into validation, verification and uncertainty qualification needed for clinical approval of these novel medical devices.
Panel (longitudinal) data enables learning the dynamics and relations of (groups of) units, strengthening the inference on both cross-sectional and dynamic parameters. The dominant approach to modelling unbounded continuous quantities of interest is to assume symmetric error terms (eg Gaussian or Student), but this assumption is often not met in practice. The idea is to explore, from a Bayesian perspective, the use of skew error distributions from well-known skewing mechanisms (eg Azzalini, 1985; Ley 2010), while allowing the parameter(s) controlling the skewness to vary over time. This implies the design of a stochastic path to control the process in time and exploring conditions for convergence.
The Brownian web is a dense system of coalescing Brownian motions in 1+1 dimensional space time. It is known to be the universal scaling limit of a wide variety of models, ranging from population genetics and interface growth to theoretically important processes such as orientated percolation and directed spanning forests. The project will develop scaling limit theory for dense interacting systems of particles in cases where the limiting particle motions are not simply Brownian motions e.g. they might feature spatial inhomogeneity or heavy tailed behaviour.
Please email the supervisor for more information
Large random trees have well-known discrete and continuous limits which have been studied since the 90's. The aim of this project is to investigate how much changes if we force the trees to be in unusual regimes.
Life cycle assessment underpins environmental decision-making across industry and policy, but its practical use is limited by incomplete process data, the manual effort required to build and validate models, and the difficulty of adapting existing assessments to new materials or products. This project develops AI methods along three complementary lines. The first uses large language model based reasoning agents to automate parts of the assessment workflow itself, including reading technical documentation, navigating supply chains, and selecting appropriate modelling assumptions. The second develops probabilistic models, including neural processes and knowledge graphs, to complete missing inventory data with calibrated uncertainty when primary data is unavailable. The third uses active learning to prioritise which data gaps most affect the reliability of an assessment. The methods will be validated on industrial case studies with partners including AMRC and international collaborators, and the underlying approach extends naturally to broader circular economy questions such as material flow tracking and end of life material reuse.
Engineering design problems, from analog circuit sizing to physical simulation, share a common structure: expensive evaluations, high dimensional design spaces, and a need to combine data driven search with domain expertise. This project develops methodology for reliable AI assisted design along three complementary lines. The first is multi-fidelity Bayesian fusion, which combines cheap approximate simulations with expensive accurate ones to reduce the cost of exploring large design spaces. The second is Bayesian optimization with structured expert priors, which encodes domain knowledge as statistical priors rather than black box features so that optimization remains sample efficient and interpretable. The third is large language model based reasoning agents that propose candidate designs and verify them against domain constraints, aiming to make agentic search auditable rather than opaque. The methods will be validated primarily on analog and mixed signal circuit design, including transistor sizing, yield estimation, and simulation acceleration, an area with an established track record of peer reviewed results at DAC, ICCAD, and ASP-DAC. The underlying methodology is not specific to circuits and extends naturally to other engineering domains with similar cost and reliability constraints.