10:45-10:50 Openning remark
《Session 1》
10:50-11:30 Spatio-temporal Coarse-grained Hawkes Processes for Aggregated Count Data
Tomoya Uda (The Graduate University for Advanced Studies, SOKENDAI)
Hawkes processes are widely used to model self-exciting event data but typically require precise event times and locations. We develop a class of stochastic processes for modeling counts aggregated over spatial regions and time intervals. The proposed model accounts for excitation within spatial and temporal bins and provides higher-order approximations to stationary moments than conventional binned Poisson approaches. Parameters of the underlying continuous-time Hawkes process are inferred from first- and second-order moment conditions of the aggregated data. We evaluate the predictive performance and computational efficiency of the proposed method using earthquake data and discuss practical challenges in its application to real-world spatio-temporal data. This is a joint work with Shinsuke Koyama.
11:30-12:10 Scalable Estimation of Crossed Random Effects Models via Multi-way Discretization
Shonosuke Sugasawa (Keio University)
Cross-classified data frequently arise in scientific fields such as education, healthcare, and social sciences. A common modeling strategy is to introduce crossed random effects within a regression framework. However, this approach often encounters serious computational bottlenecks, particularly for non-Gaussian outcomes. In this talk, we propose a scalable and flexible method that approximates the distribution of each random effect by a discrete distribution, effectively partitioning the random effects into a finite number of representative values. This approximation allows us to express the model as a multi-way discrete structure, which can be efficiently estimated using a simple and fast iterative algorithm. The proposed method accommodates a wide range of outcome models and remains applicable even in settings with more than two-way cross-classification. We theoretically establish the consistency and asymptotic normality of the estimator under general settings of classification levels. Through simulation studies and real data applications, we demonstrate the practical performance of the proposed method in logistic, Poisson, and ordered probit regression models involving cross-classified structures. This is a joint work with Shota Takeishi.
12:10-13:40 Lanch break
《Session 2》
13:40-14:20 Hierarchical Bayesian Modeling and Poststratification for Income Distributions in Cross-classified Subpopulations
Yuki Kawakubo (Chiba University)
We consider the estimation of income distributions for subpopulations defined by geographic areas, age groups, and other categorical characteristics. Household income is assumed to follow a lognormal distribution within each cross-classified cell. The cell-specific mean income and variance of log income are modeled hierarchically through effects associated with the categorical variables. The estimated cell-specific distributions are then combined using known population cell sizes to construct income distributions for subpopulations and arbitrary aggregations of cells. This poststratification approach adjusts for differences between the sample and population compositions without requiring individual prediction or simulation of nonsampled population units. We apply the proposed method to household income data from a consumer panel conducted by a private company in Japan.
14:20-15:00 Complete Sample Likelihood Estimation and Inference for Single and Multi-Stage Sample Surveys
Robert Clark (Australian National University)
In the analysis of sample surveys, accounting for informative sampling design is critical to ensure estimates of model parameters such as regression coefficients are not systematically biased. One popular solution is maximum pseudo-likelihood estimation, which incorporates weights given by functions of the selection probabilities. We propose an alternative method for analytical estimation and inference in sample surveys known as complete sample likelihood or CSL. This approach jointly postulates a population regression model for the responses, and a model for the selection probabilities conditional on the the response and measured covariates. CSL can offer potentially substantial reductions in bias and standard errors compared to maximum pseudo-likelihood estimation, while also avoiding a major drawback of existing plug-in sample likelihood techniques which require burdensome resampling techniques to perform inference. We develop general forms for CSL and theoretically justify their usage in both single- and multi-stage sample surveys, before deriving several tractable cases involving generalized linear (mixed) models commonly seen in practice with sample surveys. Simulation studies and an application to a two-stage sample survey of smoking status in New Zealand demonstrate the superior finite sample performance of CSL, and its ability to facilitate deployment of likelihood-based techniques and diagnostic tools. This is a joint work with Francis KC Hui.
15:00-15:30 Cofee break
《Session 3》
15:30-16:10 A Linear Model of Coregionalization for Spatial Generalized Estimating Equations
Toshihiro Hirano (Kanto Gakuin University)
Generalized estimating equations (GEEs) are a widely used marginal modeling approach to analyzing correlated data in many fields. In this article, focusing on multivariate spatial data, we propose a new class of GEEs which utilizes a linear model of coregionalization (LMC) as the working correlation matrix to flexibly capture spatial correlations between within and across responses. The proposed LMC-GEE adopts a two-step iterative algorithm for estimation, combining a Newton-Raphson update of the regression coefficients with maximum pseudo-likelihood estimation of the remaining parameters characterizing the marginal variance and working correlation matrix. In the latter, we leverage recent developments in the coregionalization literature to facilitate scalable computation of the (inverse) working correlation matrix. Extensive simulation studies and an analysis of multivariate abundance data from a Southern Ocean Continuous Plankton Recorder survey demonstrate that the proposed LMC-GEE performs well in terms of the point estimation and inference, as well as selection of the working correlation matrix based on adaptation of existing criteria for selecting working correlations in GEEs. This is a joint work with Quan Vu, Francis K. C. Hui, and Alan H. Welsh.
16:10-16:50 Simultaneous Inference for Latent Variable Predictions in Factor Analytic Models
Francis K. C. Hui (Australian National University)
Factor analytic models, also known as latent variable models or LVMs, are widely used for modeling multi-response data in fields such as agriculture and the social sciences. Although considerable research has been done on estimation and inference for model parameters in LVMs, most notably covariate effects and loading matrices, much less attention has been devoted to inference on the factor scores (latent variables) themselves. In particular, the task of performing inference on latent variables simultaneously across more than one cluster remains largely unexplored, e.g., we want to assess whether one cluster is statistically different from the others while controlling for Type I error rate.
In this talk, we develop a framework for simultaneous inference on factor scores in LVMs, focusing on methods for constructing simultaneous prediction regions (SPRs) based on the parametric bootstrap. We demonstrate that the proposed approach is asymptotically guaranteed to control for Type I error in a joint manner, while numerical studies comparing the proposed bootstrap SPRs to other, simpler approaches such as the Bonferroni correction and Monte Carlo simulation demonstrate its superior finite-sample performance. We apply the bootstrap SPR method to a dataset on ratings of different aspects of coffee quality, and show how it can offer differing conclusions to so-called pointwise prediction regions that do not account for simultaneity of the inference. This is a joint work with Zhining Wang, Emi Tanaka and Alan H Welsh.
16:50-17:30 Fast Covariance-free Spatio-temporal Modeling via Coarse-to-fine Learning
Daisuke Murakami (The Institute of Statistical Mathematics)
Scalable spatio-temporal modeling remains challenging because conventional methods rely on covariance models that tightly couple spatial representation and temporal inference, often leading to high computational cost. To address the difficulty, this study develops coarse-to-fine spatio-temporal modeling (CF-STM), a framework that extends coarse-to-fine spatial modeling to spatio-temporal settings. CF-STM represents latent spatial processes through multiscale locally weighted models, with temporal dependence modeled separately through state-space models defined at local centers. This covariance-free formulation achieves scalable computation while allowing temporal models to be flexibly specified without altering the spatial representation. Monte Carlo experiments demonstrate predictive performance comparable to alternative scalable space-time models at substantially lower computational cost. An application to long-term residential land price data in the Tokyo metropolitan area shows that CF-STM flexibly captures complex spatio-temporal patterns while enabling interpretable inference.
17:30-17:35 Closing Remark
Organizers
Daisuke Murakami (The Institute of Statistical Mathematics)
Shonosuke Sugasawa (Keio University)
Contact
Daisuke Murakami: dmuraka(at)ism.ac.jp
Acknowledgments
This symposium is supported by the Research Center for Risk Analysis and the Research Center for Data Assimilation at the Institute of Statistical Mathematics, as well as by JSPS KAKENHI Grant Numbers 24K00175 and 24H00546.