Subha Guha develops statistical methods for integrating complex biomedical data across studies and data sources. His methodological research focuses on causal inference, transfer learning, high-dimensional inference, Bayesian nonparametrics, hierarchical modeling, and scalable computation. A central goal is to produce efficient and generalizable inference when observational studies differ in their populations, covariate distributions, treatment patterns, or observed outcomes.
These methods are motivated by collaborative problems in cancer genomics, precision oncology, microbiome and multi-omics research, healthcare delivery, and population science. Current projects include causal meta-analysis across observational studies, weighting methods for externally specified target populations, Bayesian analysis of multimodal single-cell data, and regionalization of healthcare referral networks.
Causal inference and multi-study evidence integration
Transfer learning and generalizable statistical inference
Bayesian modeling and computational methods
High-dimensional and multimodal biomedical data integration
Cancer genomics and precision oncology
Microbiome and multi-omics data analysis
Electronic health records and complex observational studies
Healthcare delivery and population science
Statistical software and reproducible biomedical research
This research develops methods for combining evidence from observational studies that differ in population composition and data structure. Current work addresses multivariate outcomes, high-dimensional confounding, target-population specification, and weighting strategies that preserve statistical information while supporting scientifically meaningful comparisons.
This research develops interpretable Bayesian methods for high-dimensional and multimodal biomedical data. Current directions include integrative analysis of cancer genomics, single-cell RNA and chromatin-accessibility data, microbiome data, biomedical images, survival outcomes, and brain connectivity.
This research uses patient-referral flows and geographic information to study healthcare markets and patterns of care. Ongoing methodological work develops probabilistic regionalization approaches that quantify uncertainty and characterize fragmentation and patient leakage. These approaches also measure disagreement between data-derived regions and conventional healthcare referral regions.
Collaborative research applies statistical methodology to cancer treatment, musculoskeletal disease, microbiome studies, electronic health records, and healthcare delivery. These projects motivate new methods and support reproducible analyses of complex observational data.
Guha, S., and Li, Y. (2026). Bayesian estimation of propensity scores for integrating multiple cohorts with high-dimensional covariates. Statistics in Biosciences, 18, 46-67.
Gu, C., Baladandayuthapani, V., and Guha, S. (2025). Nonparametric Bayes differential analysis of multigroup DNA methylation data. Bayesian Analysis, 20, 489-518.
Guha, S., and Qiu, P. (2025). Bayesian pairwise comparison of high-dimensional images. Journal of Computational and Graphical Statistics, 34(4), 1356-1365.
Guha, S., and Li, Y. (2024). Causal meta-analysis by integrating multiple observational studies with multivariate outcomes. Biometrics, 80, ujae070.
Song, J., Guha, S., and Li, Y. (2023). Bayesian inference for high-dimensional Cox models with Gaussian and diffused-gamma priors: A case study of mortality in COVID-19 patients admitted to the ICU. Statistics in Biosciences, 1-29.
Guha, S., Jung, R., and Dunson, D. (2022). Predicting phenotypes from brain connection structure. Journal of the Royal Statistical Society: Series C, 71, 639-668.
Sachdeva, A., Ahn, C., Tiwari, R., and Guha, S. (2022). A novel approach to augment single-arm clinical studies with real-world data. Journal of Biopharmaceutical Statistics, 27, 1-17.
Guha, S., and Datta, S. (2021). A Bayesian approach to restoring the duality between principal components of a distance matrix and operational taxonomic units in microbiome analyses. In Statistical Analysis of Microbiome Data (S. Datta and S. Guha, eds.). Springer Nature.
Guha, S., and Ghosh, S. K. (2020). Probabilistic detection and estimation of conic sections from noisy data. Journal of Computational and Graphical Statistics, 29, 513-522.
Yan, D., Guha, S., Ahn, C., and Tiwari, R. (2020). Semiparametric Bayesian Markov analysis of personalized benefit-risk assessment. Annals of Applied Statistics, 14, 768-788.
Guha, S., and Baladandayuthapani, V. (2016). A nonparametric Bayesian technique for high-dimensional regression. Electronic Journal of Statistics, 10, 3374-3424.
Guha, S. (2010). Posterior simulation in countable mixture models for large datasets. Journal of the American Statistical Association, 105, 775-786.
A complete publication record is available through Google Scholar and the curriculum vitae.
National Institute of Arthritis and Musculoskeletal and Skin Diseases. Biomechanics Contributions to Symptoms and Joint Health in Individuals with Rotator Cuff Tears. R01AR084273, 2024-2028. Role: Co-Investigator.
National Institute of Arthritis and Musculoskeletal and Skin Diseases. Nervous System Influences on Recovery from Painful Rotator Cuff Tears. R01AR080058, 2023-2028. Role: Co-Investigator.
National Cancer Institute. Causal Machine Learning in Cancer Survival by Integrating Multiple High-Dimensional Observational Studies. R01CA269398, 2022-2026. Role: Multiple Principal Investigator (MPI).
National Cancer Institute. The Boston Lung Cancer Survivor Cohort. U01CA209414, 2023-2028. Role: Co-Investigator.
Health Resources and Services Administration. Improving Sexually Transmitted Infection Screening and Treatment among People Living with or at Risk for HIV. U90HA32147, 2020-2023. Role: Co-Investigator.
National Science Foundation. Collaborative Research: New Bayesian Nonparametric Paradigms of Personalized Medicine for Lung Cancer. DMS-1854003, 2015-2020. Role: Principal Investigator (PI).
National Science Foundation. Bayesian Mixture Models: Unified Theoretical Frameworks and MCMC Methods. DMS-0906734, 2009-2013. Role: Principal Investigator (PI).
Department of Health and Human Services. Statistical Informatics for Cancer Research. 2009-2013. Role: Co-Investigator.
WMAP: An R package for weighting and causal meta-analysis across multiple observational studies. Version 1.3.0 includes diagnostic tools and cross-fitted random forest estimation. The associated development repository is available on GitHub.
Integrative covariate-balancing weighting: R code implementing the weighting strategy developed for causal meta-analysis across multiple observational studies.
B-MSC: R code for Bayesian propensity-score integration across multiple cohorts with high-dimensional covariates.
sRPM: R code for Bayesian pairwise comparison of high-dimensional images.
BayesDiff: R code for nonparametric Bayesian differential analysis of multigroup DNA methylation data.
BaCon: R code for Bayesian prediction of phenotypes from brain connectivity data.
Microbiome-SVD: R code for Bayesian dimension reduction and ordination of microbiome data.
BayesConics: R code for probabilistic detection and estimation of unknown conic sections from noisy data.
NPCluster: An R package implementing VariScan, a Bayesian method for clustering, variable selection, and prediction in high-dimensional regression.
hmmSeq: An R package for Bayesian differential-expression analysis of RNA sequencing data.
glmmGS: An R package for fitting generalized linear mixed models to large datasets.