I am an Associate Professor of Statistics at North Carolina State University, where I have been working since Aug 2020. I received my Ph.D. in Statistics from the University of Illinois at Urbana-Champaign in Jul 2016. From Aug 2016 to Aug 2020, I was a tenure-track Assistant Professor of Statistics at Virginia Tech.
The primary theme of my research is developing formal inferential algorithms for network data and applying such algorithms to epidemiology, social sciences, and environmental health. I am also working on developing a statistical science of patient safety, focusing on adverse medical events due to human errors, medical devices, drug reactions, and radiation therapy. See my CV below for more on my work and background.
Email: ssengup2 ''at'' ncsu.edu
Google Scholar: https://scholar.google.com/citations?user=MXM2IiUAAAAJ
Upcoming poster presentation on uncertainty quantification for named entity recognition at the ORNL Core Universities AI Workshop (AI-CORE), Charlottesville, VA, September 14–15, 2026. I will be sharing our work with Matthew Singer and Karl Pazdernik on conformal prediction for natural language processing.
I am visiting the Department of Statistics at the University of Virginia on September 10, 2026, to give a colloquium talk, “HODOR: Causal experiments under unobserved network interference and lurking variables.” Our work develops experimental designs and inference methods for settings where treatments can affect other individuals through an unobserved network. Thanks to Jingming Wang and Zach Lubberts for the invitation!
Our paper, “Biogeochemical reactions in anaerobic digesters for biogas production yield enhanced ammonia emissions,” has been accepted for publication in Biogeochemistry. This is joint work with Viney Aneja, Swarnali Sanyal, and William Schlesinger.
Biogas recovery from animal waste can reduce methane emissions, but its broader environmental consequences require careful evaluation. Using field measurements, regression analysis, and a mass-transfer model, we find higher ammonia emissions from secondary lagoons receiving digestate from covered anaerobic digesters than from conventional lagoons. The findings highlight the importance of integrating nitrogen management into biogas systems to address potential trade-offs between renewable energy production and local environmental quality.
I had the pleasure of representing Statistics at NC State’s Research Showcase Lunch with NASA astronaut and NC State alumna Christina Koch on August 25, 2026. I shared a brief introduction to our research on discovering hidden structure in networks and quantifying uncertainty in AI systems. It was a wonderful opportunity to discuss how statistical methods help us understand the reliability of complex systems, alongside colleagues from the Colleges of Sciences and Engineering. See NC State’s coverage of her homecoming and the university’s video recap. The College of Sciences photo gallery includes our group photo and other moments from the luncheon.
I presented “HODOR: Causal experiments under unobserved network interference and lurking variables” in the session on Causal Inference with Interference at the 2026 Joint Statistical Meetings in Boston, August 2026.
I presented our work on uncertainty quantification for named entity recognition at the inaugural STAI-X conference, “Statistics and Trustworthy AI for Cross (X)-Domain Acceleration,” at Harvard University on August 1, 2026. This is joint work with Matthew Singer and Karl Pazdernik, using conformal prediction to construct prediction sets with statistical coverage guarantees for NLP tasks. Conference details.
The revised version of our paper, “Predictive Subsampling for Scalable Inference in Networks”, is now on arXiv. This is joint work with Arpan Kumar and Minh Tang.
We develop a framework for scalable estimation and two-sample testing in large networks. The central idea is to estimate the model on a small random subgraph and incorporate the remaining vertices through fast out-of-sample predictions. We establish finite-sample error bounds and consistency results that characterize the trade-offs between statistical accuracy and computational cost, and demonstrate the methods on simulated networks, DBLP coauthorship networks, and Cannes social media networks.
I participated in the Interdisciplinary Research Cluster on Causal Inference with Complex Spillovers at the Institute for Mathematical and Statistical Innovation (IMSI), University of Chicago, July 6–10, 2026. The visit provided an opportunity for sustained collaboration on causal inference when interventions have effects beyond the individuals directly receiving them.
New arXiv preprint on geopolitical alignment through covariate-assisted community detection, with Chirayata Kusari and Souvik Roy.
We develop a spectral clustering framework that combines heterogeneous network data with node-level covariates to identify community structure. The method comes with theoretical guarantees and is applied to United Nations General Assembly voting data, where voting interactions and auxiliary information together reveal patterns of geopolitical alignment.
Congratulations to Kaustav Chakraborty on completing his Ph.D. in Statistics at the University of Illinois Urbana-Champaign and joining Virginia Commonwealth University as a postdoctoral scholar! I had the pleasure of working with Kaustav and Yuguo Chen on his dissertation, “Subsampling-based Methods in Large Network Data Inference.”
I gave an invited talk on “Link prediction via Isotonic Regression” at the 7th International Symposium on Nonparametric Statistics (ISNPS 2026), Thessaloniki, Greece, June 22–26, 2026.
I am honored to have received the 2025 IISA Early Career Award in Statistics and Data Sciences in the Applications category. Many thanks to the International Indian Statistical Association for this recognition! Award announcement.
Our paper, “A label-switching algorithm for fast core-periphery identification”, with Eric Yanchenko, has been published in Network Science. We develop a fast greedy algorithm for identifying a densely connected core and a sparse periphery in networks. The method substantially reduces computational cost while delivering strong empirical performance.
I am serving on the Executive Council of the International Indian Statistical Association in 2026, including work on the association’s website and preparations for its annual conference.
I have been organizing a research roundtable in the NC State Department of Statistics, providing an interactive forum for faculty discussions about research collaboration, grant development, publishing, advising, and emerging research directions.
Upcoming talk on link prediction at the 7th International Symposium on Nonparametric Statistics (ISNPS 2026) at Thessaloniki, Greece, June 22–26, 2026.
I am visiting Japan from March 3rd to April 2nd as part of a US-Japan research collaboration project on statistical network science, jointly supported by NSF and JSPS. Thanks to Akita International University and Prof. Eric Yanchenko for hosting me. During my trip I will be visiting and giving talks at the following universities and institutions:
Akita International University
Radiation Effects Research Foundation (Hiroshima)
University of Tokyo
Nagoya University
Our paper, Scalable community detection in massive networks via predictive assignment, has been published in the Journal of the American Statistical Association (Theory and Methods). This is joint work with Subhankar Bhadra and Marianna Pensky.
Massive network datasets are becoming increasingly common in scientific applications. Existing community detection methods encounter significant computational challenges for such massive networks due to two reasons. First, the full network needs to be stored and analyzed on a single server, leading to high memory costs. Second, existing methods typically use matrix factorization or iterative optimization using the full network, resulting in high runtimes. We propose a strategy called predictive assignment to enable computationally efficient community detection while ensuring statistical accuracy. The core idea is to avoid large-scale matrix computations by breaking up the task into a smaller matrix computation plus a large number of vector computations that can be carried out in parallel. Under the proposed method, community detection is carried out on a small subgraph to estimate the relevant model parameters. Next, each remaining node is assigned to a community based on these estimates. We prove that predictive assignment achieves strong consistency under the stochastic blockmodel and its degree-corrected version. We also demonstrate the empirical performance of predictive assignment on simulated networks and two large real-world datasets: DBLP (Digital Bibliography & Library Project), a computer science bibliographical database, and the Twitch Gamers Social Network.
New AISTATS 2026 Paper on UQ for NLP: Our paper on Uncertainty Quantification for Named Entity Recognition has been accepted as a poster at the 29th International Conference on Artificial Intelligence and Statistics (AISTATS 2026). This is joint work with Matthew Singer and Karl Pazdernik and our first paper from this ongoing line of research on UQ for NLP tasks (and more broadly, ML models).
This paper develops a conformal-prediction framework for quantifying uncertainty in named entity recognition, providing statistically calibrated measures of confidence for entity-level predictions. The work was presented by Matt as a poster at AISTATS 2026 in Morocco.
Upcoming talk on clique detection at the Department of Statistics and Actuarial Science seminar series at the University of Iowa, Feb 2026, Iowa City, Iowa.
Upcoming talk on clique detection at the Statistics Department seminar series at North Carolina State University, Jan 2026, Raleigh, NC.
Upcoming talk on clique detection at the IMSI workshop on Recent Advances in Random Networks, Jan 2026, Chicago.
Our new preprint (with Matthew Singer and Karl Pazdernik) on uncertainty quantification for named entity recognition is now on arXiv: https://arxiv.org/abs/2601.16999
Named Entity Recognition (NER) serves as a foundational component in many natural language processing (NLP) pipelines. However, current NER models typically output a single predicted label sequence without any accompanying measure of uncertainty, leaving downstream applications vulnerable to cascading errors. In this paper, we introduce a general framework for adapting sequence-labeling-based NER models to produce uncertainty-aware prediction sets. These prediction sets are collections of full-sentence labelings that are guaranteed to contain the correct labeling with a user-specified confidence level. This approach serves a role analogous to confidence intervals in classical statistics by providing formal guarantees about the reliability of model predictions. Our method builds on conformal prediction, which offers finite-sample coverage guarantees under minimal assumptions. We design efficient nonconformity scoring functions to construct efficient, well-calibrated prediction sets that support both unconditional and class-conditional coverage. This framework accounts for heterogeneity across sentence length, language, entity type, and number of entities within a sentence. Empirical experiments on four NER models across three benchmark datasets demonstrate the broad applicability, validity, and efficiency of the proposed methods.
Upcoming talk on clique detection at the Joint Meetings of 2025 Taipei International Statistical Symposium and the 13th ICSA International Conference, Dec 2025, Taipei, Taiwan.
Upcoming talk on clique detection at the Banff workshop on large models, Dec 2025, Chennai, India.
Our paper, "A Unified Framework for Community Detection and Model Selection in Blockmodels" (w/ Subhankar Bhadra and Minh Tang), has been published in the Journal of Computational and Graphical Statistics.
Blockmodels are a foundational tool for modeling community structure in networks, with the stochastic blockmodel (SBM), degree-corrected blockmodel (DCBM), and popularity-adjusted blockmodel (PABM) forming a natural hierarchy of increasing generality. While community detection under these models has been extensively studied, much less attention has been paid to the model selection problem, i.e., determining which model best fits a given network. Building on recent theoretical insights about the spectral geometry of these models, we propose a unified framework for simultaneous community detection and model selection across the full blockmodel hierarchy. A key innovation is the use of loss functions that serve a dual role: they act as objective functions for community detection and as test statistics for hypothesis testing. We develop a greedy algorithm to minimize these loss functions and establish theoretical guarantees for exact label recovery and model selection consistency under each model. Extensive simulation studies demonstrate that our method achieves high accuracy in both tasks, outperforming or matching state-of-the-art alternatives. Applications to five real-world networks further illustrate the interpretability and practical utility of our approach.
Our paper, "A Bootstrap-based Method for Testing Network Similarity," (w/ Somnath Bhadra, Kaustav Chakraborty, and Soumendra Nath Lahiri) has been published in the Journal of Computational and Graphical Statistics.
In this work, we address the problem of determining whether two networks, defined on a common set of nodes, exhibit stochastic similarity. We introduce a bootstrap-based testing framework that assesses two notions of similarity: (i) equality—testing if the networks arise from the same random graph model, and (ii) scaling—testing if their probability matrices are proportional. The proposed method is versatile, accommodating various network models such as stochastic blockmodels, Chung-Lu models, and random dot product graph models. We establish the theoretical consistency of our tests and demonstrate their empirical performance through extensive simulations and a real-world application involving the Aarhus network dataset.
New arXiv pre-print on Network Cross-Validation and Model Selection via Subsampling (w/ Sayan Chakrabarty and Yuguo Chen). In this work, we introduce NETCROP (NETwork CRoss-Validation using Overlapping Partitions), a novel cross-validation procedure designed for complex and large-scale networks. NETCROP enhances computational efficiency by leveraging smaller, overlapping subnetworks for training, providing accurate model selection and parameter tuning. Our numerical results demonstrate that NETCROP often surpasses existing network cross-validation methods in both speed and accuracy.
Our work (w/ Kartik Lovekar and Subhadeep Paul) on small-world networks is now published in the Electronic Journal of Statistics. The “small-world” property—where networks have both high clustering and short paths between nodes—shows up across fields like sociology, biology, and neuroscience. But current ways of detecting small-world structure often fall short. Existing approaches mix clustering and path length into a single metric, lack statistical rigor, and rely on overly simple baseline models. In our work, we separate these two key features and define small-worldness as a formal hypothesis test. We introduce both parametric bootstrap and asymptotic tests (with theoretical guarantees) that work under flexible null models, including Erdős–Rényi. Applying these tools to real-world networks reveals a more accurate and nuanced view of the small-world phenomenon.
The revised version of our predictive assignment paper (w/ Subhankar Bhadra and Marianna Pensky) is now on arXiv. We propose a strategy called predictive assignment to scale up community detection in massive networks while ensuring statistical accuracy. First, community detection is carried out on a small subgraph to estimate the relevant model parameters. Next, each remaining node is assigned to a community based on these estimates. We prove that predictive assignment achieves strong consistency under the stochastic blockmodel and its degree-corrected version, even when the parent community detection algorithm is only weakly consistent.
Our work (w/ Indrila Ganguly and Sujit Ghosh) on subsampled residual bootstrap is now published in the Journal of Machine Learning Research. We propose a simple and versatile scalable algorithm called subsampled residual bootstrap (SRB) for generalized linear models (GLMs), a large class of regression models that includes the classical linear regression model as well as other widely used models such as logistic, Poisson and probit regression. We prove consistency and distributional results that establish that the SRB has the same theoretical guarantees under the GLM framework as the classical residual bootstrap, while being computationally much faster. We demonstrate the empirical performance of SRB via simulation studies and a real data analysis of the Forest Covertype data from the UCI Machine Learning Repository.
Our work (w/ Kaustav Chakraborty and Yuguo Chen) on scalable inference for RDPG networks is now published in the Journal of Computational and Graphical Statistics. In this article, we propose a subsampling-based method to reduce the computational cost of estimation and two-sample hypothesis testing. The idea is to divide the network into smaller subgraphs with an overlap region, then draw inference based on each subgraph, and finally combine the results together. We first develop the subsampling method for random dot product graph models, and establish theoretical consistency of the proposed method. Then we extend the subsampling method to a more general setup and establish similar theoretical properties. We demonstrate the performance of our methods through simulation experiments and real data analysis.
New NSF grant: Scalable and Generalizable Inference for Network Data. This is a single PI grant for methodological work on network inference.