publications and preprints
Publications and preprints in reversed chronological order.
Denotes joint/equal first authorship. Corresponding author.
2026
- Preprint
Rashomon-Seeded Annealing for Robust Bayesian Inference in Factorial DesignsYiyang Fan, Soumyakanti Pan, and Tyler H. McCormickarXiv preprint, 2026, SubmittedIntegrating over model uncertainty in factorial designs via Bayesian model averaging is hindered by the combinatorial explosion of interpretable interaction effects, often yielding a multimodal posterior, where standard Markov chain Monte Carlo algorithms encounter significant convergence issues. We propose a general computational framework that repurposes Rashomon sets, collections of high-performing models traditionally valued for prediction and interpretability, as a strategic "warm start" for estimating the full posterior. Our method, Rashomon-seeded annealing, initializes annealed importance sampling (AIS) by anchoring the starting density within these pre-identified, high-evidence regions while preserving global support over the entire model space. Rather than restricting inference to the Rashomon set and understating uncertainty, the AIS correction restores full posterior inference, turning the Rashomon certificate from an inferential truncation into a proposal mechanism. We demonstrate this approach using Rashomon Partition Sets (RPS) as a rigorous, certified seed constructor for factorial designs. The resulting algorithm yields consistent self-normalized posterior summaries, such as model-averaged cell means, credible intervals, and uncertainty summaries without exhaustive enumeration of the complete model space. This bridges the gap between high-evidence model discovery and rigorous Bayesian inference, and outlines a general strategy in which any high-posterior seed set can provide computational leverage for AIS-based model averaging.
@article{rhs_arxiv, author = {Fan∗, Yiyang and Pan∗, Soumyakanti and McCormick, Tyler H.}, title = {Rashomon-Seeded Annealing for Robust Bayesian Inference in Factorial Designs}, year = {2026}, journal = {arXiv preprint}, pubstate = {Submitted}, doi = {10.48550/arXiv.2606.02589}, show = {true} } - Preprint
A Multimodal Benchmark for Evaluating Cause-of-Death Inference Using Child Health and Mortality DataJunhe Yang, Soumyakanti Pan,, Hyun Seung Lim , 17 more authors , and Tyler H. McCormickmedRxiv preprint, 2026, SubmittedAccurately attributing causes of death is vital for global health, yet fewer than 5% of deaths in resource-constrained regions are medically certified. To assign causes to these unlabeled deaths at scale, practitioners traditionally rely on verbal autopsy, using supervised statistical models to classify based on structured survey data. However, modern mortality surveillance increasingly collects rich, unstructured multimodal data, such as free-text caregiver narratives and postmortem diagnostics, which traditional supervised statistical models struggle to seamlessly integrate. In this paper, we present a comprehensive, multimodal benchmark for cause-of-death classification using data from the Child Health and Mortality Prevention Surveillance (CHAMPS) network, a unique surveillance platform spanning nine countries across South Asia and Sub-Saharan Africa. Using this dataset, we introduce an evaluation framework designed to rigorously assess diagnostic reasoning, moving beyond traditional metrics that fail to capture complex clinical realities. We demonstrate the utility of this benchmark by evaluating zero-shot large language models against supervised baselines across various data modalities. Our results reveal distinct differences in how these modeling approaches synthesize unstructured medical evidence. This benchmark provide a rigorously defined resource for assessing clinical reasoning in next-generation mortality surveillance.
@article{coda_ff_medrxiv, author = {Yang∗, Junhe and Pan∗✉, Soumyakanti and Lim, Hyun Seung and Chu, Yue and Guo, Yuting and Agarwal, Nishtha and Babbar, Varun and Parikh, Gaurav Rajesh and Chen, Yiqun T. and Rees, Chris A. and Dangor, Ziyaad and Lala, Sanjay G. and Li, Zehang Richard and Clark, Samuel J. and Wu, Zhenke and Datta, Abhirup and Liu, Li and Rudin, Cynthia and Scarpino, Samuel V. and Gyori, Benjamin M. and McCormick, Tyler H.}, title = {A Multimodal Benchmark for Evaluating Cause-of-Death Inference Using Child Health and Mortality Data}, elocation-id = {2026.07.13.26357980}, journal = {medRxiv preprint}, year = {2026}, pubstate = {Submitted}, doi = {10.64898/2026.07.13.26357980}, show = {true}, } - PreprintHuman Dioxin Exposure Levels in Vietnam: A Systematic Database and Bayesian Areal KrigingSophie K. F. Michel, Soumyakanti Pan, Sudipto Banerjee , 2 more authors , and Ondine S. von Ehrenstein2026, SubmittedAccepted at International Society for Environmental Epidemiology Conference 2026, Munich, Bavaria, Germany.
Agent Orange and other herbicides used during the U.S. - Vietnam War were contaminated with the dioxin 2,3,7,8- tetrachlorodibenzo-p-dioxin (TCDD). The extreme toxicity, high environmental persistence, and strongly bioaccumulative properties of TCDD make exposure assessment in the Vietnamese population an urgent, currently unmet public health problem. We created a systematic database of all reports on TCDD measurements in human tissues and used a Bayesian areal kriging approach to assess human TCDD exposure in Vietnam. Tissue measurements from 8325 persons were identified in 58 key reports. Modeling results show increased exposure in residents of Central and South Vietnam, particularly at the Bien Hoa airbase in the Dong Nai province, and other areas affected by high-volume herbicide sprays or spills. Therefore, the ongoing environmental remediation of the Bien Hoa hotspot must be completed with high priority, while certain sprayed areas, especially lowland agricultural regions, have been neglected and should be further assessed for persistent TCDD contamination. The database, our area-level TCDD exposure estimates, and files for easily deriving additional estimates will be made publicly available, facilitating future research on TCDD exposure patterns in Vietnam, as well as further studies, such as epidemiological investigations into associated health effects.
- PreprintBirth Weight in Relation to Province-Level TCDD Contamination in Vietnam: An Epidemiological AnalysisSophie K. F. Michel, Soumyakanti Pan, Sudipto Banerjee , 3 more authors , and Ondine S. von Ehrenstein2026, SubmittedSubmitted to American Public Health Association Conference 2026, San Antonio, Texas, United States.
Background: The highly persistent and toxic dioxin congener TCDD has adverse reproductive, gestational, and transgenerational effects, as seen in animal studies. The widespread TCDD contamination in Vietnam makes epidemiologic investigations into effects on birth outcomes among the Vietnamese population urgently necessary. Methods: We used our newly developed Human TCDD Contamination in Vietnam Database and Bayesian areal kriging, to obtain province-level TCDD exposure estimates for Vietnam and linked these with birth weight data from 2,048 live births recorded as part of the 2006 and 2011 Multiple Indicator Cluster Surveys (MICS) in Vietnam. We fit Bayesian linear mean models, regressing individual birth weights modeled as a continuous outcome with a Gaussian likelihood, on continuous province-level TCDD exposure estimates, while adjusting for individual- and province-level covariates and assessing potential effect measure modification. Results: Bayesian linear regression modeling including an interaction term for urban vs. rural children showed a 69 g reduction in birthweight (95% CrI: -118 to -16 g) in rural children. Restricting the data to families in South Vietnam, a 97 g (95% CrI: -167 to -29 g) reduction in birth weight was predicted among rural children. Sensitivity analyses and model diagnostics indicated our results to be robust. Discussion: Analyses indicated reductions in birth weight with higher province-level TCDD exposure. Reductions were particularly pronounced among rural families, likely because rural living is associated with dietary practices (e.g., fishing) that would lead to higher exposure to TCDD in local foods. Public health programs (e.g., dietary behavior interventions) should be continued and expanded.
- Envcs
Bayesian Inference for Spatially-Temporally Misaligned Data Using Predictive StackingSoumyakanti Pan and Sudipto BanerjeeEnvironmetrics, 2026Air pollution remains a major environmental risk factor that is often associated with adverse health outcomes. However, quantifying and evaluating its effects on human health is challenging due to the complex nature of exposure data. This article develops a Bayesian hierarchical model to analyze spatially-temporally misaligned exposure and health outcome data. We introduce Bayesian predictive stacking, which optimally combines multiple predictive spatial-temporal models and avoids iterative estimation algorithms such as Markov chain Monte Carlo. We apply our proposed method to study the effects of ozone on asthma in the state of California.
@article{pan2026_envcs, title = {Bayesian Inference for Spatially-Temporally Misaligned Data Using Predictive Stacking}, author = {Pan, Soumyakanti and Banerjee, Sudipto}, journal = {Environmetrics}, volume = {37}, number = {2}, pages = {e70072}, year = {2026}, doi = {10.1002/env.70072}, show = {true}, }
2025
- BA
Bayesian Inference for Spatial-Temporal Non-Gaussian Data Using Predictive StackingSoumyakanti Pan, Lu Zhang, Jonathan R. Bradley, and Sudipto BanerjeeBayesian Analysis, 2025, In pressSelected as one of four papers for presentation at the "Selected Papers from Bayesian Analysis" session at ISBA 2026 World Meeting, Nagoya, Japan.Analysing non-Gaussian spatial-temporal data typically requires introducing spatial dependence in generalised linear models through the link function of an exponential family distribution. However, unlike in Gaussian likelihoods, inference is considerably encumbered by the inability to analytically integrate out the random effects and reduce the dimension of the parameter space. We devise an approach that obviates these issues by exploiting generalised conjugate multivariate distribution theory for exponential families, which enables exact sampling from analytically available posterior distributions conditional upon some fixed process parameters. We evaluate inferential performance on simulated data, compare with fully Bayesian inference using Markov chain Monte Carlo and apply our proposed method to analyse spatially-temporally referenced avian count data from the North American Breeding Bird Survey database.
@article{pan2025_stacking_ba, title = {Bayesian Inference for Spatial-Temporal Non-{G}aussian Data Using Predictive Stacking}, author = {Pan, Soumyakanti and Zhang, Lu and Bradley, Jonathan R. and Banerjee, Sudipto}, journal = {Bayesian Analysis}, publisher = {International Society for Bayesian Analysis}, year = {2025}, doi = {10.1214/25-BA1582}, pubstate = {In Press}, show = {true}, remark = {Selected as one of four papers for presentation at the "Selected Papers from Bayesian Analysis" session at ISBA 2026 World Meeting, Nagoya, Japan.}, highlight = {true}, }
2024
- Preprint
spStack: Practical Bayesian Geostatistics Using Predictive Stacking in RSoumyakanti Pan and Sudipto Banerjee2024Statistical modeling and analysis for spatially oriented point-referenced outcomes play a crucial role in diverse scientific applications such as earth and environmental sciences, ecology, epidemiology, and economics. In this article, we introduce the R package spStack that delivers fast Bayesian inference for a class of geostatistical models, where we obviate weak identifiability issues in spatial process parameters by sampling from analytically available posterior distributions conditional upon candidate values of the spatial process parameters and, subsequently, assimilate inference from these individual posterior distributions using Bayesian predictive stacking. Our proposed algorithm is executable in parallel, drastically improving runtime over traditional MCMC.
@misc{pan2024spstack, title = {spStack: Practical {B}ayesian Geostatistics Using Predictive Stacking in {R}}, author = {Pan, Soumyakanti and Banerjee, Sudipto}, year = {2024}, cran = {https://CRAN.R-project.org/package=spStack}, show = {true} } - AWEH
Bayesian Hierarchical Modeling and Inference for Mechanistic Systems in Industrial HygieneSoumyakanti Pan, Darpan Das, Gurumurthy Ramachandran, and Sudipto BanerjeeAnnals of Work Exposures and Health, 2024A series of experiments in stationary and moving passenger rail cars were conducted to measure removal rates of particles in the size ranges of SARS-CoV-2 viral aerosols and the air changes per hour provided by existing and modified air handling systems. Such methods for exposure assessments are customarily based on mechanistic models derived from physical laws of particle movement that are deterministic and do not account for measurement errors inherent in data collection. The resulting analysis compromises on reliably learning about mechanistic factors such as ventilation rates, aerosol generation rates, and filtration efficiencies from field measurements. This manuscript develops a Bayesian state-space modeling framework that synthesizes information from the mechanistic system as well as the field data by deriving a stochastic model from finite difference approximations of differential equations explaining particle concentrations. Our inferential framework trains the mechanistic system using the field measurements from the chamber experiments and delivers reliable estimates of the underlying physical process with fully model-based uncertainty quantification.
@article{pan2024mechih, title = {Bayesian Hierarchical Modeling and Inference for Mechanistic Systems in Industrial Hygiene}, author = {Pan, Soumyakanti and Das, Darpan and Ramachandran, Gurumurthy and Banerjee, Sudipto}, journal = {Annals of Work Exposures and Health}, volume = {68}, number = {8}, pages = {834--845}, year = {2024}, doi = {10.1093/annweh/wxae061}, show = {true}, }