PubMed Health⌕ Search

Biomedical subjects

Peter X K Song

Publications and source records attributed to Peter X K Song.

2 recordsLinked to original sources

Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data.

OBJECTIVE: To develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). MATERIALS AND METHODS: Large-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI ( n = 50,935) and AOU ( n = 241,563) to estimate smoking-cancer associations. RESULTS: In simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. CONCLUSION: Our framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.

Journal Article↗

Bayesian hierarchical models for multi-level repeated ordinal data using WinBUGS.

Multi-level repeated ordinal data arise if ordinal outcomes are measured repeatedly in subclusters of a cluster or on subunits of an experimental unit. If both the regression coefficients and the correlation parameters are of interest, the Bayesian hierarchical models have proved to be a powerful tool for analysis with computation being performed by Markov Chain Monte Carlo (MCMC) methods. The hierarchical models extend the random effects models by including a (usually flat) prior on the regression coefficients and parameters in the distribution of the random effects. Because the MCMC can be implemented by the widely available BUGS or WinBUGS software packages, the computation burden of MCMC has been alleviated. However, thoughtfulness is essential in order to use this software effectively to analyze such data with complex structures. For example, we may have to reparameterize the model and standardize the covariates to accelerate the convergence of the MCMC, and then carefully monitor the convergence of the Markov chain. This article aims at resolving these issues in the application of the WinBUGS through the analysis of a real multi-level ordinal data. In addition, we extend the hierarchical model to include a wider class of distributions for the random effects. We propose to use the deviance information criterion (DIC) for model selection. We show that the WinBUGS software can readily implement such extensions and the DIC criterion.

Bayes Theorem↗