PubMed HealthSearch

SEARCH · PubMed Health

Results for “Uncertainty Quantification”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

16 recordsLinked to original sources

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem

Estimating protein isoform abundances with [Formula: see text].

A single gene can encode multiple versions of a protein, dubbed isoforms, with varying functionality. Cellular control of isoform abundances is critical for multiple aspects of biology and is only partially regulated by transcript levels. While long-read sequencing facilitates transcript quantification, quantifying the resulting protein isoforms on a large scale is a major challenge, complicating biological interpretation of transcript alterations. Standard "bottom up" mass spectrometry can assess only short portions of isoforms called peptides, and these peptides often map onto more than one isoform. We introduce [Formula: see text] (Protein isoform Abundance Quantification), a Bayesian method that leverages multiomic information from the peptidome and transcriptome to provide accurate estimates of isoform abundance even when peptide mapping is ambiguous. [Formula: see text] offers several advantages over existing methods in a unified framework. It provides uncertainty quantification, integrates multiomic information for improved accuracy, and provides a rigorous framework for hypothesis testing. Extensive simulations show that [Formula: see text] consistently outperforms competing methods in detecting differentially abundant protein isoforms and estimating their abundances. We use [Formula: see text] to investigate differences in isoform abundance levels between people with schizophrenia and control subjects, confirming a long-held hypothesis that levels of the C4A isoform of Complement Component 4 are increased in schizophrenia while C4B is not. These results demonstrate that [Formula: see text] can identify significant variations in isoform abundance levels not previously possible.

Protein Isoforms

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

Improving the reliability of polygenic risk score-based prediction for cardiovascular and renal complications across ancestries in type 2 diabetes using Mondrian Cross-Conformal Prediction.

Polygenic risk scores (PRS) developed in European populations often show reduced predictive performance in non-European populations, limiting their clinical utility. This lack of transferability across ancestries remains a major challenge in genomic medicine and raises concerns about health equity. We aimed to evaluate whether uncertainty-aware prediction, implemented through Mondrian Cross-Conformal Prediction, improves the performance and reliability of polygenic risk score-based predictions across ancestries for nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes in a multi-ethnic cohort. We leveraged Mondrian Cross-Conformal Prediction (MCCP), an uncertainty quantification framework, combined with logistic regression applied to a multi-polygenic risk score (multiPRS) to predict the risk of nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes. Two training frameworks were evaluated: one using 4,098 individuals with type 2 diabetes of European ancestry from the ADVANCE trial for training and 17,574 White British, 1,145 South Asian, and 749 African UK Biobank participants for testing; and another using the 17,574 White British UK Biobank participants for training and the South Asian and African participants for testing. Logistic regression provided robust baseline performance across populations. On top of this baseline, MCCP did not improve performance but added capabilities absent from probability-based stratification: for each individual, it issued a prediction together with an explicit confidence and credibility level; it allowed a tolerated error level to be set in advance and delivered prediction sets respecting it in the majority of settings; and it flagged individuals for whom no reliable prediction could be made. Applying MCCP to PRS-based prediction thus enables uncertainty-aware risk stratification and improves the reliability of risk prediction across ancestries, providing a more equitable framework for clinical use.

Female

Measuring economic efficiency in adult intensive care units: A systematic review of methods, metrics, and evidence.

OBJECTIVES: Intensive care units (ICUs) consume substantial hospital resources, yet "efficiency" is inconsistently defined and measured. This study systematically reviewed how economic efficiency has been conceptualised and quantified in adult ICUs and appraised the quality of evidence. METHODS: Following PRISMA 2020 and a PROSPERO-registered protocol (CRD420251107866), we searched MEDLINE, Embase, CINAHL, Cochrane Library and Web of Science (2000-August 2025), plus global grey sources. Eligible studies explicitly defined efficiency and reported an efficiency metric/model linking ICU inputs (e.g., staff, beds/capacity, time, consumables, or costs) to outputs/outcomes (e.g., throughput/discharges, length of stay/resource use, risk-adjusted mortality). Dual independent screening and extraction were performed. Study quality was appraised using MMAT, and findings were synthesised narratively (SWiM), given heterogeneity. RESULTS: 39 studies (2001-2025) from 17 countries were included, all from high-income or upper-middle-income settings. Four methodological families were identified: (1) frontier modelling (predominantly DEA; occasional SFA/RFDH), (2) benchmarking indicators (risk-adjusted mortality and LOS/resource-use ratios; "efficiency matrix" quadrant classification), (3) cost-outcome evaluations, and (4) operational/process metrics. Across families, variation in decision-making units, input/output selection, and risk adjustment limited comparability; long-term and patient-reported outcomes were absent, and equity considerations were uncommon. CONCLUSIONS: ICU efficiency research is feasible but fragmented and often methodologically limited. Standardised definitions, validated risk adjustment, uncertainty quantification, and inclusion of patient-centred and equity-relevant outcomes are needed before efficiency metrics can reliably inform value-based decision making.

Intensive Care Units

Robust and accurate Bayesian inference of genome-wide genealogies for hundreds of genomes.

The Ancestral Recombination Graph (ARG), which describes the genealogical history of a sample of genomes, is a vital tool in population genomics and biomedical research. Recent advancements have substantially increased ARG reconstruction scalability, but they rely on approximations that can reduce accuracy, especially under model misspecification. Moreover, they reconstruct only a single ARG topology and cannot quantify the considerable uncertainty associated with ARG inferences. Here, to address these challenges, we introduce SINGER (sampling and inferring of genealogies with recombination), a method that accelerates ARG sampling from the posterior distribution by two orders of magnitude, enabling accurate inference and uncertainty quantification for hundreds of whole-genome sequences. Through extensive simulations, we demonstrate SINGER's enhanced accuracy and robustness to model misspecification compared to existing methods. We demonstrate the utility of SINGER by applying it to individuals of British and African descent within the 1000 Genomes Project, identifying signals of population differentiation, archaic introgression and strong support for ancient polymorphism in the human leukocyte antigen region shared across primates.

Humans

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics

Calibrated Prediction Intervals for Polygenic Scores: Updated Comparisons, Contextual Calibration, and Data Normalization.

Calibrated prediction intervals for polygenic scores (PGS) are essential for communicating individual-level uncertainty in genomic medicine. We present updated comparisons of two methods for constructing such intervals: CalPred, a parametric approach, and PredInterval, a non-parametric approach. Our results show that both methods can achieve calibrated coverage, although CalPred additionally requires a sufficiently large calibration set. The two methods also exhibit complementary trade-offs with respect to dataset size and risk identification. We further show that contextual calibration, as introduced in Hou et al. and followed in Shi et al., is most naturally achieved through appropriate phenotype normalization and data preprocessing. Apparent miscalibration can arise from inadequate normalization or from providing contextual information to some methods but not others. In UK Biobank, standard GWAS phenotype normalization procedures are sufficient to achieve contextual calibration for traits analyzed. In the extreme simulations of Hou et al. and Shi et al., supplying contextual covariates to PredInterval restores contextual calibration without normalization, and appropriate normalization can achieve contextual calibration without supplying covariates, while also substantially improving upstream tasks including association power and PGS accuracy. Together, these results underscore the central role of phenotype normalization and data preprocessing in GWAS analyses, including reliable uncertainty quantification for PGS.

Journal Article

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Development and validation of a liquid chromatography-tandem mass spectrometry method for the quantification of twenty-five steroids in equine serum.

Steroids are potential biomarkers for monitoring equine pregnancy. However, immunoassays currently used for their quantification suffer from cross-reactivity and limited specificity, thus requiring more accurate methods. This study reports the development and validation of a robust liquid chromatography-tandem mass spectrometry (LC-MS/MS) method for simultaneous quantification of 25 steroids covering the main biosynthetic pathways of progestogens, corticosteroids, androgens, and estrogens. Steroids were extracted by protein precipitation followed by evaporation, derivatization, and reconstitution before LC-MS/MS analysis. A surrogate matrix was used for calibration and validation to avoid endogenous interference. Validation was performed according to and partly adapted from Clinical and Laboratory Standards Institute guidelines (CLSI), including linearity, trueness, precision, limits of detection and quantification, measurement uncertainty, recovery, matrix effects, carryover, selectivity, and stability. Calibration curves were fitted using the best-performing weighted linear or quadratic regression model, yielding excellent linearity (R2&#xa0;>&#xa0;0.990), trueness between -9.0% and 2.3%, and intra- and inter-day precision <6.3%. Lower limits of quantification ranged from 2.07 to 2250&#xa0;pg/mL depending on physiological analytes concentration. Extraction recovery averaged 24.3-114.9%, matrix effects were acceptable, and accuracy ranged from 94.4% to 98.9%. No carryover or interferences were detected. Measurement uncertainty remained <15%. This study presents the first LC-MS/MS method partially validated per CLSI criteria for the quantification of 24 steroids in equine serum. The method offers a sensitive and specific alternative to immunoassays and provides a robust tool for equine steroid profiling with potential applications in pregnancy monitoring, placentitis diagnosis, and fetal sex determination.

Animals

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq

An integrated multiscale air quality modelling framework for industrial park pollution: Linking local emissions to regional transport.

Capturing the spatiotemporal distribution of pollutants in industrial parks remains challenging for regional air quality models because of their coarse resolution (3 km), resulting in uncertainties in local emission quantification. To address this, we developed the Integrated Multiscale Air Quality Modelling System for Industry (IAQMS-Industry), coupling the regional Nested Air Quality Prediction Modelling System (NAQPMS) with a city-scale chemical transport model. This framework integrates point-source locations and Gaussian plume dispersion to simulate particulate matter with a diameter smaller than 2.5 micrometres (PM2.5) at 100 m resolution. Applied to the Beijing Yi Zhuang and Tangshan industrial parks and evaluated against observations. The coupled model achieved a normalized mean bias (NMB) ranging from 3.1 % to 6.2 %, improving upon NAQPMS (-16.9 % to -7.7 %). Spatial analysis revealed that coarse regional grids underestimated the PM2.5&#x200b; concentrations at industrial sites by smoothing gradients, whereas IAQMS-Industry successfully resolved spatial patterns. Industrial point emissions accounted for 22.9 %-26.4 % of PM2.5 in the coupled model, which was significantly greater than the regional model estimates of 1.6 %-13.7 %. These findings indicate that regional models overestimate pollutant dispersion processes in industrial parks while underestimating local industrial impacts. By explicitly resolving point-source dynamics and linking them to regional transport, IAQMS-Industry provides a robust tool for designing targeted emission controls in industrial cities and balancing local air quality improvements with minimized regional pollution outflow. This study underscores the necessity of multiscale modelling for accurate source apportionment and informed environmental governance in industrial zones.

Air Pollution

Global disparities in COVID-19 vaccine coverage associated with trajectories of SARS-CoV-2 adaptation.

BACKGROUND: Vaccination serves as an effective intervention for health promotion and disease prevention across the socioecological systems and has played an important role during the COVID-19 pandemic. However, global disparities in vaccine coverage have increased uncertainty about the trajectories of viral adaptation, and the potential interplay between SARS-CoV-2 adaptation and vaccine rollout warrants further quantification. METHODS: Using over 13&#xa0;million SARS-CoV-2 genomes across 86 countries from March 2020 to September 2022, we analyzed nonlinear associations between SARS-CoV-2 adaptation and vaccination coverage, considering public health and social measures, international travel, and infection dynamics, before and after the emergence of Omicron. Additionally, we examined the relationship between SARS-CoV-2 adaptation and COVID-19 mortality. RESULTS: During the pre-Omicron period, we found positive associations between nonsynonymous to synonymous divergence (dN/dS) ratios in the S1 subunit and medium levels of adjusted vaccine coverage (effect size: 0.96 [95% CI 0.47, 1.45]), while the association became insignificant at high levels (effect size: -1.89 [95% CI -4.20, 0.43]). However, no significant associations were found when Omicron dominated, possibly due to the immune escape ability of Omicron variants and the complex immune landscape shaped by mass hybrid immunity. Moreover, we observed evidence of dynamic interdependence and positive correlations between COVID-19 mortality and SARS-CoV-2 adaptation, with COVID-19 mortality interpreted as a proxy for uncontrolled viral spread. CONCLUSIONS: Our findings suggest a complex nonlinear relationship between vaccine-induced immunity and SARS-CoV-2 adaptation, with high vaccine coverage potentially linked to lower positive selection. We also observed directional coupling between COVID-19 mortality and SARS-CoV-2 adaptation. This may have implications for fair and fast vaccination in pandemic preparedness and response. CLINICAL TRIAL NUMBER: Not applicable.

Humans

Bayesian estimation of allele-specific expression in the presence of phasing uncertainty.

MOTIVATION: Allele-specific expression (ASE) analyses aim to detect imbalanced expression of maternal versus paternal copies of an autosomal gene. Such allelic imbalance can result from a variety of cis-acting causes, including disruptive mutations within one copy of a gene that impact the stability of transcripts, as well as regulatory variants outside the gene that impact transcription initiation. Current methods for ASE estimation suffer from a number of shortcomings, such as relying on only one variant within a gene, assuming perfect phasing information across multiple variants within a gene, or failing to account for alignment biases and possible genotyping errors. RESULTS: We developed BEASTIE, a Bayesian hierarchical model designed for precise ASE quantification at the gene level, based on given genotypes and RNA-Seq data. BEASTIE addresses the complexities of allelic mapping bias, genotyping error, and phasing errors by incorporating empirical phasing error rates derived from Genome-in-a-Bottle individual NA12878. BEASTIE surpasses existing methods in accuracy, especially in scenarios with high phasing errors. This improvement is critical for identifying rare genetic variants often obscured by such errors. Through rigorous validation on simulated data and application to real data from the 1000 Genomes Project, we establish the robustness of BEASTIE. These findings underscore the value of BEASTIE in revealing patterns of ASE across gene sets and pathways. AVAILABILITY AND IMPLEMENTATION: The software is freely available from Github (https://github.com/x811zou/BEASTIE); and Zendo (DOI: 10.5281/zenodo.15062124).

Bayes Theorem

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Quantification of triglyceride transport in blood plasma: a critical analysis.

Reliable and precise quantification of endogenous triglyceride transport in man has not been possible with simple means to date. Direct measurement of net splanchnic secretion of triglyceride fatty acids (TGFA) in very low density lipoproteins (VLDL) provides the must unambiguous information, but precision is low. Coupling infusion of labeled fatty acid with sampling of arterial and hepatic venous blood increases precision; however, the contribution of precursors other than plasma free fatty acids (FFA) must be assessed. Measurement of the rate of hydrolysis of plasma triglycerides after displacing lipases into the blood with heparin holds promise as a simple, nonisotopic method, but it has not been carefully validated and heparin itself alters FFA and triglyceride transport. Multicompartmental analysis following pulse injection of labeled fatty acid offers a practical approach, but uncertainties about the number and location of interacting compartments have made it impossible to determine an absolute value for transport. Reinjection of biologically labeled plasma VLDL is impractical for large scale use, and validity of this approach remains uncertain because of heterogeneity of VLDL-triglycerides and their complex metabolic behavior. Methods to label VLDL-triglycerides in vitro deserve more study as does labeling of other components, such as the B-apoprotein. Such approaches will require rigorous comparison with biologically labeled material as well as careful assessment of alterations in kinetic behavior that may occur when VLDL are separated from blood plasma.

Adult