PubMed HealthSearch

SEARCH · PubMed Health

Results for “Uncertainty modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

Inferring the demographic history of Chinese and Indian rhesus macaque (Macaca mulatta) populations from PacBio HiFi long-read sequencing data.

The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time ∼140,000 generations ago from an ancestral population of ∼65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of ∼220,000 individuals in the Chinese population and ∼14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

Cercopithecidae

Active learning of enhancers and silencers in the developing neural retina.

Deep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.

Retina

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar

Tomato bacterial wilt disease outbreaks are accompanied by an increase in soil antibiotic resistance.

The presence of soil-borne disease obstacles and antibiotic resistance genes (ARGs) in soil leads to serious economic losses and health risks to humans. One area in need of attention is the evolution of ARGs as pathogenic soil gradually develops, which introduces uncertainty to the dynamic ability of conventional farming models to predict ARGs. Here, we investigated variations in tomato bacterial wilt disease accompanied by the resistome by metagenomic analysis in soils over 13 seasons of monoculture. The results showed that the abundance and diversity of ARGs and mobile genetic elements (MGEs) exhibited a significant and positive correlation with R. solanacearum. Furthermore, the binning approach indicated that fluoroquinolone (qepA), tetracycline (tetA), multidrug resistance genes (MDR, mdtA, acrB, mexB, mexE), and β-lactamases (ampC, blaGOB) carried by the pathogen itself were responsible for the increase in overall soil ARGs. The relationships between pathogens and related ARGs that might underlie the breakdown of soil ARGs were further studied in R. solanacearum invasion pot experiments. This study revealed the dynamics of soil ARGs as soil-borne diseases develop, indicating that these ecological trends can be anticipated. Overall, this study enhances our understanding of the factors driving ARGs in disease-causing soils.

Soil Microbiology

Robust and accurate Bayesian inference of genome-wide genealogies for hundreds of genomes.

The Ancestral Recombination Graph (ARG), which describes the genealogical history of a sample of genomes, is a vital tool in population genomics and biomedical research. Recent advancements have substantially increased ARG reconstruction scalability, but they rely on approximations that can reduce accuracy, especially under model misspecification. Moreover, they reconstruct only a single ARG topology and cannot quantify the considerable uncertainty associated with ARG inferences. Here, to address these challenges, we introduce SINGER (sampling and inferring of genealogies with recombination), a method that accelerates ARG sampling from the posterior distribution by two orders of magnitude, enabling accurate inference and uncertainty quantification for hundreds of whole-genome sequences. Through extensive simulations, we demonstrate SINGER's enhanced accuracy and robustness to model misspecification compared to existing methods. We demonstrate the utility of SINGER by applying it to individuals of British and African descent within the 1000 Genomes Project, identifying signals of population differentiation, archaic introgression and strong support for ancient polymorphism in the human leukocyte antigen region shared across primates.

Humans

Toward Class Imbalance and Uncertainty in Powder XRD Analysis: A Dual-Channel Fusion Network for Space Group Classification.

Accurate identification of space groups from powder X-ray diffraction (pXRD) is essential for understanding crystal structures and accelerating materials discovery. However, this task remains highly challenging due to inherent peak overlap, experimental noise, and the complexity of the 230-class classification problem. To address the critical issues of class imbalance and data scarcity, we first design a general physics-informed data augmentation pipeline. We then propose a dual-channel fusion uncertainty-aware network (DFUN) for automated space group classification. The DFUN architecture integrates two complementary feature representations: convolutional features extracted directly from raw diffraction profiles and domain-specific peak descriptors. These distinct representations are adaptively fused through a gating mechanism. Furthermore, to mitigate the inherent long-tailed distribution of crystallographic data, we employ a hybrid loss function that combines Focal Loss with Label Smoothing. Finally, we incorporate Monte Carlo Dropout to provide predictive uncertainty estimation, thereby enabling not only accurate classification but also a crucial assessment of the model's reliability. Evaluated on large-scale simulated data and two public data sets (opXRD and RRUFF), DFUN outperforms the evaluated baseline methods across the reported metrics. The framework also provides uncertainty-aware predictions, establishing DFUN as a robust and interpretable solution for high-throughput automated crystallographic analysis from powder diffraction.

Uncertainty

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R²=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans

The use of simulation modelling in the management of brucellosis eradication.

Problems faced by Government veterinarians in the planning and implemention of the national bovine brucellosis eradication campaign are described. These problems stem from the uncertainty associated with the epidemiology of the disease and its initial status in eradication areas. They are compounded by contraints on abattoir capacity, finance and time. A computer stimulation model is discussed which is designed to assist campaign organisers to cope with these problems. Inputs related to expected test and slaughter rates, district disease prevalence and proposed campaign intensity are entered into the model. Output from the model inlcudes distributions giving predicted testing workloads, cattle slaughtered and disease status over the course of the campaign. The operation of the model is illustrated using data from the Bangalow district of New South Wales.

Animals

Genomic and Developmental Models to Predict Cognitive and Adaptive Outcomes in Autistic Children.

IMPORTANCE: Although early signs of autism are often observed between 18 and 36 months of age, there is considerable uncertainty regarding future development. Clinicians lack predictive tools to identify those who will later be diagnosed with co-occurring intellectual disability (ID). OBJECTIVE: To predict ID in children diagnosed with autism. DESIGN, SETTING, AND PARTICIPANTS: This prognostic study involved the development and validation of models integrating genetic variants and developmental milestones to predict ID. Models were trained, cross-validated, and tested for generalizability across 3 autism cohorts: Simons Foundation Powering Autism Research (SPARK), Simons Simplex Collection, and MSSNG. Autistic participants were assessed older than 6 years of age for ID. Study data were analyzed from January 2023 to July 2024. EXPOSURES: Ages at attaining early developmental milestones, occurrence of language regression, polygenic scores for cognitive ability and autism, rare copy number variants, de novo loss-of-function and missense variants impacting constrained genes. MAIN OUTCOMES AND MEASURES: The out-of-sample performance of predictive models was assessed using the area under the receiver operating characteristic curve (AUROC), positive predictive values (PPVs), and negative predictive values (NPVs). RESULTS: A total of 5633 autistic participants (4574 male [81.2%]) were included in this analysis. On average, participants were diagnosed with autism at 4 (IQR, 3-7) years of age and assessed for ID at 11 (8-14) years of age, with 1159 participants (20.6%) being diagnosed with ID. The model integrating all predictors yielded an AUROC of 0.653 (95% CI, 0.625-0.681), and this predictive performance was cross-validated and generalized across cohorts. This modest performance reflected that only a subset of individuals carried large-effect variants, high polygenic scores, or presented delayed milestones. However, combinations of genetic variants that are typically not considered clinically relevant by diagnostic laboratories achieved PPVs of 55% and correctly identified 10% of individuals developing ID. The addition of polygenic scores to developmental milestones specifically improved NPVs rather than PPVs. Notably, the ability to stratify ID probabilities using genetic variants was up to 2-fold higher in individuals with delayed milestones compared with those with typical development. CONCLUSIONS AND RELEVANCE: Results of this prognostic study suggest that the growing number of neurodevelopmental condition-associated variants cannot, in most cases, be used alone for predicting ID. However, models combining different classes of variants with developmental milestones provide clinically relevant individual-level predictions that could be useful for targeting early interventions.

Humans

Toward AI Virtual Cells for Hepatology: Representation, Generation, Dynamics, and Intervention in Single-Cell Models.

``Single-cell and spatial atlases describe the healthy and diseased liver at high resolution, including lobular hepatocyte zonation, fibrotic macrophage-stellate niches, cholangiocyte reactions, immune remodeling, and hepatocellular carcinoma ecosystems. These maps show where cell states occur but do not, by themselves, predict whether liver injury will progress or how the liver will respond to an untested drug, toxicant, or genetic perturbation. In this review, we organize current approaches toward an AI Virtual Cell (AIVC) for the liver into three complementary modeling routes. Generative models represent cell states, dynamics and transport models infer state transitions, and pretrained or foundation models test whether learned representations transfer across donors, etiologies, disease stages, and platforms. Perturbation-response prediction serves as a cross-cutting assessment of whether these layers can predict responses to untested genetic, chemical, inflammatory, or metabolic interventions. Available evidence can be categorized as direct liver validation, liver-included benchmarks, general single-cell evidence, and conceptual applications. Published models demonstrate individual components, including atlas integration, inferred trajectories, transferable representations, and retrospective response programs. However, these models do not constitute a prospectively validated liver simulator. At minimum, evaluation should include donor-, etiology-, stage-, platform-, and perturbation-level hold-outs. Model performance should be reported using response direction, recovery of differentially expressed genes and rare states, and calibrated uncertainty. Claims about tissue- or function-level prediction additionally require independent spatial, histologic, metabolic, and functional readouts. Near-term use should prioritize experiment selection and hypothesis generation, whereas clinical decision support remains a longer-term objective.

AI Virtual Cell

Kinetic model for production and metabolism of very low density lipoprotein triglycerides. Evidence for a slow production pathway and results for normolipidemic subjects.

A model for the synthesis and degradation of very low density lipoprotein triglyceride (VLDL-TG) in man is proposed to explain plasma VLDL-TG radioactivity data from studies conducted over a 48-h interval after injection of glycerol labeled with 14C, 3H, or both. The curve describing the radioactivity of plasma VLDL triglycerides reaches a maximum at about 2 h, after which the decay is biphasic in all cases; the late curvature becoming evident only after 8--12 h. To fit the complex curve, it was necessary to postulate two pathways for the incorporation of plasma glycerol into VLDL-TG, one much slower than the other. A process of stepwise delipidation of VLDL in the plasma compartment, previously proposed for VLDL apoprotein models, was also necessary. Predicted VLDL-TG synthesis rates calculated with this model can differ significantly from those based on experiments of shorter duration in which the slow VLDL-TG component is not apparent. The results of these studies strongly support the interpretation that the late, slow component of the VLDL-TG activity curve is predominantly due to the slowly turning-over precursor compartment in the conversion pathway and is not due either to a slow compartment in the labeled precursor, plasma free glycerol, or to an exchange of plasma VLDL-TG with an extravascular compartment. It also cannot, in these studies, be attributed to a slowly turning-over VLDL-TG moiety in the plasma. The model was tested with data from 59 studies including normal subjects and patients with obesity and(or) various forms of hyperlipoproteinemia. Good fits were obtained in all cases, and the estimated parameter values and their uncertainties for 13 normolipemic nonobese subjects are presented. Sensitivty testing was carried out to determine how critical various parameter estimations are to the assumptions introduced in the modeling.

Glycerol