PubMed HealthSearch

SEARCH · PubMed Health

Results for “Uncertainty modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq

Robust optimisation for photon radiotherapy: A scoping review of models, paradigms, and reporting.

BACKGROUND AND PURPOSE: Robust optimisation offers an alternative to conventional margin-based photon radiotherapy planning by explicitly modelling uncertainty, but practice is variable and not standardised. MATERIALS AND METHODS: A scoping review was conducted to map robust optimisation for photon external beam radiotherapy. Electronic searches of Scopus, PubMed and Google Scholar (2000-2025, English language) identified planning studies that incorporated modelled uncertainties into the optimisation process and reported at least one robustness-related outcome. Data were charted on clinical context, uncertainty models, optimisation paradigms, robustness metrics and evidence for clinical implementation. RESULTS: Seventy-one studies were included. Most investigated prostate, breast or lung cancer and used intensity-modulated radiotherapy or volumetric-modulated arc therapy in commercial or research treatment planning systems. Scenario-based worst-case (minimax) optimisation was the dominant paradigm in clinically oriented work, while chance-constrained, conditional value at-risk, distributionally robust and adaptive formulations were confined to small methodological series. Uncertainty modelling focused mainly on rigid set-up error; fewer studies incorporated respiratory motion, inter-fraction anatomical change, dose-calculation uncertainty or biological variation. Robustness was evaluated with diverse scenario-based dose-volume metrics, probabilistic coverage measures, composite robustness indices and, less often, biological endpoints. Direct clinical implementation reports were scarce. CONCLUSION: Robust photon planning is technically feasible and generally maintains or improves target coverage and organ sparing compared with margin-based planning. However, heterogeneity in uncertainty models, optimisation configuration and robustness reporting limits comparison and synthesis. Pragmatic minimum standards are proposed to support future consensus and wider clinical adoption.

Humans

Modeling unknowns: A vision for uncertainty-aware machine learning in healthcare.

The integration of machine learning (ML) into healthcare is accelerating, driven by the proliferation of biomedical data and the promise of data-driven clinical support. A key challenge in this context is managing the pervasive uncertainty inherent in medical reasoning and decision-making. Despite its recognized importance, uncertainty is often underrepresented in the design and evaluation of clinical AI systems. Here we report an editorial overview of a special issue dedicated to uncertainty modeling in medical AI, which gathers theoretical, methodological, and practical contributions addressing this critical gap. Across these works, authors reveal that fewer than 4% of studies address uncertainty explicitly, and propose alternative design principles-such as optimizing for clinical net benefit or embedding explainability with confidence estimates. Notable contributions include the RelAI system for real-time prediction reliability, empirical findings on how uncertainty communication shapes clinical interpretation, and benchmarks for out-of-distribution detection in tabular data. Furthermore, this issue highlights the use of causal reasoning and anomaly detection to enhance system robustness and accountability. Together, these studies argue that representing, communicating, and operationalizing uncertainty are essential not only for clinical safety but also for building trust in AI-driven care. This special issue thus repositions uncertainty from a limitation to a foundational asset in the responsible deployment of ML in healthcare.

Machine Learning

Scale reliant mixed effects models enhance microbiome data analysis.

Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional - they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package.

Microbiota

scRNA-seq and bulk RNA-seq reveal the characteristics of macrophage copper metabolism and establish a risk signature in hepatocellular carcinoma.

BACKGROUND: Hepatocellular carcinoma (HCC) is a prevalent malignancy with an urgent need for improved prognostic stratification and treatment-response prediction. This study aimed to explore a macrophage copper metabolism-associated prognostic model and to investigate the relationship between this risk model and the tumor immune microenvironment. METHODS: The FindClusters function was used to analyze cell clusters, and CellChat and CellPhoneDB/LIANA were employed for cell-cell communication analysis. Copper metabolism-related genes were sourced from the MSigDB database. A prognostic risk model was established using least absolute shrinkage and selection operator (LASSO) analysis and multivariate Cox regression analysis, and a nomogram was constructed by integrating the prognostic model with clinicopathological factors. Additional analyses were performed to map the seven model genes in single-cell data, assess model uncertainty and robustness, evaluate macrophage/copper/cuproptosis-related transcriptional programs, and examine the correlations between risk score, immune infiltration and predicted drug sensitivity. RESULTS: Using single-cell RNA sequencing (scRNA-seq) data, we identified four macrophage subpopulations. Macrophages with high SPP1 expression showed close interaction with T cell populations and were associated with copper ion metabolism. By incorporating 141 copper metabolism-related genes and using The Cancer Genome Atlas Liver Hepatocellular Carcinoma (TCGA-LIHC) cohort, we constructed a seven-gene risk prediction model. Additional single-cell mapping showed that the model genes were detectable in the HCC single-cell dataset and showed a macrophage-associated expression pattern. The model showed moderate prognostic discrimination in TCGA-LIHC, whereas its external performance was heterogeneous and remained evaluable across external cohorts, with performance varying among datasets. Immune and mechanism-related analyses suggested that the risk signature was associated with macrophage-related infiltration, copper metabolism and cuproptosis-related transcriptional programs. Drug sensitivity analysis nominated Daporinad as a computationally predicted candidate compound, supporting Daporinad as a pharmacogenomic candidate for follow-up investigation. CONCLUSIONS: By integrating scRNA-seq and bulk RNA sequencing (RNA-seq) data, we constructed a macrophage copper metabolism-associated prognostic signature for HCC. The risk score was associated with survival, immune microenvironment features and predicted drug response, providing a transcriptomic framework for risk stratification and therapeutic hypothesis generation.

Hepatocellular carcinoma (HCC)

PaNDA: Efficient Optimization of Phylogenetic Diversity in Networks.

Phylogenetic diversity (PD) plays an important role in biodiversity, conservation, and evolutionary studies by measuring the diversity of a set of taxa based on their phylogenetic relationships. In phylogenetic trees, a subset of k taxa with maximum PD can be found by a simple and efficient greedy algorithm. However, this algorithmic tractability is lost when considering phylogenetic networks, which incorporate reticulate evolutionary events such as hybridization and horizontal gene transfer. To address this challenge, we introduce PaNDA (Phylogenetic Network Diversity Algorithms), the first software package and interactive graphical user-interface for exploring, visualizing, and maximizing diversity in phylogenetic networks. PaNDA includes a novel algorithm to find a subset of k taxa with maximum diversity, running in polynomial time for networks of bounded scanwidth, a measure of tree-likeness of a network that grows slower than the well-known level measure. This algorithm considers the variant of PD on networks in which the branch lengths of all paths from the root to the selected taxa contribute towards their diversity. We demonstrate the scalability of this algorithm on simulated networks, successfully analyzing level-15 networks with up to 200 taxa in seconds. We also provide a proof-of-concept analysis using a phylogenetic network on Xiphophorus species, illustrating how the tool can support diversity studies based on real genomic data. The software is easily installable and freely available at https://github.com/nholtgrefe/panda. Additionally, we extend the definition of PD to semi-directed phylogenetic networks, which are mixed graphs increasingly used in phylogenetic analysis to model uncertainty of the root location. We prove that finding a subset of k taxa with maximum diversity remains NP-hard on semi-directed networks, but do present a polynomial-time algorithm for networks with bounded level.

network

Population-scale detection of methylation outliers from long-read genome sequencing.

BACKGROUND: Aberrant DNA methylation can mediate the functional effects of rare genetic variation and contribute to imprinting disorders, repeat expansion diseases, and other pathogenic regulatory mechanisms. Long-read sequencing technologies now enable genome-wide detection of CpG methylation alongside genetic variation from a single assay. However, methods for systematic identification and interpretation of methylation outliers from long-read sequencing data remain limited. METHODS: We developed METAFORA, a computational workflow for detecting methylation outlier regions from PacBio and Oxford Nanopore long-read sequencing data. METAFORA constructs population-level methylation references, segments the genome into correlated CpG blocks, infers technical and biological sources of variation through hidden factor estimation, models uncertainty due to variable depth sequencing, and computes covariate-adjusted methylation outlier scores for individual samples. We applied METAFORA across large long-read sequencing cohorts and integrated methylation outliers with multi-omic data. METAFORA is implemented as a snakemake workflow available at https://github.com/tjense25/METAFORA. RESULTS: METAFORA identified methylation outlier regions associated with rare structural variants, tandem repeat expansions, and imprinting abnormalities. We found outlier regions were enriched for molecular outliers across transcriptomic and chromatin accessibility datasets, supporting their functional relevance in gene regulation. In a representative case, METAFORA identified an imprinting defect affecting the GNAS locus associated with an STX16 deletion. CONCLUSIONS: METAFORA enables scalable detection and interpretation of methylation outliers from long-read sequencing data and provides a framework for integrating epigenetic outliers with genomic and multi-omic analyses. These approaches may improve interpretation of rare regulatory variation and support discovery of clinically relevant epigenetic abnormalities in genomic medicine.

DNA methylation

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

Rational design of high-productivity perfusion processes for CHO Cells: From growth inhibitory strategies to model-driven optimization.

While perfusion culture for Chinese hamster ovary (CHO) cells offers advantages such as continuous operation and flexibility, it suffers from product loss through cell bleeding and difficulties in reaching high productivity due to sustained rapid cell growth. Growth inhibitory strategies are widely used to enhance productivity in fed‑batch processes; however, their practical implementation and comparative effectiveness in perfusion processes remain insufficiently explored. Meanwhile, process development often relies on costly trial‑and‑error approaches. Here, we systematically compared three growth inhibitory strategies in perfusion culture-low cell‑specific perfusion rate (CSPR), sodium butyrate, and mild hypothermia-with respect to cell growth, metabolism, productivity, and product quality. Genome‑scale metabolic flux sampling analysis revealed that low‑CSPR and sodium butyrate induce a convergent up‑regulation of energy metabolism, correlating with greater gains in specific productivity (qp). Building on this insight, we developed a growth‑kinetic model for the combined low‑CSPR + butyrate strategy, incorporating parameter uncertainty. This model‑guided framework enabled the rational design of two distinct high‑productivity perfusion processes: a sustained mode that achieved robust long‑term stability alongside substantial productivity gains, and a high‑intensity mode that pushed qp and daily volumetric titer to their maxima, with increases of up to 108.94% and 190.36%, respectively, in a model CHO cell line with a moderate baseline productivity. Our study provides a proof‑of‑concept framework for perfusion intensification, from strategy selection to rational process design.

Animals

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Nerpa 2: probabilistic linking of biosynthetic gene clusters to nonribosomal peptides.

MOTIVATION: Nonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways. RESULTS: We present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC-product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry. AVAILABILITY AND IMPLEMENTATION: Nerpa 2 is freely available at https://github.com/gurevichlab/nerpa.

Multigene Family

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses

Diagnosing the undiagnosed: AI-enhanced multimodal modeling for placental mesenchymal dysplasia in high-risk pregnancies.

Placental mesenchymal dysplasia (PMD) is a rare vascular placental disorder that mimics molar pregnancy but often coexists with a viable fetus, making its misdiagnosis potentially devastating. In high-risk pregnancies, artificial intelligence (AI)-enhanced multimodal modeling - incorporating imaging, genomics, proteomics, and clinical features - offers a transformative diagnostic strategy. Leveraging Bayesian hyperparameter optimization for model refinement, this approach improves diagnostic accuracy while reducing uncertainty and clinician hesitation. Recent clinical studies support its efficacy and interpretability through SHAP and LIME models, while real-time surgical enhancements using Bayesian methods highlight its broader clinical utility. Despite current challenges such as data heterogeneity and integration barriers, multimodal AI provides unprecedented resolution in placental analysis, enabling precise differentiation between PMD and similar fetopathies. Ultimately, this advancement supports timely, non-invasive diagnosis, personalized management, and emotionally informed decision-making aligned with ethical AI implementation standards.

Bayesian optimization

An integrated multiscale air quality modelling framework for industrial park pollution: Linking local emissions to regional transport.

Capturing the spatiotemporal distribution of pollutants in industrial parks remains challenging for regional air quality models because of their coarse resolution (3 km), resulting in uncertainties in local emission quantification. To address this, we developed the Integrated Multiscale Air Quality Modelling System for Industry (IAQMS-Industry), coupling the regional Nested Air Quality Prediction Modelling System (NAQPMS) with a city-scale chemical transport model. This framework integrates point-source locations and Gaussian plume dispersion to simulate particulate matter with a diameter smaller than 2.5 micrometres (PM2.5) at 100 m resolution. Applied to the Beijing Yi Zhuang and Tangshan industrial parks and evaluated against observations. The coupled model achieved a normalized mean bias (NMB) ranging from 3.1 % to 6.2 %, improving upon NAQPMS (-16.9 % to -7.7 %). Spatial analysis revealed that coarse regional grids underestimated the PM2.5​ concentrations at industrial sites by smoothing gradients, whereas IAQMS-Industry successfully resolved spatial patterns. Industrial point emissions accounted for 22.9 %-26.4 % of PM2.5 in the coupled model, which was significantly greater than the regional model estimates of 1.6 %-13.7 %. These findings indicate that regional models overestimate pollutant dispersion processes in industrial parks while underestimating local industrial impacts. By explicitly resolving point-source dynamics and linking them to regional transport, IAQMS-Industry provides a robust tool for designing targeted emission controls in industrial cities and balancing local air quality improvements with minimized regional pollution outflow. This study underscores the necessity of multiscale modelling for accurate source apportionment and informed environmental governance in industrial zones.

Air Pollution

Adverse pregnancy outcomes and long-term cardiovascular disease risk.

Pregnancy provides a unique physiological stress test for the cardiovascular system, during which, adverse pregnancy outcomes (APOs) can unmask latent susceptibility to future disease. Common complications, including hypertensive disorders of pregnancy (HDP), gestational diabetes, and preterm birth (delivery before 37 weeks' gestation), identify women at substantially higher long-term risk of cardiovascular morbidity and mortality compared with women without a history of APOs. These excess risks likely reflect the combined effects of pre-existing cardiometabolic and genetic susceptibility, as well as the haemodynamic and metabolic stressors of pregnancy, heralding accelerated risk factor trajectories, relative impairment in endothelial and microvascular function, and early disease onset. This final Review in the Series extends the focus from cardiovascular disease during pregnancy and HDP to the long-term cardiovascular implications of APOs after delivery. We synthesise epidemiological data quantifying cardiovascular risk across major APO phenotypes and emerging evidence linking maternal APO history with cardiometabolic risk trajectories in offspring. We also delineate putative mechanistic pathways and summarise guidelines and consensus-informed recommendations for short-term and long-term follow-up after APOs. Finally, we propose practical approaches for integrating APO history into cardiovascular disease risk assessment and guideline-directed prevention across the female life course. We highlight key knowledge gaps, including uncertainty about optimal follow-up models, the limitations of current risk-stratification tools, and the absence of APO-specific prevention trials. We also outline priorities for mechanistic and implementation research. Positioning APOs as early, sex-specific indicators of cardiovascular risk offers a key window of opportunity to shift prevention upstream and improve cardiovascular health outcomes for women.

Humans

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

Inferring the demographic history of Chinese and Indian rhesus macaque (Macaca mulatta) populations from PacBio HiFi long-read sequencing data.

The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time ∼140,000 generations ago from an ancestral population of ∼65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of ∼220,000 individuals in the Chinese population and ∼14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

Cercopithecidae