PubMed HealthSearch

SEARCH · PubMed Health

Results for “probabilistic modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Deficiency in POLE Exonuclease Causes Synthetic Lethality in Highly Aneuploid Cancer Cells.

UNLABELLED: Aneuploidy is a hallmark of cancer and is associated with drug resistance and poor clinical outcomes across diverse cancer types. However, no therapies have been clinically established to target highly aneuploid tumors. By analyzing nearly half a million tumor samples subjected to comprehensive genomic profiling, we identified a striking mutual exclusivity between POLE exonuclease domain mutations and high aneuploidy burden. This observation was independently validated using data from The Cancer Genome Atlas (TCGA) and the Cancer Cell Line Encyclopedia (CCLE). Probabilistic modeling revealed that the elevated quantity and unique spectrum of mutations induced by POLE exonuclease deficiency increase the likelihood of inactivating essential genes on chromosome arms harboring losses, leading to a synthetic lethal phenotype in highly aneuploid cells. Functional experiments demonstrated that POLE exonuclease activity is essential for the viability of highly aneuploid cancer cell lines but dispensable in diploid cells. These findings suggest that selective inhibition of POLE exonuclease activity may represent a promising therapeutic strategy for targeting highly aneuploid tumors. SIGNIFICANCE: An integrated approach using large-scale genomic analyses, probabilistic modeling and functional validation identified POLE exonuclease as a potential synthetic lethal target to overcome cancer aneuploidy.

Humans

A probabilistic generative model for quantification of DNA modifications enables analysis of demethylation pathways.

We present a generative model, Lux, to quantify DNA methylation modifications from any combination of bisulfite sequencing approaches, including reduced, oxidative, TET-assisted, chemical-modification assisted, and methylase-assisted bisulfite sequencing data. Lux models all cytosine modifications (C, 5mC, 5hmC, 5fC, and 5caC) simultaneously together with experimental parameters, including bisulfite conversion and oxidation efficiencies, as well as various chemical labeling and protection steps. We show that Lux improves the quantification and comparison of cytosine modification levels and that Lux can process any oxidized methylcytosine sequencing data sets to quantify all cytosine modifications. Analysis of targeted data from Tet2-knockdown embryonic stem cells and T cells during development demonstrates DNA modification quantification at unprecedented detail, quantifies active demethylation pathways and reveals 5hmC localization in putative regulatory regions.

5-Methylcytosine

Probability of conduction deficit as related to fiber length in random-distribution models of peripheral neuropathies.

This paper presents a set of probabilistic models which reproduce the proximodistal gradient of sensory deficit in peripheral neuropathies, on the basis of the occurrence of axonal dysfunction as a result of randomly distributed abnormalities. The models, which are based on conduction block, loss of temporal coherence, and weak interactions between nerve fibers, demonstrate that randomly distributed axonal dysfunction provides a sufficient condition for distal sensory deficit. The models predict a marked reduction in the length for normal sensory conduction with small increases in the probability of axomal dysfunction, providing a possible correlate for the rapid clinical progression of some neuropathies. The hypothesis that weak interactions between fibers result in paresthesiae in peripheral neuropathies is also discussed.

Humans

Bayesian inference of lineage trees by joint analysis of single-cell multimodal lineage-tracing data with BiLinT.

The advent of single-cell lineage-tracing technologies has enabled the simultaneous profiling of gene expression and lineage barcodes. However, accurate, high-resolution reconstruction of cell lineage trees remains challenging because most existing approaches treat these modalities separately and therefore fail to fully exploit their complementary information. Here we present BiLinT, a Bayesian framework that jointly models multimodal single-cell lineage-tracing data for lineage tree reconstruction. BiLinT integrates barcode evolution (a continuous-time Markov chain) with gene expression dynamics (an Ornstein-Uhlenbeck process) within a unified probabilistic model. Across synthetic and real data sets, BiLinT provides accurate lineage-tree reconstruction and reveals differentiation-associated clonal structure and developmental fate biases.

Journal Article

Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.

Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

Gene Regulatory Networks

Metax enables accurate cross-domain taxonomic profiling of metagenomes.

Taxonomic profiling is fundamental to microbiome research, yet achieving high species-level accuracy remains challenging for complex communities that span bacteria, viruses, eukaryotes, and archaea, and these limitations are exacerbated in low-biomass, host-dominated samples. We introduce Metax, a cross-domain taxonomic profiler that integrates coverage-based probabilistic modeling with an expectation-maximization framework to distinguish true microbial signals from artifacts. Across >600 samples from host-associated, environmental, wastewater, and low-biomass clinical settings, including benchmarks with limited reference representation, Metax improved profiling accuracy, achieving on average 55% higher F1 scores and 45% lower Bray-Curtis dissimilarity than other methods. Moreover, this broad evaluation demonstrated that Metax resolved bacterial and viral signatures of peri-implantitis in oral microbiomes and revealed signals suggestive of reagent-borne contaminants and reference misassemblies in plasma-cell-free DNA. By leveraging genome-wide coverage evidence, Metax enables robust cross-domain profiling across diverse sample types and sequencing depths, including settings where reference databases are highly incomplete.

abundance estimation

Multimodal profiling reveals tissue-directed signatures of human immune cells altered with age.

The immune system comprises multiple cell lineages and subsets maintained in tissues throughout the lifespan, with unknown effects of tissue and age on immune cell function. Here we comprehensively profiled RNA and surface protein expression of over 1.25 million immune cells from blood and lymphoid and mucosal tissues from 24 organ donors aged 20-75 years. We annotated major lineages (T cells, B cells, innate lymphoid cells and myeloid cells) and corresponding subsets using a multimodal classifier and probabilistic modeling for comparison across tissue sites and age. We identified dominant site-specific effects on immune cell composition and function across lineages; age-associated effects were manifested by site and lineage for macrophages in mucosal sites, B cells in lymphoid organs, and circulating T cells and natural killer cells across blood and tissues. Our results reveal tissue-specific signatures of immune homeostasis throughout the body, from which to define immune pathologies across the human lifespan.

Humans

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

PhyClone: accurate Bayesian reconstruction of cancer phylogenies from bulk sequencing.

MOTIVATION: Cancer is driven by somatic mutations that result in the expansion of genomically distinct sub-populations of cells called clones. Identifying the clonal composition of tumours and understanding the evolutionary relationships between clones is a crucial task in cancer genomics. Bulk DNA sequencing is commonly used for studying the clonal composition of tumours, but it is challenging to infer the genetic relationship between different clones due to the mixture of different cell populations. RESULTS: In this work, we introduce a new probabilistic model called PhyClone that can infer clonal phylogenies from bulk-sequencing data. We demonstrate the performance of PhyClone on simulated and real-world datasets and show that it outperforms previous methods in terms of accuracy and sample scalability. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at: https://github.com/Roth-Lab/PhyClone under the GPL v3.0 license.

Neoplasms

Simple scaling laws control the genetic architectures of human complex traits.

Genome-wide association studies have revealed that the genetic architectures of complex traits vary widely, including in terms of the numbers, effect sizes, and allele frequencies of significant hits. However, at present we lack a principled way of understanding the similarities and differences among traits. Here, we describe a probabilistic model that combines the effects of mutation, drift, and stabilizing selection at individual sites with a genome-scale model of phenotypic variation. In this model, the architecture of a trait arises from the distribution of selection coefficients of mutations and from two scaling parameters. We fit this model for 95 highly polygenic quantitative traits of different kinds from the UK Biobank. Notably, we infer that all these traits have fairly similar, though not identical, distributions of selection coefficients. This similarity suggests that differences in architectures of highly polygenic traits arise mainly from the two scaling parameters: the mutational target size and heritability per site, which vary by orders of magnitude among traits. When these two scale factors are accounted for, we find that the architectures of all 95 traits are very similar.

Humans

Novel Predictive Spatial Biomarker in Non-Small Cell Lung Carcinoma: The Diversity of Niches Unlocking Treatment Sensitivity (DONUTS).

Probabilistic spatial modelling techniques developed on large-scale tumor-immune Atlases (~35M individually mapped cells; 50,000 high power fields) were used to characterize predictive features of treatment-responsive lung cancer. We identified CD8+FoxP3+ cell density as a robust pre-treatment biomarker for outcomes across disease stages and therapy types. In parallel, single-cell RNAseq studies of CD8+FoxP3+ T-cells revealed an activated, early effector phenotype, substantiating an anti-tumor role, and contrasting with CD4+FoxP3+ T-regulatory cells. A spatial biomarker was developed using an empirical probabilistic model to define the immediate cell neighbors or niche surrounding CD8+FoxP3+ cells and proximity to the tumor-stromal boundary. The resultant 'Diversity of Niches Unlocking Treatment Sensitivity (DONUTS)' are more prevalent than the CD8+FoxP3+ cells themselves, mitigating sampling error in small biopsies. Further, the DONUTS only require four markers, are additive to PD-L1, and associate with tertiary lymphoid structure counts. Taken together, the DONUTS represent a next-generation predictive biomarker poised for clinical implementation.

AstroPath

An Application of Iterative Health Economic Evaluation: An Update on the Early Cost Effectiveness of Whole-Genome Sequencing in Advanced Non-small-Cell Lung Cancer.

OBJECTIVE: Whole genome sequencing (WGS) can identify more druggable targets than the standard of care (SoC) panels, however, its health effects and costs are highly uncertain. Given the rapidly evolving treatment landscape and pricing, an iterative approach is crucial to continuously reassess evidence and adapt economic models. Our objective was to update a previously developed economic model for WGS. METHODS: We used a structured approach to identify and report model elements requiring updates, based on established tools and methodological guidance, and applied it to the probabilistic decision model by Simons et al.(2021), which compared SoC, WGS, and SoC followed by WGS in patients with inoperable stage IIIB, C/IV NSCLC in the Dutch setting. RESULTS: Updates included a new treatment (sotorasib), revised drug and diagnostic costs, and adherence to the latest guidelines. Drug and WGS diagnostics costs fell by 8% and 26%, respectively. SoC diagnostic prices increased by 17%. We explored the impact of the prevalence of druggable targets, effectiveness of off-label treatments, (academic-specific) diagnostic costs, and price negotiations. The ICER of WGS versus SoC decreased from €737,197 to €419,053/QALY. WGS would become cost-effective if diagnostic costs descended from €2,180 to €1,246 or if additional druggable targets were identified in ≥3.3% of patients. CONCLUSION: Our structured approach effectively identified items in the original analysis requiring updates and provides a foundation for further developing a checklist to guide iterative HTA. Continued monitoring and assessment of new treatment options, the dynamic diagnostics and costs throughout the life-cycle remain necessary to determine when WGS can be considered cost-effective.

NSCLC

Localized PD-1 CAR T therapy reprograms neuroinflammation.

B cell-depleting therapies are effective in multiple sclerosis (MS), yet some patients relapse, underscoring the need for more precise interventions. To identify new therapeutic targets, we generated a single-cell RNA sequencing (scRNA-seq) atlas of cerebrospinal fluid (CSF), brain, and blood from non-inflammatory controls and patients with MS or other neuroinflammatory diseases. We found disease-associated enrichment of class-switched immunoglobulin G+ (IgG+) B cells and plasma cells in MS CSF. Unbiased analysis identified a rare disease-enriched subset of activated, T cell receptor (TCR)-restricted, PD-1+ T follicular helper-like cells with B cell-recruiting features. To target this population, we developed PD-1-directed chimeric antigen receptor (CAR) T cells that selectively depleted pathogenic PD-1+ CD4 T cells and locally released IL-10. This strategy attenuated central nervous system (CNS) inflammation, reprogrammed the local immune milieu, and improved clinical outcomes across murine neuroinflammation models. These findings define a CNS-localized adaptive immune circuit in MS and nominate programmable PD-1 CAR T cells as a strategy to disrupt it.

Animals

Leveraging functional annotations to map rare variants associated with Alzheimer disease with gruyere.

Increased availability of whole-genome sequencing (WGS) has facilitated the study of rare variants (RVs) in complex diseases. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most do not fully leverage the availability of variant-level functional annotations. We propose genome-wide rare variant enrichment evaluation (gruyere), an empirical Bayesian framework that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimer's Disease Sequencing Project to identify Alzheimer disease (AD)-associated genes and annotations. Growing evidence suggests that the disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not examined rare non-coding effects that incorporate such cell-type-specific information. To address this gap, we (1) define per-gene non-coding RV test sets using predicted enhancer and promoter regions in microglia and other brain cell types (oligodendrocytes, astrocytes, and neurons) and (2) include cell-type-specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 13 significant genetic associations not detected by other RV methods, four of which remain significant in omnibus tests. We find that deep-learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a robust framework incorporating functional annotations, coding RVs, and cell-type-associated non-coding RVs to perform genome-wide association tests, uncovering AD-relevant genes and annotations.

Alzheimer Disease

Network methods for diagonal integration of unpaired single-cell multiomics data: a review.

MOTIVATION: Advances in single-cell sequencing have enabled multiomics profiling at unprecedented resolution; however, mass spectrometry-based single-cell proteomics (scMS) remains inherently destructive, precluding simultaneous transcriptomic capture. Unlike antibody-based methods such as CITE-seq, which permit paired profiling but are restricted to targeted protein panels, scMS provides unbiased, genome-scale coverage of the intracellular proteome yet necessitates post hoc integration of unpaired datasets. This diagonal integration challenge, where transcriptomes and proteomes are measured in separate cells lacking shared anchors, remains underserved by existing reviews, which focus predominantly on vertical integration strategies enabled by non-destructive assays. RESULTS: We survey the complete computational pipeline for constructing mechanistic proteogenomic networks from unpaired single-cell data, covering: (i) unimodal network inference such as knowledge-based approaches, probabilistic graphical models, temporal directionality inference, and generative and foundation model strategies that establish the transcriptomic scaffold; (ii) cross-modal integration architectures such as network propagation, graph neural networks (scMRDR, scmFormer, scCotag), and consensus frameworks designed explicitly for the unpaired proteomics setting; and (iii) benchmarking paradigms spanning network reconstruction (BEELINE, GRETA, CausalBench) and multi-task integration evaluation (scMultiBench, SCMMIB), with guidance on metric selection under network sparsity and class imbalance. We identify three principal axes of future development: generative proteomic translation from transcriptomic precursors, inductive prior embedding in next-generation architectures, and perturbation-based causal benchmarking. AVAILABILITY AND IMPLEMENTATION: This is a review article; no novel software is distributed. A curated benchmark resource table, methods starter guide, and per-method bottleneck annotations are provided in the Supplementary Material.

Multiomics

A divide and conquer strategy for recapitulating whole genome 3D structure using Hi-C data.

The three dimensional (3D) spatial organization of the genome is closely linked to biological functions and can be captured by Hi-C assays through interrogating genome-wide chromatin interactions. Methodologies for inferring 3D structures from Hi-C data summarized as a two-dimensional (2D) contact matrix can be broadly placed within the paradigms of optimization-based and sampling-based. Many optimization-based methods are capable of constructing whole genome 3D structures but do not account for spatial dependency in the 2D data matrix nor cell heterogeneity in bulk Hi-C data, which provide an average over millions of cells. Sampling-based methods, on the other hand, are probabilistic model-based and can account for not only dependency, heterogeneity, but also other features inherent in Hi-C data, such as over-dispersion and sparsity. However, whole-genome 3D structure recapitulation is too computationally expensive for sampling-based methods, while chromosome-by-chromosome strategies for sampling-based methods ignore important information on inter-chromosomal contacts. To address these issues, we propose the truncated Random effect EXpression-cut and paste (tREX-cap) method, which applies the tREX model within a divide and conquer strategy. The resulting method inherits the good data-feature-cognizant properties of tREX and, in the meantime, can efficiently infer the whole genome 3D structure. We demonstrate the performance of tREX-cap through an extensive simulation study and analyses of a Hi-C lymphoblastoid dataset and a Hi-C IMR90 dataset.

Humans