PubMed HealthSearch

SEARCH · PubMed Health

Results for “probabilistic modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

Dynamic decision models for clinical diagnosis.

A unified approach to clinical decision-making is presented. This combines partially observable Markovian decision processes (Markov or semi-Markov) with cause-effect models as a probabilistic representation of the diagnostic process. Pattern recognition techniques are used in a first stage of system state identification. This new class of dynamic models has a direct application to medical diagnosis and treatment and specific physiological examples are emphasised. The methodology is given for combining the patient state of health, the clinician's state of knowledge of the cause-effect representation from the observation space (measurements), feature selection using pattern recognition techniques and, finally, the treatment decisions with which to restore the patient to a more desirable state of health. A cost functional for the decision process has then to be optimised according to some pre-assigned objective function (social return from the patient state of health or treatment cost for the patient), when the process has an infinite time horizon.

Computers

[Methodological problems of clinical trials of psychotropic drugs (author's transl)].

Up to now, clinical studies only succeeded in differentiating between great categories of psychotropic drugs, but failed to prove finer differences of effects within these categories of substances. Two points of the testing-method are discussed: 1. problems which arise when rating pathological behaviour and 2. problems of sampling psychiatric patients. A great part of symptoms that clinicians and psychologists used to consider as relevant proved to be extremely rare. Total scores cannot be taken as a measure of the therapeutic effect, because they don't express adequately the degree of severity of the illness before and after treatment, and there is a symptom that is independent from the observer, i.e. the frequency of the symptom in different clinical pictures. The frequency is an inverse ratio to the specificity of the symptom. It is then argued that even in clinical studies, it would be possible to choose among the variety of descriptive symptoms those which fulfil requirements of the probabilistic test-model of Rasch and to take only those symptoms to characterize the degree of severity of psychic disturbance in trials with psychotropic drugs. Conclusions are then drawn from a study including three groups of physicians (specialists for internal diseases, psychiatrists and general practitioners): failure to differentiate between placebo and a Minor-Tranquilizer was not due to the inefficiency of the drug, but ought to be attributed to the lack of sharpness of the observations made by untrained judges. A significant difference between placebo and the Minor-Tranquilizer was yet found, but only in the group of psychiatrists. The comparison of the first 13 and the last 13 cases in the two remaining groups, however, reveals a learning process in the course of the study. The main problems of sampling are discussed, i.e.: the loss of information as a consequence of taking the mean in a group of psychiatric patients, the role of biological rhythms, which was hitherto insufficiently considered, and finally it is demonstrated in connection with two selected cases of depressive patients that enormous difference of psychophysiological responsiveness can be hidden behind very similar clinical pictures. It is pointed out that the existing research strategy is adjusted to great samples, which were composed on the basis of behavioral characteristics, and that it failed to differentiate subtle effects of psychotropic drugs. Only experiments involving a much greater display, which take into account all aspects of observation of the selected single cases and longitudinal studies can answer the question which is the right medicine for the right patient. Psychophysiological and biochemical methods have here priority over other methods.

Drug Evaluation

From genes to trajectories: mapping genetic influences on Huntington's disease progression.

MOTIVATION: There are many diseases with established genetic factors, such as Huntington's disease (HD), that are characterized by variable rates of progression. However, beyond the contribution of the known genetic factors - in this case the Huntingtin (HTT) gene - the impact of the full human genome on the natural progression of such diseases throughout a patient's life remains largely unknown. The increased availability of genome wide association (GWA) data in HD gene expansion carriers (HDGECs), combined with the clinical assessment scores on the same set of patients, has provided a perfect opportunity to assess the potentially broader genetic impact on the natural progression of HD. RESULTS: We present a genetics-driven, probabilistic disease progression model designed to identify and investigate the ways in which a range of genetic factors affect the natural progression of HD. When applied to a clinico-genomic HD dataset, our model identified several single nucleotide polymorphisms (SNPs) with previously unreported effects on disease progression that act at distinct stages and with varying magnitudes. This discovery may shed light on the potential mechanistic impact of previously unidentified genes on HD that may have implications for clinical management. As increasing amounts of GWA data become available more generally, we anticipate that this modeling framework will be broadly applicable to other diseases with strong genetic components. AVAILABILITY AND IMPLEMENTATION: The source code for IHDPM is available at https://github.com/BiomedSciAI/IHDPM.

Huntington Disease

Medical maxims: two views of science.

Clinicians are beings with finite minds and thus need to use simplified models of the world in making decisions [1]. However, these models need not be oversimplified, as are the models encouraged by the prevalent view of science, the Mechanistic Paradigm, which are articulated as rigid maxims. A more current view of science, the Probabilistic paradigm, encourages more complex models, which can be articulated as the more flexible maxims used with insight by the wise clinician.

Models, Biological

Storing covariance with nonlinearly interacting neurons.

A time-dependent, nonlinear model of neuronal interaction which was probabilistically analyzed in a previous article is shown here to be a natural generalization of the Hartline-Ratliff model of the Limulus retina. Although the primary physical variables in the model are the membrane potentials of neurons, the equations which govern the means and covariances of the membrane potentials are coupled through the average firing rates; as a consequence, the average firing rates control the selective storage and retrieval of covariance information. Motor learning in the cerebellar cortex is treated as a problem of covariance storage, and a predicition is made for the underlying synaptic plasticity: the change in synaptic strength between a parallel fiber and a Purkinje cell should be proportional to the covariance between discharges in the parallel fiber and the climbing fiber. Unlike previous proposals for synaptic plasticity, this prediction requires both facilitation and depression to occur (under different conditions) at the same synapse.

Cerebellum

The scientific status of the Rorschach.

Demonstrated that the interpretation of projective test data is semantic, not probabilistic. The clinician does not employ the model of statistical inference in evaluating the meaning of test responses, although he may employ probabilistic rules as guides to interpreation. Procedural rules for interpreting meaning, the nature of clinical diagnosis from psychological tests, and the meaning of prediction as a clinical activity are discussed. It was concluded that clinicians do not make inferences, in the mathematical sense of this term.

Association

High-Purity Monovalent Functionalization of Carbon Nanotubes.

Single-walled carbon nanotubes (SWCNTs) show promise for probing molecular interactions at single-molecule resolution, yet generating SWCNT populations bearing a single defined functional tag remains challenging because surface functionalization is inherently stochastic. Here, we present a batch-scale strategy to produce predominantly singly tagged SWCNTs by leveraging the stochastic adsorption of single-stranded DNA (ssDNA). Specifically, SWCNTs are dispersed using a mixture of unmodified ssDNA (um-ssDNA) and a minor fraction of modified ssDNA (m-ssDNA) carrying an affinity handle. We developed a probabilistic ssDNA-SWCNT binding model that predicts the distribution of m-ssDNA per nanotube as a function of the input minor-strand fraction p = m-ssDNA/total ssDNA, enabling selection of conditions that maximize single-tag purity. Using magnetic-bead capture via a biotin affinity interaction and subsequent release, we isolate SWCNTs with 97.6% predicted single-tag purity at 2% recovery. Single-molecule fluorescence imaging further supports predominantly single-label occupancy under the model-selected conditions. Thus, this approach provides a general route to SWCNTs bearing a single molecular handle for downstream conjugation and assembly, supporting diverse future applications in SWCNT-based nanotechnologies.

Nanotubes, Carbon

A review of parameter values used to assess the transport of plutonium, uranium, and thorium in terrestrial food chains.

A general methodology of predicting the food chain transport of atmospherically deposited radionuclides is reviewed with an emphasis on variation in parameter values important for realistic behavioral characterization of environmental releases of plutonium, uranium, and thorium. Parameters important to generic simulations of food chain transport, given a known constant deposition onto vegetation, include: fractional interception of particulates by vegetation, vegetation density, effective half-life of contamination on vegetation, soil-to-plant transfer factors, consumption rates by cattle and man, and transfer of nuclides from forage to meat and from forage to milk. Variation in these parameters, which has been encountered in field studies, is summarized. A partial reduction in the variation of predicted concentrations of actinides in foods can be accomplished by more accurately determining critical parameter values like fractional interception of deposition by vegetation, vegetation biomass, and the effective half-life of contamination on vegetation. Field research describing the site dependency and time dependency of probability density functions for model parameter values is needed to make probabilistic predictions concerning Pu, U, and Th transport in food chains and to reduce the uncertainty associated with model predictions and generic assessments of environmental impact.

Biological Transport

Accelerating inference in genomic and proteomic foundation models via speculative decoding.

MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2×-1.4×), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.

Genomics

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics

LAML-Pro: joint maximum likelihood inference of cell genotypes and cell lineage trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (i) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (ii) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈25%-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is implemented in C++ and is available as both a command-line interface and as a Python library at: github.com/raphael-group/LAML-Pro.

Cell Lineage

Role of nuclear size in cell growth initiation.

Swiss 3T3 cells arrested in B0 (quiescent state) by reducing serum content of the medium all contain the same amount of DNA but vary in nuclear volume over approximately a twofold range. By use of flow microfluorimetry, scatterplots of nuclear volume versus DNA content were obtained in intervals after serum stimulation. The earliest cells to enter DNA synthesis were those with the largest nuclei, whereas cells with the smallest nuclei were among the latest. Regulation of cellular transit from G0 to the S phase was therefore, at least in part, deterministic, since all G0 cells did not have equal probabilities of entry into S at a given moment. All cells having the same nuclear volume did not initiate DNA synthesis at the same moment; therefore, factors other than nuclear volume must also influence this timing. Nuclear volume correlated with the maximum rate at which cells could enter S. The kinetic model of the cell cycle postulating a probabilistic event as solely responsible for entry into S thus appears too simple.

Animals

Deconvolution of evolutionary architecture unmasks a high-risk, subclonal-rich subtype in treatment-naive small cell lung cancer.

BACKGROUND: Intratumoral heterogeneity (ITH) drives therapeutic resistance in small cell lung cancer (SCLC). However, conventional single-sample analysis has limited horizontal, cross-patient comparisons, leaving the overarching evolutionary architecture in treatment-naive tumors poorly understood. This study aims to deconvolve these architectures to identify clinically relevant evolutionary subtypes. METHODS: We analyzed whole-exome sequencing data from 41 treatment-naive SCLC patients. To overcome the cross-patient comparability bottleneck, we developed a novel probabilistic framework using a refined Gaussian Mixture Model (GMM). This standardized subclonal structures into four hierarchical strata, enabling the identification of evolutionary subtypes via unsupervised clustering. To address the scarcity of SCLC public data, prognostic concordance was robustly explored in The Cancer Genome Atlas (TCGA) lung squamous cell carcinoma (LUSC) based on shared smoking etiology, with lung adenocarcinoma (LUAD) serving as a negative control. RESULTS: The cohort robustly segregated into "Clonal-dominant" (Group 1, n=28) and "Subclonal-rich" (Group 2, n=13) subtypes. Group 1 evolution was primarily driven by tobacco signatures (SBS4). Conversely, Group 2 exhibited late-stage acquisition of a DNA mismatch repair deficiency (MMRd) signature (SBS15), fueling trace subclonal diversification. Clinically, Group 2 demonstrated a significantly lower objective response rate (ORR) to platinum-based regimens (25.0% vs. 81.3%, P=0.02). Furthermore, the Subclonal-rich architecture independently predicted inferior overall survival (OS) [adjusted hazard ratio (adj. HR) =2.93, P=0.02], driven predominantly by limited-stage disease. Cross-cancer analysis validated this histology-dependent, high-heterogeneity adverse pattern in early-stage LUSC but not in LUAD. CONCLUSIONS: This hypothesis-generating study demonstrates that a "Subclonal-rich" architecture, driven by acquired MMRd, identifies high-risk, chemo-resistant SCLC. Our GMM approach suggests that pre-existing heterogeneity may serve as a potential, histology-dependent prognostic marker that warrants prospective validation for tailoring future therapeutic regimens.

Gaussian Mixture Model (GMM)

LAML-Pro: Joint Maximum Likelihood Inference of Cell Genotypes and Cell Lineage Trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (1) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (2) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈ 25-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is freely available at: github.com/raphael-group/LAML-Pro.

Journal Article