PubMed HealthSearch

SEARCH · PubMed Health

Results for “Model interpretability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

GT-Mamba: a Topology-Aware Graph-State space model for robust and interpretable epigenetic age prediction.

MOTIVATION: Current epigenetic clocks face a trade-off between predictive accuracy and biological interpretability, often relying on dataset-specific correction to generalize across cohorts. We propose GT-Mamba, a novel architecture that integrates a Structure-Aware Graph Transformer with the Mamba state space model. This design captures CpG topological correlations and genome-wide long-range dependencies. RESULTS: GT-Mamba demonstrates strong out-of-the-box robustness across heterogeneous independent validation cohorts, achieving a weighted average MAE of 4.43 years. Notably, it effectively generalizes to EPIC 850k arrays despite partial feature missingness, and maintains consistent performance across homologous age distribution shifts (MAE 2.94 years in a young cohort). Ablation studies confirm that graph topology contributes to improved robustness against noise. Mechanistic analysis suggests that the model captures methylation patterns associated with both developmental and functional processes. AVAILABILITY: Source code and pre-trained models are freely available at https://github.com/NENUBioCompute/GT-Mamba and archived on Zenodo (DOI: 10.5281/zenodo.19703155).

Epigenesis, Genetic

Electrical properties of frog skeletal muscle fibers interpreted with a mesh model of the tubular system.

This paper presents the construction, derivation, and test of a mesh model for the electrical properties of the transverse tubular system (T-system) in skeletal muscle. We model the irregular system of tubules as a random network of miniature transmission lines, using differential equations to describe the potential between the nodes and difference equations to describe the potential at the nodes. The solution to the equations can be accurately represented in several approximate forms with simple physical and graphical interpretations. All the parameters of the solution are specified by impedance and morphometric measurements. The effect of wide circumferential spacing between T-system openings is analyzed and the resulting restricted mesh model is shown to be approximated by a mesh with an access resistance. The continuous limit of the mesh model is shown to have the same form as the disk model of the T-system, but with a different expression for the tortuosity factor. The physical meaning of the tortuosity factor is examined, and a short derivation of the disk model is presented that gives results identical to the continuous limit of the mesh model. Both the mesh and restricted mesh models are compared with experimental data on the impedance of muscle fibers of the frog sartorius. The derived value for the resistivity of the lumen of the tubules is not too different from that of the bathing solution, the difference probably arising from the sensitivity of this value to errors in the morphometric measurements.

Animals

Factors influencing the enhancement of the new iron triangle in healthcare organisations.

PURPOSE: A new paradigm, "healthcare's new iron triangle," has been developed to emphasise the technological perspective of healthcare delivery, focusing on automation, value and empathy. The study aims to build a conceptual model and to identify factors for the enhancement of the new iron triangle in healthcare organisations. DESIGN/METHODOLOGY/APPROACH: The healthcare organisation is the primary focus point of the current study. To determine the factors, a survey of the literature and healthcare experts' opinions was conducted. The healthcare professionals validated the identified factors. Data for this study were gathered using a closed-ended questionnaire and scheduled interviews. The study employed "Total Interpretive Structural Modeling methodology and Matriced' Impacts Croise´s Multiplication Appliqué´ a UN Classement/Cross-Impact Matrix Multiplication Applied to a Classification (MICMAC) analysis" to address the "why" and "how" the factors interact and prioritise the identified factors. FINDINGS: The study found that organisational structure (F8), artificial intelligence (F1), innovation (F2) and human resources (F5) are the driving or key factors of the study. RESEARCH LIMITATIONS/IMPLICATIONS: The study primarily focused on identifying factors for the enhancement of a new iron triangle in healthcare organisations. The scope could eventually be expanded to explore more areas. PRACTICAL IMPLICATIONS: Academics and other stakeholders will have a better understanding of the key drivers for the enhancement of the new iron triangle in healthcare organisations. ORIGINALITY/VALUE: In this study, total interpretive structural modeling and cross-impact MICMAC analysis are proposed as an innovative approach to address the new iron triangle in healthcare organisations.

Humans

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans

Metabolism of totally ischemic excised dog heart. II. Interpretation of a computer model.

Analysis of the ischemic dog heart preparation described in the preceding paper indicates that it is an analogue in slow motion of the tissue in the center of a cardiac infarct. It is respiring very slowly and not capable of performing mechanical work. Glycolysis starts up with both glucose and glycogen as inputs. Later hexokinase and to some extent phosphofructokinase become limiting owing to inhibitor accumulation or acidosis. Metabolism then results primarily from cAMP-driven glycogenolysis, largely limited by the glycogen debranching enzymes at later times, with accumultion not only of lactate and alpha-glycerophosphate but of glucose as well. Amino acid levels oscillate with time while fatty acids accumulate at late times. The elevation of cAMP at later times may involve disturbances in its metabolism as well as mechanisms such as adenosine accumulation that are more important in cardiac ischemia than in normal heart. The clinical implications of this behavior are discussed.

Amino Acids

[Electronic measuring and calculating devices for arcogrammetric model diagnosis and for the interpretation of teleradiographs].

A newly developed method to mesure different parameters from plaster models of the teeth, from dental radiographs and from cephalometric X-rays by means of a 4-K minicomputer and on-line linear transducers are described. Programs in connection to Arcogrammetrics [Herren] are presented. The electronic devices allow storage of these parameters in order to make drawings of the actual dental arches and of predicted arches, as well as to trace growth and/or progress in orthodontic treatment.

Cephalometry

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy.

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Databases, Protein

Predicting food taste with bound-driven optimization.

The prediction of sensory attributes from ingredient-level formulations is an emerging challenge at the intersection of food science and artificial intelligence. We address the fundamental question of whether the taste of a food can be predicted from its ingredients by treating recipes as composite materials. We apply Hashin-Shtrikman (HS) and Reuss-Voigt (RV) bounds, techniques originally developed for elastic moduli, as a null-hypothesis additive baseline for five taste dimensions (sweetness, sourness, bitterness, umami, saltiness) on a curated dataset of 70 recipes decomposed into 115 distinct ingredients scored against a library of 209 ingredient-level taste references with trained-panel ground truth. This baseline systematically under-predicts perceived taste: 77% of actual taste values exceeded the HS upper bound, with the exceedance rate ranging from 26% (bitterness) to 97% (saltiness). We traced this gap to specific processing chemistry (Maillard reactions, caramelization, evaporative concentration, protein hydrolysis, and nucleotide synergy) and introduced a hybrid model that augments the HS baseline with eight chemistry-proxy features encoding these mechanisms. Our results show that our interpretable hybrid model eliminates the systematic bias and reduces mean absolute error by 27%-62% for sweetness, sourness, umami, and saltiness while using only 10 interpretable features, achieving performance comparable to a black-box Lasso regression on 115 per-ingredient features. We further demonstrate constrained inverse design via Differential Evolution, recovering ingredient formulations that match target taste profiles subject to compositional bounds. Our work demonstrates how key chemical processes during food preparation can inform and augment physics-based and machine learning models, providing a quantitative fingerprint of processing chemistry's contribution to taste perception and paving the way for model-driven food formulation with targeted sensory characteristics.

Composite material bounds

General theory of critical periods and development of obesity.

The general systems theory (GST), the general theory of organization (GTO), and the general theory of critical periods (GTCP) have been applied to some nutritional problems. This theoretical approach seems to be in good agreement with most of the data of the literature and with the personal experience, pointing at the possibility to use a simple general model to interpret the complex problem of obesity.

Critical Period, Psychological

Intrauterine pressure wave form characteristics in hypocontractile labor before and after oxytocin administration.

The data demonstrate that the contractions of hypocontractile active labor and normal spontaneous labor are different in several measures in addition to maximal amplitude. Furthermore, when the pathophysiology is corrected by the use of oxytocin, the contractions resemble those of normal spontaneous labor except in the maximal rate of tension development. Our data tend to support the subcellular model of uterine contractility, although the incompleteness of these models limits interpretation.

Adrenocorticotropic Hormone

Prediction of antimicrobial minimum inhibitory concentration from bacterial genomes using a scalable and interpretable machine learning approach.

Although machine learning models can predict antimicrobial susceptibility from bacterial whole genome sequencing (WGS), state-of-the-art approaches are computationally demanding or dependent on knowledge of genetic resistance determinants. Here, we describe an efficient data-driven approach to predicting minimum inhibitory concentration (MIC) by progressively extending and refining predictive genome segments, independent of prior knowledge of resistance determinants. Resultant models had high interpretability - known and potentially novel resistance determinants were captured. Using 762 clinical E. coli strains, 71.6% of predictions were within one dilution of the measured MIC. Models trained with this algorithm generalised better onto external data (F1 score = 0.85) compared with alternative models trained on annotated resistance determinants (F1 = 0.82) or k-mer counts (F1 = 0.74). Computational demands were low (RAM usage 23.6GB vs 38.8GB for k-mer model). These advantages represent an important advance in predicting antimicrobial susceptibility from WGS, with potential applications for clinical diagnostics, drug development, and surveillance.

Journal Article

Evaluating the pathogenic significance of unique chromosomal variants in craniosynostosis using patient-derived induced pluripotent stem cells and mouse modelling.

PURPOSE: Unravelling causal links between unique structural/copy-number variants (SV/CNV) and associated phenotypes is essential for correct genetic counselling. We investigated two families in which patients with craniosynostosis had SV/CNV potentially dysregulating a fibroblast growth factor (FGF)-encoding gene; a 730 kb dup(4)(q21.21) including FGF5; and a complex 568 kb interspersed 13q12.11 duplication, located 841 kb from FGF9. METHODS: We combined bioinformatic predictions of altered topologically-associating domain (TAD) structure, with experimental analysis (RNA- and ATAC- [assay for transposase-accessible chromatin] sequencing) of patient induced pluripotent stem cell lines (iPSCs) differentiated to neural crest (NCC) and osteoprogenitor (OPC) identities. For the dup(4)(q21.21) we generated a mouse bearing an equivalent rearrangement using CRISPR-Cas9 targeting. RESULTS: TAD analysis suggested potential dysregulation of the FGF5/FGF9 gene by bringing it into a novel genomic milieu. The RNA- and ATAC-seq assays demonstrated FGF5/FGF9 upregulation (2.7-18x) and local opening of chromatin, in 3/4 cell lines. For the dup(4)(q21.21), a causal role was supported by the mouse model, whereas interpretation of the 13q12.11 SV is confounded by a co-existing FOXP2 pathogenic variant. CONCLUSION: Patient iPSC-differentiated NCC and OPC lines, combined with TAD-based modelling to generate testable functional hypotheses, provide valuable functional evidence when evaluating causation of unique SV/CNV in craniosynostosis.

copy-number variant

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92&#x2009;million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study&#xa0;(GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that&#xa0;guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

Decoding TnsC Filament Assembly in CRISPR-Associated Transposons Using Interpretable Deep Learning and Molecular Simulations.

CRISPR-associated transposons (CASTs) enable programmable DNA integration, yet how the TnsC regulator forms processive filaments on DNA to coordinate RNA-guided transposition in type V-K CAST systems remains unknown. Here, we integrate large-scale molecular simulations, interpretable deep learning using graph attention networks (GATs), and causal inference analyses to define the molecular determinants of TnsC filament nucleation and elongation. We show that TnsC nucleates by inducing localized DNA deformation that propagates along extended filaments, with Granger causality revealing that TnsC motions precede and predict DNA deformation. Interpretable GAT models demonstrate that elongation is determined during early recognition between incoming and DNA-bound subunits, followed by structural reorganization that regenerates the recruitment interface and enables processive assembly. These results elucidate the molecular mechanism of processive TnsC filament assembly and explain why isolated TnsC filaments preferentially elongate in the 5' &#x2192; 3' direction, while accessory transposition factors can reshape the interaction landscape and alter filament growth polarity. Together, these findings advance our understanding of CAST function and inform the engineering of programmable DNA integration platforms. Beyond CAST systems, this work introduces an interpretable GAT approach as a general and transferable deep learning strategy for uncovering molecular mechanisms in biological systems, while demonstrating the power of causal inference for dissecting directional relationships in molecular dynamics.

Deep Learning

Interpretation of osmotic pressure in solutions of one and two nondiffusible components.

Osmotic pressure data from aqueous solutions of nondiffusible serum albumin (BSA), chondroitin sulfate (CHS), and dextran T110 (D110), taken singly and in binary combinations, were interpreted in terms of excluded volume. The principal solvent was phosphate-buffered saline, pH 7.2, at 23 degrees C. Osmotic pressures were measured with a membrane osmometer fitted with Amicon PM-10 membranes. Data from each solution were fit by stepwise regression with a three- or four-term polynomial in integral powers of total nondiffusible solute concentration in accordance with the general solution theory of McMillan and Mayer (1945, J. Chem. Phys. 13:276) as extended by Yamakawa (1971, Modern Theory of Polymer Solutions, Harper & Row, New York). The date display a high internal consistency, and the results correlate well with published molecular weights and exclusion data where available. Number average molecular weights calculated from the "first virial coefficients" are: BSA, 67,000 +/- 11%; D110, 76,000 +/- 11%, CHS, 39,000 +/- 6%. Excluded volumes (in cubic centimeters per molecule) calculated from the "second virial coefficients" are: BSA, 0.97 X 10(-18); D110, 3.04 X 10(-18); CHS, 14.3 X 10(-18); BSA-D110, 6.8 X 10(-18); BSA-CHS, 7.8 X 10(-18). Uncertainty is about 30%. An empirical model for interpretation of calculated excluded volumes is proposed. It appears that CHS has the "largest" exclusion effect of the three molecules.

Animals