PubMed HealthSearch

SEARCH · PubMed Health

Results for “Model interpretability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

Reliability-aware hierarchical learning for Chagas disease screening from 12-lead ECGs: tackling label uncertainty and class imbalance.

Objective.Chagas disease, a neglected tropical disease (NTD) with significant cardiovascular impact, remains underdiagnosed in resource-limited regions. Electrocardiogram (ECG) screening offers a low-cost tool for detecting cardiac involvement, yet algorithm development is challenged by label noise, data scarcity, and the latent nature of infection. This study proposes a robust ECG-based screening framework that explicitly addresses these constraints.Approach.We introduce aReliability-Aware Hierarchical Learningstrategy that calibrates supervision according to data provenance, prioritizing serology-confirmed labels over noisy self-reports. To mitigate data scarcity, we compare a specialized convolutional neural network (CNN) trained from scratch with a transfer learning approach based on a Spatio-Temporal ECG foundation Model (FM). Performance is evaluated across varying data scales, and the representation structure is analyzed to interpret model behavior.Main results.On the official hidden test set of the George B. Moody PhysioNet/Computing in Cardiology Challenge 2025, our approach achieved a Challenge Score of 0.163. We observe that while the specialized CNN performs competitively in data-rich regimes, the FM exhibits superior robustness in extreme low-resource settings. Furthermore, performance reaches a plateau imposed by underlying disease physiology. Bimodal score distributions suggest that models distinguish established cardiomyopathy from indeterminate infection, which remains electrophysiologically indistinguishable from healthy controls.Significance.These findings clarify both the potential and intrinsic limits of ECG-based AI screening for NTD-associated cardiac involvement. Reliability-aware supervision and data-efficient transfer learning provide a practical framework toward scalable and clinically meaningful ECG screening systems in resource-constrained environments.

Humans

Foundations of Artificial Intelligence in Hepatology: What a Clinician Needs to Know.

This review focuses on foundational knowledge about artificial intelligence (AI) in hepatology, exploring how AI, including machine learning and deep learning, leverages large-scale clinical data to transform the diagnosis, risk assessment, prognostication, and management of liver diseases. Online resources are described to offer fundamental AI knowledge and essential technical skills and to facilitate clinician participation across the entire AI lifecycle, ensuring they contribute not only as end users but also in development and deployment. Unlike traditional statistical approaches that prioritize interpretable parameters and clinical insight, AI focuses on maximizing predictive accuracy by identifying complex, often non-linear patterns using high-dimensional data, albeit often at the cost of model interpretability. AI is demonstrating clinical utility in liver histopathology and radiological imaging, significantly improving detection accuracy for cirrhosis, clinically significant portal hypertension, and hepatocellular carcinoma. Beyond diagnostics, AI-driven prediction models are emerging to provide personalized risk stratification for the development of liver-related complications and treatment guidance, based on complex data including longitudinal laboratory results, comorbidities, and co-medication use to monitor disease progression and therapy response. The field is rapidly expanding into novel areas such as analyzing patient-reported outcomes, genomic data, and real-time liver function monitoring, offering deeper mechanistic insights alongside clinical tools. Despite the potential to revolutionize hepatology practice and research, successful integration into routine care faces challenges. These include seamless workflow integration with existing electronic health records, establishing clear liability frameworks, and guaranteeing protection of patient privacy. Addressing these hurdles requires collaborative efforts from clinicians, researchers, and regulators to develop best practices and governance. Understanding the transformative capabilities, current applications, emerging frontiers, and essential implementation considerations is crucial for clinicians navigating the evolving AI landscape and responsibly utilizing its power for improved patient outcomes.

PROBAST+AI

A comprehensive review of AI innovations for tackling antimicrobial resistance.

Antimicrobial resistance (AMR) represents a major global public health concern, rendering available antimicrobials ineffective and leading to infections that are difficult to treat. Artificial intelligence (AI) has been increasingly applied across the AMR continuum, including resistance prediction, rapid diagnostics, new antimicrobial discovery, drug repurposing, antimicrobial surveillance, and clinical decision support. In this review, we aim to highlight recent developments in the use of artificial intelligence (AI) to address antimicrobial resistance (AMR). In addition, we review computational methods that help interpret genomic, phenomic, clinical, and epidemiological data to support the development of treatment strategies and novel antimicrobial agents. The key issues addressed include data quality, model interpretability, external validation, regulatory requirements, privacy, and fairness. While AI is not a complete solution to AMR, it can certainly strengthen the global AMR response by complementing key areas of AMR such as antimicrobial stewardship, infection prevention, laboratory diagnostics, and global surveillance.

Antimicrobial resistance (AMR)

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95 % CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article

Graph neural network-based risk stratification of prostate cancer using gene expression and SHAP interpretability.

Accurate risk stratification is essential for guiding treatment decisions and preventing over treatment of prostate cancer, which remains one of the most prevalent cancers among adult men. While the Gleason score, obtained from prostate biopsies, is routinely used to assess tumor aggressiveness, the biopsy procedure carries risks such as pain, infection, and, in some cases, serious complications such as sepsis. In this study, we proposed an artificial intelligence-based framework that integrates mRNA expression profiles with functional interaction networks to classify prostate cancer patients into low-, medium-, and high-risk groups defined by Gleason scores. The pipeline comprised five steps: (1) data collection from The Cancer Genome Atlas (TCGA), (2) preprocessing of gene expression data, (3) two-stage feature selection to identify informative biomarkers, (4) risk classification using a dual-branch graph neural network (GNN) that combines gene-gene interaction graphs with sample-level expression features, and (5) model interpretation using SHAP to quantify feature contributions. Differentially expressed genes were identified in the High (ASPN, GMNN, PEBP4, C2, KNCK17), Medium (C2, IGSF1, ASPN, CDKN3, AMH), and Low (TNMD, VWA5B2, ST6GALNAC5, CYP3A5, PHGR1) risk groups, underscoring the molecular heterogeneity of disease progression. On an independent held-out test set, the model achieved AUCs of 0.86, 0.88, and 0.95 for the low-, medium-, and high-risk groups, respectively, with an overall accuracy of 80%. These results suggest that combining GNN-based modeling with explainable AI can capture both global and local molecular patterns relevant to tumor aggressiveness. However, as the model was developed and evaluated solely on the TCGA cohort, the findings should be regarded as exploratory, and external validation will be required to establish generalizability. Within these limitations, the proposed framework highlights the potential of molecular profiling and graph-based deep learning to support more precise, potentially less invasive, risk assessment and individualized treatment planning in prostate cancer.

Prostatic Neoplasms

Structure-resolved virus-host interactomics by cross-linking mass spectrometry.

Viruses depend on host protein networks to replicate, assemble progeny, and spread between cells and organisms. Defining these virus-host protein interactions is challenging because they are highly dependent on infection stage, cell type, species, and because mechanistic interpretation requires information about structural interfaces and conformational states. Cross-linking mass spectrometry (XL-MS) addresses these challenges by adding a spatial and structural dimension to virus-host interactomics in native systems. In this review, we discuss how XL-MS has advanced from targeted analysis of viral protein complexes to structure-resolved mapping of virion architecture and infected-cell virus-host interactomes. We highlight how XL-MS complements AP-MS, cryo-EM/cryo-ET, quantitative proteomics, genetic perturbation, and structure prediction to connect physical proximity with molecular mechanisms. Finally, we discuss current limitations in sensitivity, chemical coverage, temporal resolution, and model interpretation, and outline how future quantitative and integrative XL-MS workflows may enable systems-level structural virology.

Mass Spectrometry

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans

Biological autoxidation. I. Decontrolled iron: an ultimate carcinogen and toxicant: an hypothesis.

Ionic iron at physiological pH hydrolyzes into insoluble aggregates, which disperse on slight acidification. Uncontrolled ionic iron promotes autoxidation, which crosslinks biomolecules and produces destructive activated oxygen. Defenses against autoxidative crosslinking include: 1. ferritin, the macromolecular scavenger of iron; 2. metabolic turnover, which prevents irreversible crosslinking through early catabolic degradation and replacement; and 3. enzymatic deactivation of oxygen. I am proposing that the anticrosslinking defenses are defeated by transient actions of metabolic perturbations, toxicants, oxidants and "foreign bodies", which cause oxidative crosslinking of proteins and lipids into irreversible tissue imprint: indigestible bodies containing porous limited-access spaces (LASs). The pores exclude the macromolecular ferritin and the digestive and antiautoxidation enzymes but admit ionic iron which, sheltered from ferritin, accumulates into decontrolled-iron pathogen (DIP). DIP utilizes the energy of ambient pH fluctuations to erupt from the LAS, swamp the available ferritin, poison the surroundings, catalyze autoxidation and crosslink cell components into additional LAS carriers. With time and sufficient promotion by pH fluctuations or metal-complexing agents, DIP and LAS expand. DIP injures through heavy-metal inhibition of life processes and catalysis of autoxidation. Typically, carcinogenic initiators are protein denaturants, cell poisons, "foreign bodies" and autoxidation catalysts. These are DIP-initiating properties, and DIP may be a preneoplastic stage of carcinogenesis. A DIP-model interpretation is given for the growth of asbestos bodies. DIP is an inorganic parasite. It may envelope and attack phagocytized particles.

Animals

Efficient Detection and Characterization of Targets of Natural Selection Using Transfer Learning.

Natural selection leaves detectable patterns of altered spatial diversity within genomes, and identifying affected regions is crucial for understanding species evolution. Recently, machine learning approaches applied to raw population genomic data have been developed to uncover these adaptive signatures. Convolutional neural networks (CNNs) are particularly effective for this task, as they handle large data arrays while maintaining element correlations. However, shallow CNNs may miss complex patterns due to their limited capacity, while deep CNNs can capture these patterns but require extensive data and computational power. Transfer learning addresses these challenges by utilizing a deep CNN pretrained on a large dataset as a feature extraction tool for downstream classification and evolutionary parameter prediction. This approach reduces extensive training data generation requirements and computational needs while maintaining high performance. In this study, we developed TrIdent, a tool that uses transfer learning to enhance detection of adaptive genomic regions from image representations of multilocus variation. We evaluated TrIdent across various genetic, demographic, and adaptive settings, in addition to unphased data and other confounding factors. TrIdent demonstrated improved detection of adaptive regions compared to recent methods using similar data representations. We further explored model interpretability through class activation maps and adapted TrIdent to infer selection parameters for identified adaptive candidates. Using whole-genome haplotype data from European and African populations, TrIdent effectively recapitulated known sweep candidates and identified novel cancer, and other disease-associated genes as potential sweeps.

Selection, Genetic

Efficient detection and characterization of targets of natural selection using transfer learning.

Natural selection leaves detectable patterns of altered spatial diversity within genomes, and identifying affected regions is crucial for understanding species evolution. Recently, machine learning approaches applied to raw population genomic data have been developed to uncover these adaptive signatures. Convolutional neural networks (CNNs) are particularly effective for this task, as they handle large data arrays while maintaining element correlations. However, shallow CNNs may miss complex patterns due to their limited capacity, while deep CNNs can capture these patterns but require extensive data and computational power. Transfer learning addresses these challenges by utilizing a deep CNN pre-trained on a large dataset as a feature extraction tool for downstream classification and evolutionary parameter prediction. This approach reduces extensive training data generation requirements and computational needs while maintaining high performance. In this study, we developed TrIdent, a tool that uses transfer learning to enhance detection of adaptive genomic regions from image representations of multilocus variation. We evaluated TrIdent across various genetic, demographic, and adaptive settings, in addition to unphased data and other confounding factors. TrIdent demonstrated improved detection of adaptive regions compared to recent methods using similar data representations. We further explored model interpretability through class activation maps and adapted TrIdent to infer selection parameters for identified adaptive candidates. Using whole-genome haplotype data from European and African populations, TrIdent effectively recapitulated known sweep candidates and identified novel cancer, and other disease-associated genes as potential sweeps.

Journal Article