PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning explanation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning↗

Spatial concordance metrics and related risk factors of brain-peripheral barrier axes: unveiling distinct concordance patterns for mental and neurological axes.

Numerous studies have documented bidirectional interactions between the central nervous system and barrier organs (skin, gut, and lung). While genome-wide association studies have revealed shared genetic factors across brain-peripheral barrier axes, investigating these connections from an environmental perspective in large populations remains difficult. Using data from the Global Burden of Disease (GBD) 2023, I extracted annual incidence rates for 56 diseases related to brain-peripheral barrier axes and exposure rates for the 70 most detailed risk factors across 204 countries and territories. By categorizing regional incidence rates into four quartiles for each disease, I pinpointed regions with concordance of these axes and constructed a spatial atlas of disease concordance within the brain-peripheral barrier axis from a macro-epidemiologic view. Subsequently, I calculated global spatial concordance percentages for each axis, which allowed the comparatively assessment of concordance patterns across different axes, specific diseases, and their variations over time, across the lifespan, and by gender. Finally, I applied machine learning models and Shapley additive explanations to identify risk factors related to the spatial concordance of each axis. From 1990 to 2023, the overall trend for most brain-peripheral barrier axis pairs remained stable. Spatial concordance patterns showed dynamic fluctuations across the lifespan, followed by a convergence toward stability in older age. Several risk factors are related to most brain-peripheral barrier axes. Notably, the mental and neurological axes exhibited distinct concordance patterns. Compared with neurological axes, concordance within mental axes showed a broader and more dispersed geographic distribution, with greater variation across sexes and over time. Furthermore, concordance percentages of mental and neurological axes exhibited opposing age-related trends, contrasting disease spectra for peripheral conditions, and inverse relationships with alcohol and sodium consumption. Those divergences suggest distinct mechanisms underlying the brain-peripheral barrier axes in mental and neurological diseases. Related risk factors offer population-based hypotheses for further investigation in individual-level studies.

Humans↗

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans↗

KG-Microbe: Building modular and scalable knowledge graphs for microbiome and microbial sciences.

BACKGROUND: The integration of many disparate forms of data is essential for understanding the microbial world and its interaction with the environment and human health. Doing so is particularly challenging in the context of microbe-host and microbe-microbe interactions that contribute to health or environmental outcomes. There are thousands of relevant microbial species, and millions of interactions among those microbes and with their environment or host. Integrated information (e.g., about host and microbial physiology, genetics, and metabolism) facilitates deeper understanding of complex mechanisms and helps interpret correlative results. RESULTS: The KG-Microbe construction framework is a novel approach to harmonizing bacterial and archaeal data in the form of a findable, accessible, interoperable, reusable and AI-ready knowledge graph (KG). Starting from a core KG with organismal traits, environments, and growth preferences and the integration of established ontologies, the framework generates a hierarchy of related KGs targeting specific use cases, including the human microbiome in the context of disease, or environmental microbiomes. The framework supports customizable taxa subsets representing communities or clades of interest. Evaluations of the KG-Microbe KGs through a series of competency questions demonstrate the accuracy and effectiveness of the data harmonization, and the utility of the resulting KGs in studies of inflammatory bowel disease and Parkinson's disease. Finally, the predictive and environmental capabilities of the KGs are demonstrated by predicting growth preferences using graph features. CONCLUSIONS: The KG-Microbe framework unifies microbial contexts in a single resource to support integrative analyses across biomedical, host, and environmental domains. KG-Microbe is a flexible, modular enabling technology for humans and machine learning methods to uncover candidate mechanistic explanations of microbial associations.

Microbiota↗

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans↗

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi‑omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans↗

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms↗

A machine learning method for extracting symbolic knowledge from recurrent neural networks.

Neural networks do not readily provide an explanation of the knowledge stored in their weights as part of their information processing. Until recently, neural networks were considered to be black boxes, with the knowledge stored in their weights not readily accessible. Since then, research has resulted in a number of algorithms for extracting knowledge in symbolic form from trained neural networks. This article addresses the extraction of knowledge in symbolic form from recurrent neural networks trained to behave like deterministic finite-state automata (DFAs). To date, methods used to extract knowledge from such networks have relied on the hypothesis that networks' states tend to cluster and that clusters of network states correspond to DFA states. The computational complexity of such a cluster analysis has led to heuristics that either limit the number of clusters that may form during training or limit the exploration of the space of hidden recurrent state neurons. These limitations, while necessary, may lead to decreased fidelity, in which the extracted knowledge may not model the true behavior of a trained network, perhaps not even for the training set. The method proposed here uses a polynomial time, symbolic learning algorithm to infer DFAs solely from the observation of a trained network's input-output behavior. Thus, this method has the potential to increase the fidelity of the extracted knowledge.

Algorithms↗

Rule generation for protein secondary structure prediction with support vector machines and decision tree.

Support vector machines (SVMs) have shown strong generalization ability in a number of application areas, including protein structure prediction. However, the poor comprehensibility hinders the success of the SVM for protein structure prediction. The explanation of how a decision made is important for accepting the machine learning technology, especially for applications such as bioinformatics. The reasonable interpretation is not only useful to guide the "wet experiments," but also the extracted rules are helpful to integrate computational intelligence with symbolic AI systems for advanced deduction. On the other hand, a decision tree has good comprehensibility. In this paper, a novel approach to rule generation for protein secondary structure prediction by integrating merits of both the SVM and decision tree is presented. This approach combines the SVM with decision tree into a new algorithm called SVM_ DT, which proceeds in three steps. This algorithm first trains an SVM. Then, a new training set is generated through careful selection from the output of the SVM. Finally, the obtained training set is used to train a decision tree learning system and to extract the corresponding rule sets. The results of the experiments of protein secondary structure prediction on RS126 data set show that the comprehensibility of SVM_DT is much better than that of the SVM. Moreover, the generalization ability of SVM_DT is better than that of C4.5 decision trees and is similar to that of the SVM. Hence, SVM_DT can be used not only for prediction, but also for guiding biological experiments.

Algorithms↗

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation↗

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans↗

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article↗

Learning Petri net models of non-linear gene interactions.

Understanding how an individual's genetic make-up influences their risk of disease is a problem of paramount importance. Although machine-learning techniques are able to uncover the relationships between genotype and disease, the problem of automatically building the best biochemical model or "explanation" of the relationship has received less attention. In this paper, I describe a method based on random hill climbing that automatically builds Petri net models of non-linear (or multi-factorial) disease-causing gene-gene interactions. Petri nets are a suitable formalism for this problem, because they are used to model concurrent, dynamic processes analogous to biochemical reaction networks. I show that this method is routinely able to identify perfect Petri net models for three disease-causing gene-gene interactions recently reported in the literature.

Algorithms↗

Proteome Analyst: custom predictions with explanations in a web-based tool for high-throughput proteome annotations.

Proteome Analyst (PA) (http://www.cs.ualberta.ca/~bioinfo/PA/) is a publicly available, high-throughput, web-based system for predicting various properties of each protein in an entire proteome. Using machine-learned classifiers, PA can predict, for example, the GeneQuiz general function and Gene Ontology (GO) molecular function of a protein. In addition, PA is currently the most accurate and most comprehensive system for predicting subcellular localization, the location within a cell where a protein performs its main function. Two other capabilities of PA are notable. First, PA can create a custom classifier to predict a new property, without requiring any programming, based on labeled training data (i.e. a set of examples, each with the correct classification label) provided by a user. PA has been used to create custom classifiers for potassium-ion channel proteins and other general function ontologies. Second, PA provides a sophisticated explanation feature that shows why one prediction is chosen over another. The PA system produces a Naïve Bayes classifier, which is amenable to a graphical and interactive approach to explanations for its predictions; transparent predictions increase the user's confidence in, and understanding of, PA.

Internet↗

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Likelihood linkage analysis (LLA) classification method: an example treated by hand.

This paper describes a very general method of data analysis using a hierarchical classification. The data can be provided by observation, experiment or knowledge; their nature can be numerical, qualitative or logical. First, the classical view of the context of data representation, in which the algorithm of hierarchical ascendant construction of the classification tree is set, is treated in a synthetic manner. The main notion in our method is one of 'similarity'. This must be elaborated in the best way, taking into account the mathematical nature of the objects to be compared. Here we adopt a set of theoretical and combinatorial representation of the descriptive attributes, which are interpreted in terms of relations. Then we introduce a probability scale for similarity measurement by using a likelihood concept. The largest part of the paper concerns an illustrating example, moderately sized, detailing very minutely the different steps and the different calculations assumed by the method. The data structure handled with this example is the simplest possible. Then, general aspects and methodological extensions are evoked. We end by indicating the interest of the described approach in future works, in which we are involved, concerning typological organization of genetic sequences. We emphasize the 'explanation' aspect of the obtained results, with respect to a given description. For this purpose, classifications (on the object set and on the attribute set) on the one hand and machine learning techniques on the other, intervene efficiently.

Algorithms↗

Combining the performance strengths of the logistic regression and neural network models: a medical outcomes approach.

The assessment of medical outcomes is important in the effort to contain costs, streamline patient management, and codify medical practices. As such, it is necessary to develop predictive models that will make accurate predictions of these outcomes. The neural network methodology has often been shown to perform as well, if not better, than the logistic regression methodology in terms of sample predictive performance. However, the logistic regression method is capable of providing an explanation regarding the relationship(s) between variables. This explanation is often crucial to understanding the clinical underpinnings of the disease process. Given the respective strengths of the methodologies in question, the combined use of a statistical (i.e., logistic regression) and machine learning (i.e., neural network) technology in the classification of medical outcomes is warranted under appropriate conditions. The study discusses these conditions and describes an approach for combining the strengths of the models.

Artificial Intelligence↗