PubMed HealthSearch

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.

OBJECTIVE: To quantify and improve the performance of standard rheumatoid arthritis (RA) algorithms in a biobank setting. METHODS: This retrospective cohort study within the Mayo Clinic (MC) Biobank and MC Tapestry Study identified RA cases by presence of at least two RA codes OR positive anti-cyclic citrullinated peptide antibodies (CCP) plus disease-modifying anti-rheumatic drug (DMARD) prescription as of 7/18/2022. Rheumatology physicians manually verified all RA cases using RA criteria and/or rheumatology physician diagnosis plus DMARD use. All other biobank participants served as non-RA controls. We defined seropositivity as rheumatoid factor and/or anti-CCP positivity. We assessed rules-based and Electronic Medical Records and Genomics (eMERGE) RA algorithms using positive predictive value (PPV). Finally, we developed a novel RA algorithm using a LASSO-based machine learning approach with five-fold cross validation. RESULTS: We identified 1,316 confirmed RA cases (968 MC Biobank, 348 Tapestry, 70 % seropositive) and 82,123 non-RA controls (mean age 65, 61 % female). The PPV of 3 RA codes was 43 %, codes plus DMARD was 54 %, and codes plus DMARD plus seropositivity was 85 %. The PPV of eMERGE was 77 %. Available in the MC Biobank, self-reported RA (PPV 10 %) only minimally improved algorithm performance (PPV from 83 % to 85 %), whereas family history of RA (PPV 3 %) worsened performance. At 90 % PPV, the novel RA algorithm incorporating key variables such as anti-CCP and DMARD use increased sensitivity by 4-11 % compared to eMERGE. CONCLUSION: Rules-based and eMERGE RA algorithms had worse performance in biobank than administrative settings. Our novel RA algorithm outperformed these standard algorithms.

Humans

Integrative machine learning and transcriptomic analysis reveals molecular mechanisms underlying low survival rate in larval Chinese Bahaba (Bahaba taipingensis).

Chinese Bahaba (Bahaba taipingensis) is a Class I protected marine fish endemic to China. Low larvae survival during artificial breeding severely hinder population recovery. To investigate the molecular mechanism of high mortality in larval fish, this study performed RNA-seq on liver from naturally deceased (ND) and mass-dead (MD) individuals, combined with least absolute shrinkage and selection operator (LASSO) regression and random forest (RF) algorithms to screen for core signature genes. A total of 873 differentially expressed genes (DEGs) were identified, including 112 upregulated and 761 downregulated genes. GO and KEGG enrichment analyses revealed significant enrichment in amino acid metabolism disorders, one‑carbon folate pool impairment, PPAR signaling abnormalities, ECM-receptor interaction, focal adhesion pathway, indicating widespread metabolic suppression accompanied by extracellular matrix remodeling and signaling disturbances in the livers of MD fish. MAD pre-filtering combined with dual machine learning algorithms yielded 18 robust core signature genes, among which SLC38A4, MMP1, FADD, FKBP5, and APOB were consistently identified as high-frequency core genes by both algorithms. SLC38A4 exhibited the highest importance score in the RF model and was significantly downregulated, making it the primary molecule distinguishing ND from MD phenotypes. ROC curve analysis showed that both models achieved an AUC of 1.000 (95% CI lower bound: 0.610), confirming the precise discriminatory ability of the core genes. GSEA further demonstrated significant enrichment of this core gene set in ND samples. This study provides the first systematic elucidation of the molecular mechanisms underlying liver dysfunction in low survival rate B. taipingensis, characterized by amino acid transport impairment, metabolic reprogramming, and structural remodeling, offering theoretical foundations for health assessment, early mortality risk warning, and artificial breeding conservation of this species.

Animals

PGS-GS: a framework integrating polygenic scores and genomic selection in animal breeding.

Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.

Animals

miRNA Target Prediction: An Overview of the Past and Current Tools.

MicroRNAs (miRNAs) are among the most studied molecules in recent years, and since their discovery, many miRNAs have been identified across various species. As members of the non-coding RNA family, miRNAs are key players in post-transcriptional gene regulation. These molecules can inhibit translation or promote degradation of messenger RNA (mRNA) by binding to the 3' untranslated region (UTR) of mRNA, thereby influencing almost all biological processes. To identify a miRNA's biological role, it is essential to predict the target sites to which it binds, a goal made possible through bioinformatics tools. This chapter discusses the bioinformatics tools commonly used for this purpose. Also, it analyzes the main factors considered in target prediction, such as seed match, free energy, conservation, site accessibility, multiple binding site contribution, and machine learning and deep learning approaches. Understanding the principles underlying these predictive methodologies is crucial for advancing one's biological research on miRNAs.

MicroRNAs

An inductive algorithm approach to knowledge acquisition for expert system development. A pilot study.

Knowledge acquisition, which consists of knowledge elicitation and knowledge representation, often is considered the weakest link in the design of expert systems. Systems frequently are built on the knowledge of one expert and require extensive use of knowledge engineering techniques to elicit this knowledge from the expert. Inductive algorithms are a potential alternative method of knowledge acquisition for expert system development. The aim of this pilot study was to examine the feasibility of applying machine learning techniques, specifically, inductive algorithms, to an existing research database as a method for knowledge elicitation and knowledge representation for expert system development. Two inductive algorithms (C4 and Classification and Regression Trees [CART]) that generate decision trees were selected for the analysis using a data set of 201 patients hospitalized for Pneumocystis carinii pneumonia. Neither C4 nor CART produced trees with an accuracy that was significantly better than the baseline accuracy (71.3%) for prediction of outcome in the data set. The mean accuracy of the C4 decision trees was below baseline and the mean accuracy of CART decision trees was 74.6%. The experts found both algorithms comprehensible, but not adequate, and identified important missing predictor variables. The study findings suggest that additional research is needed to examine the appropriate use of inductive algorithms in the transformation of nursing data and information into nursing knowledge.

Algorithms

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n&#xa0;=&#xa0;5), support vector machines (n&#xa0;=&#xa0;4), k-nearest neighbor (n&#xa0;=&#xa0;3), decision trees (n&#xa0;=&#xa0;3), random forests (n&#xa0;=&#xa0;5), neural networks (n&#xa0;=&#xa0;2), linear discriminant analysis (n&#xa0;=&#xa0;1), and pre-trained CNNs (n&#xa0;=&#xa0;1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n&#xa0;=&#xa0;12 to n&#xa0;=&#xa0;39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans

CCT2 defines a highly cisplatin-resistant and poor-prognosis subtype of lung adenocarcinoma.

Cisplatin-based chemotherapy is a standard treatment for lung adenocarcinoma (LUAD), yet acquired cisplatin resistance remains a marked cause of treatment failure. The molecular mechanisms driving cisplatin resistance in LUAD have not been fully elucidated. The present study integrated bulk transcriptomic data, genomic mutation profiles and single-cell RNA sequencing data to systematically investigate cisplatin resistance in LUAD. Resistance-associated genes were identified through differential expression, survival analysis and database integration. Unsupervised clustering was used to define cisplatin resistance-associated subtypes. Functional characteristics were explored using pathway enrichment, immune infiltration, tumor mutation burden and weighted gene co-expression network analysis. A machine learning framework incorporating 101 algorithms was applied to identify key genes and construct a prognostic model. Single-cell analyses and in vitro experiments were performed to validate the biological role of the core gene. Molecular docking and molecular dynamics simulations were conducted to identify potential therapeutic compounds. A total of two molecular subtypes with distinct cisplatin resistance levels and prognostic outcomes were identified. The high-resistance subtype exhibited enhanced cell cycle activity, DNA repair signaling and immune heterogeneity. Machine learning analysis revealed a five-gene signature, with chaperonin-containing TCP1 subunit 2 (CCT2) emerging as a key regulator of cisplatin resistance. Single-cell analyses showed that CCT2 was predominantly enriched in resistant epithelial cell subpopulations. Functional experiments demonstrated that CCT2 knockdown significantly inhibited cell proliferation and enhanced cisplatin sensitivity in LUAD cell lines. A number of candidate compounds targeting CCT2 exhibited stable binding in silico. The present findings identified CCT2 as a key mediator of cisplatin resistance in LUAD and provided potential therapeutic strategies to overcome chemotherapy resistance.

chaperonin-containing TCP-1 subunit 2

How not to be seen: predicting unseen enzyme functions using contrastive learning.

MOTIVATION: Predicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. RESULTS: Here we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AVAILABILITY AND IMPLEMENTATION: EnzPlacer and data is available at https://github.com/drxiangma/EnzPlacer under a GPL3 license.

Enzymes

Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning framework.

MOTIVATION: Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. RESULTS: This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. AVAILABILITY AND IMPLEMENTATION: The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L.

Peptides

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis.

MOTIVATION: Sexual dimorphism is a fundamental biological determinant driving systematic differences in disease susceptibility, progression, and clinical outcomes. However, current sex-combined AI-based genomic models often exhibit algorithmic bias and fail to capture these sex-specific mechanisms, creating a critical barrier to unbiased precision medicine. Ensuring fairness in the context of sexual dimorphism requires understanding and addressing the distinct biological mechanisms functioning in each sex, rather than focusing solely on equalizing predictive performance. RESULTS: We propose a fairness-aware supervised hierarchical contrastive learning approach, called FairHICON, to discover unbiased sex-common and sex-specific predictive features. Evaluations on cancer and asthma transcriptomic datasets demonstrate that FairHICON significantly outperforms state-of-the-art benchmarks, improving predictive performance by up to 9% while effectively reducing the performance gap between male and female sexes. Furthermore, prognostic validation confirms that the identified sex-specific pathways stratify patient survival significantly better within their corresponding sex groups. This validates FairHICON to elucidate the molecular heterogeneity of sexual dimorphism, advancing inclusive precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data is available at https://github.com/datax-lab/FairHICON.

Sex Characteristics

PharaCon: a new framework for identifying bacteriophages via conditional representation learning.

MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.

Bacteriophages

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Support vector machine classification of 18F-FDG PET scans across subtypes of amyotrophic lateral sclerosis.

PURPOSE: While 18F-FDG PET imaging has demonstrated diagnostic value in people with Amyotrophic Lateral Sclerosis (PwALS) and group-level differences were identified between different disease subtypes (e.g., genetic and clinical variants), refining and validating a machine-learning-based subject-level diagnostic algorithm may improve the general applicability and reliability of 18F-FDG PET as a diagnostic tool in ALS. In this study, we employed support vector machines (SVM) to further explore the diagnostic potential of 18F-FDG PET in ALS, alongside its ability to classify between different genetic subtypes or clinical phenotypes. METHODS: 18F-FDG PET data of 36 healthy volunteers (HV), 25 people with ALS-mimicking diseases (Mimics), and 167 PwALS, grouped by genetic status (e.g., sporadic (sALS) or carrying a C9orf72 hexanucleotide repeat expansion (ALSC9orf72RE) and onset (bulbar or spinal) type, acquired with Biograph 'TruePoint' PET/CT scanner, were included in the study (Dataset 1). A second dataset of 183 PwALS and 31 Mimics acquired with Biograph 'HiRez' scanner was included as an independent cross-validation set (Dataset 2). PET images were spatially normalised to MNI space to fit linear SVMs with cross-validation. Only age-matched groups were considered to eliminate age-related effects. RESULTS: For Dataset 1, the linear SVM resulted in an average accuracy of 0.86 for the classification of ALS vs. HV, 0.53 for ALS vs. Mimics, 0.83 for ALSC9orf72RE vs. sALS, and 0.58 for bulbar vs. spinal onset. These findings were corroborated with Dataset2, with an accuracy of up to 0.76 for ALSC9orf72RE vs. sALS, and 0.59 for bulbar vs. spinal. CONCLUSION: 18F-FDG brain PET imaging, combined with SVM and age-matching, can distinguish between ALSC9orf72RE and sALS with good accuracy, but lacks sufficient discriminative power to differentiate between ALS and Mimics and between different sites of onset.

Humans

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). &#xa9; 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans

Pattern recognition analysis of a set of mutagenic aliphatic N-nitrosamines.

A set of 21 mutagenic aliphatic N-nitrosamines were subjected to a pattern recognition analysis using ADAPT software. Four descriptors based on molecular connectivity, geometry and sigma charge on nitrogen were capable of achieving a 100% classification using the linear learning machine or iterative least squares algorithms. Three descriptors were capable of a 90.5% and two descriptors of a 85.7% overall correct classification. Three of the four descriptors were each capable of classifying 15 of the 16 active chemicals while it required three of the four descriptors to classify correctly two of the five inactive chemicals. These results are in concert with previous observations that molecular connectivity, geometry, and sigma charge on nitrogen are powerful descriptors for separating active from inactive mutagenic and carcinogenic N-nitrosamines.

Mutagens