PubMed HealthSearch

SEARCH · PubMed Health

Results for “deep learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Phylogenetic Methods Meet Deep Learning.

Deep learning (DL) has been widely used in various scientific fields, but its integration into phylogenetics has been slower, primarily due to the complex nature of phylogenetic data. The studies that apply DL to sequencing data often limit analyses to four-taxon trees. Many of these studies serve as "proof of principle" and perform similarly to traditional phylogeny reconstruction methods. New ways of using training data, such as encoding with compact bijective ladderized vectors or transformers, enable the handling of much larger trees and genomic data sets. This short perspective focuses on the application of DL in phylogenetics, introducing prevalent DL architectures. We highlight potential problems in the field by discussing the risks of using simulation-based training data and emphasize the importance of reproducibility and robustness in computational estimates. Finally, we explore promising research areas, including the combination of phylogenetics and population genetics in DL, the analysis of neighbor dependencies, and the potential to significantly reduce computational cost compared to traditional methods. This perspective illustrates the potential of DL in complementing traditional phylogeny reconstruction methods and aiding the advancement of phylogenetic analysis, especially in performing computationally demanding tasks such as model selection or estimating branch support values.

Humans

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

Towards mechanistic models of mutational effects: Deep learning on Alzheimer's Aβ peptide.

Deep Mutational Scanning (DMS) has enabled multiplexed measurement of mutational effects on protein properties, including kinematics and self-organization, with unprecedented resolution. However, potential bottlenecks of DMS characterization include experimental design, data quality, and depth of mutational coverage. Here, we apply deep learning to comprehensively model the mutational effect of the Alzheimer's Disease associated peptide Aβ42 on aggregation-related biochemical traits from DMS measurements. Among tested neural network architectures, Convolutional Neural Networks and Recurrent Neural Networks are found to be the most cost-effective models with high performance even under insufficiently-sampled DMS studies. While sequence features are essential for satisfactory prediction from neural networks, geometric-structural features further enhance the prediction performance. Notably, we demonstrate how mechanistic insights into phenotype may be extracted from the neural networks themselves suitably designed. This methodological benefit is particularly relevant for biochemical systems displaying a strong coupling between structure and phenotype such as the conformation of Aβ42 aggregate and nucleation, as shown here using a Graph Convolutional Neural Network (GCN) developed from the protein atomic structure input. In addition to accurate imputation of missing values (which here ranged up to 55% of all phenotype values at key residues), the mutationally-defined nucleation phenotype generated from a GCN shows improved resolution for identifying known disease-causing mutations relative to the original DMS phenotype. Our study suggests that neural network derived sequence-phenotype mapping can be exploited not only to provide direct support for protein engineering or genome editing but also to facilitate therapeutic design with the gained perspectives from biological modeling.

Alzheimer's disease

Development and Validation an Integrated Deep Learning Model to Assist Eosinophilic Chronic Rhinosinusitis Diagnosis: A Multicenter Study.

BACKGROUND: The assessment of eosinophilic chronic rhinosinusitis (eCRS) lacks accurate non-invasive preoperative prediction methods, relying primarily on invasive histopathological sections. This study aims to use computed tomography (CT) images and clinical parameters to develop an integrated deep learning model for the preoperative identification of eCRS and further explore the biological basis of its predictions. METHODS: A total of 1098 patients with sinus CT images were included from two hospitals and were divided into training, internal, and external test sets. The region of interest of sinus lesions was manually outlined by an experienced radiologist. We utilized three deep learning models (3D-ResNet, 3D-Xception, and HR-Net) to extract features from CT images and calculate deep learning scores. The clinical signature and deep learning score were inputted into a support vector machine for classification. The receiver operating characteristic curve, sensitivity, specificity, and accuracy were used to evaluate the integrated deep learning model. Additionally, proteomic analysis was performed on 34 patients to explore the biological basis of the model's predictions. RESULTS: The area under the curve of the integrated deep learning model to predict eCRS was 0.851 (95% confidence interval [CI]: 0.77-0.93) and 0.821 (95% CI: 0.78-0.86) in the internal and external test sets. Proteomic analysis revealed that in patients predicted to be eCRS, 594 genes were dysregulated, and some of them were associated with pathways and biological processes such as chemokine signaling pathway. CONCLUSIONS: The proposed integrated deep learning model could effectively predict eCRS patients. This study provided a non-invasive way of identifying eCRS to facilitate personalized therapy, which will pave the way toward precision medicine for CRS.

Humans

Expert opinion elicitation for assisting deep learning based Lyme disease classifier with patient data.

BACKGROUND: Diagnosing erythema migrans (EM) skin lesion, the most common early symptom of Lyme disease, using deep learning techniques can be effective to prevent long-term complications. Existing works on deep learning based EM recognition only utilizes lesion image due to the lack of a dataset of Lyme disease related images with associated patient data. Doctors rely on patient information about the background of the skin lesion to confirm their diagnosis. To assist deep learning model with a probability score calculated from patient data, this study elicited opinions from fifteen expert doctors. To the best of our knowledge, this is the first expert elicitation work to calculate Lyme disease probability from patient data. METHODS: For the elicitation process, a questionnaire with questions and possible answers related to EM was prepared. Doctors provided relative weights to different answers to the questions. We converted doctors' evaluations to probability scores using Gaussian mixture based density estimation. We exploited formal concept analysis and decision tree for elicited model validation and explanation. We also proposed an algorithm for combining independent probability estimates from multiple modalities, such as merging the EM probability score from a deep learning image classifier with the elicited score from patient data. RESULTS: We successfully elicited opinions from fifteen expert doctors to create a model for obtaining EM probability scores from patient data. CONCLUSIONS: The elicited probability score and the proposed algorithm can be utilized to make image based deep learning Lyme disease pre-scanners robust. The proposed elicitation and validation process is easy for doctors to follow and can help address related medical diagnosis problems where it is challenging to collect patient data.

Humans

The Use of Deep Learning in RNA Therapeutic Development.

Ribonucleic acid (RNA)-based therapeutics have emerged as promising methods of disease treatment due to their ability to target the human genome and influence protein production, their versatility, and their relative lack of toxicity compared to other gene therapies. However, the RNA therapeutic design space is extremely large, encompassing multiple variables, including codon identities, secondary structure, and design of specific regions. RNA therapeutic optimization is difficult due to the impracticality of exploring such a vast design space experimentally. To address this limitation, deep learning methods have been employed to optimize RNA therapeutic development. In this review, we examine the application of deep learning models across three key aspects of RNA therapeutic development (RNA structure prediction, CRISPR activity, and RNA delivery), highlighting major contributions in these fields and analyzing how deep learning model architectures could affect model performance. We then discuss challenges associated with using deep learning for RNA therapeutics, such as computational and data limitations. Finally, we offer perspectives on areas for future exploration, such as emerging model architectures and methods of integration with more advanced high-throughput screening techniques. Ultimately, this review provides an overview of how deep learning is used in RNA therapeutic development and how it can evolve in the future.

Deep Learning

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans

Development and validation of a deep learning model based on cascade mask regional convolutional neural network to noninvasively and accurately identify human round spermatids.

INTRODUCTION: The difficulty of identifying human round spermatids (hRSs) has impeded applications of the human round spermatid injection (ROSI) technique. RSs can be accurately screened through flow cytometric analysis utilizing the Hoechst fluorescence profile reflecting DNA, but this method is not suitable for isolating hRSs due to the toxicity associated with Hoechst staining. OBJECTIVE: To evaluate the capacity of a deep learning model grounded in a cascade mask region-based convolutional neural network (R-CNN) for the noninvasive and accurate identification of hRSs. METHODS: In this study, we presented the development and validation of a deep learning model for identifying hRSs through the analysis of 3457 optical light microscope images of sorted hRSs obtained via flow cytometric analysis. The model's accuracy and specificity were evaluated by calculating the mean average precision (mAP). Furthermore, a double-blind experiment was conducted to access the reliability of the proposed model in accurately identifying hRSs. It detected the expression of protamine (PRM1) and/or peanut lectin (PNA), which are established markers for RSs. RESULTS: Our deep learning-based model demonstrated a high precision, achieving a mAP of over 0.80 for isolating hRSs in test datasets. The expression of PRM1 and/or PNA was observed in all cells noninvasively selected by our AI model during an independent double-blind test. This phenomenon confirmed the accuracy and effectiveness of the proposed model. The model's capability for noninvasive and accurate isolation of hRSs among spermatogenic cells highlighted its robustness and generalizability for clinical applications. CONCLUSION: The deep learning AI model based on a cascade R-CNN has the ability to accurately identify hRSs among spermatogenic cells. The application of this noninvasive method, which requires no additional procedures in clinical practice, is able to facilitate the widespread implementation of ROSI technique. Therefore, it can provide patients with spermatogenic arrest the opportunity to become biological fathers.

Humans

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans

Prediction of Atrial Fibrillation From the ECG in the Community Using Deep Learning: A Multinational Study.

BACKGROUND: We aimed to refine and validate a deep neural network model from the ECG to predict atrial fibrillation (AF) risk, using samples from diverse backgrounds: the Framingham Heart Study (FHS), UK Biobank, and Estudo Longitudinal da Saúde do Adulto (ELSA-Brasil). We compared the model's performance to the clinical Cohorts for Heart and Aging Research in Genomic Epidemiology consortium (CHARGE-AF) risk score and evaluated the association with other cardiovascular outcomes. METHODS: The ECG-derived deep-learning prediction of AF (ECG-AF) model was refined using 60% of FHS samples free of AF. Its performance was then tested in the remaining FHS samples, UK Biobank, and ELSA-Brasil, with discrimination assessed by the area under the receiver operating characteristic curve. The association of ECG-AF with cardiovascular outcomes was assessed using Cox proportional hazards models. RESULTS: The study sample included 10 097 FHS participants (mean age 53±12 years; 54.9% women), 49 280 participants from the UK Biobank (mean age 64±8 years, 47.9% women), and 12 284 participants from ELSA-Brasil (mean age 53±8 years, 54.7% women). The ECG-AF model showed moderate discrimination for incident AF (area under the curve, 0.82 [95% CI, 0.80-0.84]) in the FHS, comparable to the CHARGE-AF score (area under the curve, 0.83 [95% CI, 0.81-0.85]), and incremental when combined (area under the curve, 0.85 [95% CI, 0.83-0.87]). In UK Biobank and ELSA-Brasil, combining ECG-AF and CHARGE also improved prediction. Higher ECG-AF scores were associated with increased risks of heart failure, myocardial infarction, stroke, and all-cause mortality in all 3 cohorts. CONCLUSIONS: In multinational cohort studies, the single-input ECG-AF deep neural network model demonstrated good performance in predicting AF and other cardiovascular outcomes, comparable to a multivariable clinical risk score, with improved performance when combined.

Humans

EPIPDLF: a pretrained deep learning framework for predicting enhancer-promoter interactions.

MOTIVATION: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene expression, homeostasis, and disease development across various biological processes. With advancing research, it has been uncovered that distal enhancers may engage with nearby promoters to modulate the expression of target genes. This discovery holds significant implications for deepening our comprehension of various biological mechanisms. In recent years, numerous high-throughput wet-lab techniques have been created to detect possible interactions between enhancers and promoters. However, these experimental methods are often time-intensive and costly. RESULTS: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF, which utilizes advanced deep learning techniques to predict EPIs based solely on genomic sequences in an interpretable manner. Comparative evaluations across six benchmark datasets demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally, by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned features, aiding in the identification and biological analysis of important sequences. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/xzc196/EPIPDLF.

Deep Learning

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Deep learning techniques in predicting BRAF mutation status in cutaneous melanoma from histopathologic images.

AIMS: To develop and validate a deep learning framework for discriminating BRAF mutation status in cutaneous melanoma from routine H&E whole-slide images (WSIs) as a proof-of-concept complementary approach alongside molecular testing. METHODS: We built a two-stage pipeline comprising U-Net-based tumour segmentation followed by an Inception v3 classifier. In total, 272 institutional melanoma cases with confirmed BRAF status were used for model development (training and internal validation). Generalisability was assessed in an external test set of 76 cutaneous melanoma cases from the Cancer Genome Atlas (TCGA). Dermatopathologist-defined tumour-rich regions of interest were used to train and evaluate segmentation. WSIs were processed at 20×magnification using 512×512 tiles; slide-level mutation probabilities were obtained by averaging the predicted probabilities across all tumour-enriched tiles. RESULTS: Inception v3 achieved area under the receiver operating characteristic curve values of 0.973 (training), 0.954 (validation) and 0.915 (TCGA testing) and outperformed a ResNet50 baseline, showing stable external generalisation. Performance remained robust in advanced pathological T-category primary tumours (pT3-T4). Tumour probability heatmaps supported spatial interpretability by localising regions contributing most strongly to predicted mutation status. CONCLUSIONS: Deep learning applied to routine H&E WSIs can infer BRAF mutation status in cutaneous melanoma with consistent performance across institutional and external cohorts. Given the observed external sensitivity and negative predictive value, the model is not suitable for rule-out use or for deferring/omitting molecular testing. Any workflow integration is future work and would require prospective validation and calibration of probability outputs in real-world clinical series.

Artificial Intelligence

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000 cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT > 2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared

Deep Learning on Histologic Slides Accurately Predicts Consensus Molecular Subtypes and Spatial Heterogeneity in Colon Cancer.

Colon cancer (CC) is the third most prevalent cancer type. It is highly heterogeneous, particularly in terms of molecular profiles, which have both prognostic and predictive impacts on the treatment efficacy. However, CC treatment in adjuvant situations is currently guided solely by T and N staging. In this context, consensus molecular subtypes (CMSs) were introduced to stratify patients with CC based on molecular profiles. Recent studies have shown that CMS can be heterogeneous in CC, leading to a worse prognosis. This study focused on predicting CMS and its heterogeneity in CC using deep learning on digitized hematoxylin and eosin ± saffron-stained whole-slide images. Data and whole-slide images of 1996 patients from the PETACC-8, The Cancer Genome Atlas-COAD, and PRODIGE-13 cohorts were used. The model is trained to predict a 4-dimensional CMS vector, reflecting intratumor heterogeneity (ITH). It comprises a self-supervised model for embedding image patches into vectors and a weakly supervised model predicting CMS calls. Ground-truth CMS scores are obtained with the CMSclassifier package. Interpretability analyses are performed at the slide and patch levels. For homogeneous tumors, the model trained on PETACC-8 achieves 93.0% (±1.4%) macroaverage area under the curve in internal cross-validation and 94.4% macroaverage area under the curve in external validation over PRODIGE-13, whereas the The Cancer Genome Atlas-COAD model reaches 85.4% (±3.0%) in cross-validation and 92.4% over PRODIGE-13. The trained models also provide spatial distributions of CMS across tumor slides and associate specific histologic features with each CMS. Finally, the models are able to predict ITH. The results show that a deep learning model trained on routine histology slides is capable of providing an efficient and robust method for predicting CMS and characterizing a patient's ITH, paving the way for the routine consideration of CMS/ITH in clinical decision making in the adjuvant setting.

Humans

RCoxNet: A Deep Learning Framework Integrating Random Walk with Restart, Mutation, and Clinical Data for Cancer Survival Prediction.

Accurate survival prediction in cancer remains challenging due to the sparsity of somatic mutation profiles and the failure of existing models to capture higher-order gene-gene dependencies. Network diffusion methods such as Random Walk with Restart (RWR) can propagate mutation signals across protein-protein interaction (PPI) networks to address sparsity, yet their integration within a deep learning Cox survival framework has not been comprehensively benchmarked across multiple cancer cohorts. We present RCoxNet, a deep learning framework that maps somatic mutation profiles onto a ConsensusPathDB-derived PPI network via RWR, selects prognostic genes by log-rank filtering, and processes network-informed mutation scores through three fully connected hidden layers feeding into a Cox proportional hazards output. RCoxNet was evaluated on The Cancer Genome Atlas (TCGA) cohorts for four cancer types (breast invasive carcinoma [BRCA], lung adenocarcinoma [LUNG], glioblastoma multiforme [GBM], and ovarian serous cystadenocarcinoma [OV]) using 20 independent random splits. The model achieved mean C-index values of 0.807 ± 0.044 (BRCA), 0.750 ± 0.039 (LUNG), 0.704 ± 0.041 (GBM), and 0.668 ± 0.036 (OV), consistently outperforming DeepSurv, Cox-nnet, SurvivalNet, Cox Elastic-Net (Cox-EN), and DeepHit, with statistically significant gains over Cox-EN, Cox-nnet, SurvivalNet, and DeepHit across the majority of cohorts. RCoxNet demonstrates that embedding sparse mutation profiles into a PPI network context substantially improves cancer survival prediction and yields biologically interpretable prognostic features relevant to precision oncology.

cancer survival prediction