PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article

Machine learning-driven spleen imaging and genomics uncover a splenic connection to coronary artery disease.

Despite advances in managing traditional risk factors, coronary artery disease (CAD) remains the leading cause of mortality. Circulating hematopoietic cells influence risk for CAD separately from traditional risk factors, but the role of a key regulating organ, the spleen, is unknown. The understudied spleen is a representation of the hematopoietic system optimally suited for unbiased radiologic investigations toward mechanistic insights. Here, we leveraged deep learning to extract 107 splenic radiomic features from abdominal magnetic resonance imaging (MRI) scans of 42,059 UK Biobank participants and of 2745 Mass General Brigham Biobank (MGBB) participants. Of these, 10 features from UK Biobank were associated with CAD. Genome-wide association analysis of CAD-associated features identified 219 loci, including 9p21. Variants at 9p21, the strongest yet mechanistically elusive CAD locus, were associated with splenic features such as run-length nonuniformity, reflecting heterogeneity of continuous texture regions. Research MRI findings were consistent internally, but external clinical validation highlighted challenges in translating analyses of abdominal MRI scans to routine clinical practice because of variability in imaging protocols and greater clinical heterogeneity among patients. Our study, combining deep learning with genomics, presents a framework to uncover potential splenic involvement in CAD and emphasizes translational gaps between research and clinical radiomics.

Humans

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan–Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans

A CFH- and SPINT2-based prognostic signature for cholangiocarcinoma.

BACKGROUND: Cholangiocarcinoma (CCA) is a highly malignant tumor with a poor prognosis, and reliable biomarkers for postoperative risk stratification remain limited. This study aimed to develop and validate a CFH- and SPINT2-based prognostic signature to support postoperative risk stratification and inform adjuvant therapy selection in CCA through integrative machine learning and single-cell transcriptomics. METHODS: Differentially expressed genes were screened from GSE26566. Integrative machine learning (least absolute shrinkage and selection operator-Cox, random forest, and univariate Cox regression) was performed in the training cohort (GSE89749; n=115) to construct a risk model, which was externally validated in two independent cohorts: cohort 1 (E-MTAB-6389; n=75) and cohort 2 [The Cancer Genome Atlas Cholangiocarcinoma (TCGA-CHOL) data set; n=36]. Systematic analysis was conducted and included examinations of immune infiltration [via single-sample gene set enrichment analysis (ssGSEA)], pathway enrichment (via hallmark GSEA), cellular localization (via single-cell RNA sequencing), and drug sensitivity (via the Genomics of Drug Sensitivity in Cancer 2 database). RESULTS: Two genes, CFH and SPINT2, were identified and incorporated into a prognostic risk score. High-risk patients in the training cohort had a significantly worse overall survival (log-rank P=0.02). External validation was performed in two independent cohorts. In validation cohort 1, the risk group was an independent prognostic factor [hazard ratio =2.27, 95% confidence interval (CI): 1.18-4.37; P=0.01]. In validation cohort 2, the model demonstrated acceptable discriminative ability (concordance index =0.721; 3-year area under the curve =0.692). The high-risk group exhibited an immunosuppressive microenvironment characterized by increased infiltration of macrophages and myeloid-derived suppressor cells, along with the activation of epithelial-mesenchymal transition, inflammatory response, and NF-κB signaling pathways. Single-cell analysis revealed a cell-type-specific expression pattern: CFH was predominantly expressed in fibroblasts, while SPINT2 was mainly expressed in malignant cells. Drug sensitivity analysis demonstrated that the high-risk group was more sensitive to gemcitabine, cisplatin, poly(ADP-ribose) polymerase (PARP) inhibitors, and mammalian target of rapamycin (mTOR) inhibitors, whereas the low-risk group was more sensitive to lapatinib. CONCLUSIONS: The CFH- and SPINT2-based prognostic signature may serve as an independent biomarker for postoperative risk stratification in CCA. High-risk patients, characterized by fibroblast-derived CFH enrichment and malignant-cell SPINT2 loss, exhibit an immunosuppressive microenvironment and may be more suitable for gemcitabine-based chemotherapy or PARP/mTOR inhibitors, whereas low-risk patients may benefit from less intensive adjuvant strategies or HER2/EGFR-targeted lapatinib. Prospective validation is warranted before clinical implementation.

Cholangiocarcinoma (CCA)

Integrative multi-omics analysis unravels the metabolic landscape and reveals serum biomarkers for early diagnosis of hyperuricemia.

BACKGROUND: Hyperuricemia (HUA) is a major risk factor for gout and multiple metabolic disorders. Although serum uric acid (UA) is the gold standard for HUA diagnosis, it fails to reflect early metabolic disturbances and shows limited predictive value for asymptomatic HUA. This study sought to elucidate the pathological mechanisms underlying HUA and identify novel diagnostic biomarkers beyond UA. METHODS: This study enrolled 195 patients with HUA and 98 healthy controls. Global metabolomics and proteomics profiling were performed to characterize molecular alterations underlying HUA. Based on the biological relevance of the shared dysregulated pathways, a pathway correlation network was constructed to elucidate the pathological mechanisms driving HUA initiation and progression. Furthermore, diagnostic biomarkers for HUA were identified using machine learning algorithms, and were validated with an external cohort. RESULTS: HUA patients exhibited distinct metabolic and proteomic profiles compared with healthy controls. Integrated multi-omics pathway analysis revealed that peroxisome proliferators-activated receptor signaling pathway, arachidonic acid metabolism, purine metabolism, pyrimidine metabolism and sphingolipid signaling pathway were significantly dysregulated in HUA. Among them, arachidonic acid metabolism was identified as a hub pathway involved in HUA progression. Furthermore, a metabolite panel consisting of cysteine-S-sulfate, glycerophosphocholine and 4-hydroxyphenylpyruvic acid was screened by machine learning and validated in an independent cohort, which showed slightly higher diagnostic performance for HUA than UA. CONCLUSIONS: This study reveals the core metabolic and protein regulatory networks of HUA, and identifies a novel serum metabolite panel for the diagnosis of HUA. These findings provide new insights for improved clinical diagnosis and management.

Humans

Serum Proteomic Profiling Implicates a Dysregulated Neurohormonal-Inflammatory Axis in Post-Fontan Sinus Tachycardia.

BACKGROUND: Postoperative sinus tachycardia is a poorly understood complication following the Fontan procedure. The molecular signaling cascades triggering acute tachycardia remain uncharacterized, limiting therapeutic innovation. Here, we present a retrospective study leveraging serum proteomics and machine learning to identify the molecular drivers of postoperative Fontan sinus tachycardia. METHODS: We integrated a clinically relevant ovine Fontan model with continuous telemetric heart rate monitoring and human patient data. Serum proteomics coupled with least absolute shrinkage and selection operator and Boruta machine learning algorithms were used to identify protein panels predictive of postoperative sinus tachycardia. Cross-species validation was performed by comparing proteomic signatures from sheep and pediatric patients undergoing Glenn or Fontan surgery. RESULTS: Ovine Fontan animals demonstrated significant heart rate elevation beginning on postoperative day 1, peaking at postoperative day 3 (159.4±11.7 bpm versus preoperative, 105.3±10.5 bpm; P=0.0002), before trending toward baseline by postoperative day 10. This pattern was mirrored in human patients with a more modest magnitude. Surgical controls did not exhibit tachycardia. The principal component most correlated with heart rate (principal component 1: r=0.78, P=2.2×10-4) was enriched for inflammatory and neural pathways. The Boruta algorithm identified an 11-protein panel with strong predictive power (area under the receiver operating characteristic curve, 0.963). Cross-species comparison demonstrated that angiotensinogen, angiotensin-converting enzyme, and pentraxin 3 were similarly dysregulated in both species postoperatively. CONCLUSIONS: This study provides molecular evidence implicating a dysregulated neurohormonal-inflammatory axis in acute postoperative Fontan sinus tachycardia and establishes a foundation for developing targeted diagnostics and therapeutics for this complication.

Animals

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n = 42, Healthy: n = 54), Franzosa et al. (IBD: n = 164, Healthy: n = 56), and Yachida et al. (CRC: n = 150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans

Toward multistrategy parallel and distributed learning in sequence analysis.

Machine learning techniques have been shown to be effective in sequence analysis tasks. However, current learning algorithms, which are typically serial main-memory-based, are not capable of handling the vast amounts of information being generated by the Human Genome Project. The multistrategy parallel learning approach presented in this paper is an attempt to scale existing learning algorithms. Learning speed is improved through running multiple learning processes in parallel and prediction accuracy is improved through multiple learners. Our approaches are independent of the learning algorithms used. This paper focuses on one of the MSPL approaches and preliminary empirical results that we present are encouraging.

Algorithms

The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis.

BACKGROUND: Current endoscopic retrograde cholangiopancreatography (ERCP) and cholangioscopic-based diagnostic sampling for indeterminant biliary strictures remain suboptimal. Artificial intelligence (AI)-based algorithms by means of computer vision in machine learning have been applied to cholangioscopy in an effort to improve diagnostic yield. The aim of this study was to perform a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based diagnostic performance of AI-associated cholangioscopic diagnosis of indeterminant or malignant biliary strictures. METHODS: Individualized searches were developed in accordance with PRISMA and MOOSE guidelines, and meta-analysis according to Cochrane Diagnostic Test Accuracy working group methodology. A bivariate model was used to compute pooled sensitivity and specificity, likelihood ratio, diagnostic odds ratio, and summary receiver operating characteristics curve (SROC). RESULTS: Five studies (n=675 lesions; a total of 2,685,674 cholangioscopic images) were included. All but one study analyzed a deep learning AI-based system using a convoluted neural network (CNN) with an average image processing speed of 30 to 60 frames per second. The pooled sensitivity and specificity were 95% (95% CI: 85-98) and 88% (95% CI: 76-94), with a diagnostic accuracy (SROC) of 97% (95% CI: 95-98). Sensitivity analysis of CNN studies (4 studies, 538 patients) demonstrated a pooled sensitivity, specificity, and accuracy (SROC) of 95% (95% CI: 82-99), 88% (95% CI: 72-95), and 97% (95% CI: 95-98), respectively. CONCLUSIONS: Artificial intelligence-based machine learning of cholangioscopy images appears to be a promising modality for the diagnosis of indeterminant and malignant biliary strictures.

Humans

Instrumented Walkway Gait Analysis Predicts Fallers in Neurological Disorders: Identifying Digital Biomarkers for Balance Monitoring.

Assessing balance is crucial in neurological rehabilitation, yet while wearable sensors enable real-world monitoring, identifying reliable digital biomarkers remains challenging. This study utilized a high-fidelity instrumented walkway to determine which gait parameters best predict balance impairment, providing robust targets for future wearable applications. We analyzed 49 steady-state gait metrics from 140 individuals with diverse neurological conditions. Using statistical analysis and machine learning, we evaluated these parameters against objective force plate sway scores and clinical fall-history labels. Group analysis identified 16 parameters significantly distinguishing fallers from non-fallers, and a neural network classified fallers with an area under the curve of 0.75. Across all analytical approaches, overall gait variability, e.g., Stride Width S.D. and the Gait Variability Index, emerged as a universal predictor of balance impairment and fall risk. Furthermore, while traditional linear models emphasized spatial postural control, machine learning classification uniquely identified inter-limb asymmetry as a premier driver of fall prediction. These findings indicate that instrumented gait analysis effectively identifies digital biomarkers for balance deficits. Isolating these specific metrics provides a clear blueprint for meaningful metrics required for continuous objective monitoring and future development of personalized, adaptive rehabilitation strategies.

Humans

NanoSSL: attention mechanism-based self-supervised learning method for protein identification using nanopores.

MOTIVATION: Nanopores are cutting-edge interdisciplinary tools that can analyze biomolecules at the single-molecule level for many applications, e.g. DNA sequencing. Efforts are underway to extend nanopores to proteomics, including the development of machine learning algorithms for protein sequencing and identification. However, single-molecule data are intrinsically noisy and hard to process. Moreover, the development and performance of machine learning for nanopore is jeopardized by data scarcity. Self-supervised learning is an emerging method that may yield advantages in nanopore scenarios. RESULTS: We propose and experimentally validate Nanopore analysis using Self-Supervised Learning (NanoSSL), a generative self-supervised learning framework based on attention mechanisms for the identification of protein signals from nanopores. Leveraging a two-step approach consisting of self-supervised pre-training and supervised fine-tuning, NanoSSL learns useful feature representations from empirical data to facilitate downstream classification tasks. Inspired by the concept of fragmentation in conventional protein sequencing technologies, during pretraining each translocation event is split into multiple non-overlapping fragments of equal size, some of which are randomly masked and reconstructed using a masked autoencoder. Learning the feature representations of the reconstructed nanopore events facilitates molecular identification in fine-tuning. In this study, we retested a publicly available nanopore multiplexed protein sensing dataset for model iteration, and subsequently measured Alzheimer's disease biomarker Aβ1-42 using homemade solid-state nanopores. Empirical results indicated NanoSSL achieved an unprecedented performance across four metrics: accuracy, precision, recall, and F1 score, when classifying two mutated Aβ1-42, E22G and G37R. The self-supervised learning and attention mechanism were verified as the source of performance gains. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://doi.org/10.5281/zenodo.17172822.

Nanopores

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans