PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Random forests”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Prediction of genomewide conserved epitope profiles of HIV-1: classifier choice and peptide representation.

Identification of peptides binding to Major Histocompatibility Complex (MHC) molecules is important for accelerating vaccine development and improving immunotherapy. Accordingly, a wide variety of prediction methods have been applied in this context. In this paper, we introduce (tree-based) ensemble classifiers for such problems and contrast their predictive performance with forefront existing methods for both MHC class I and class II molecules. In addition, we investigate the impact of differing peptide representation schemes on performance. Finally, classifier predictions are used to conduct genomewide scans of a diverse collection of HIV-1 strains, enabling assessment of epitope conservation. We investigated all combinations of six classification methods (classification trees, artificial neural networks, support vector machines, as well as the more recently devised ensemble methods (bagging, random forests, boosting) with four peptide representation schemes (amino acid sequence, select biophysical properties, select quantitative structure-activity relationship (QSAR) descriptors, and the combination of the latter two) in predicting peptide binding to an MHC class I molecule (HLA-A2) and MHC class II molecule (HLA-DR4). Our results show that the ensemble methods are consistently more accurate than the other three alternatives. Furthermore, they are robust with respect to parameter tuning. Among the four representation schemes, the amino acid sequence representation gave consistently (across classifiers) best results. This finding obviates the need for feature selection strategies incurred by use of biophysical and/or QSAR properties. We obtained, and aligned, a diverse set of 32 HIV-1 genomes and pursued genomewide HLA-DR4 epitope profiling by querying with respect to classifier predictions, as obtained under each of the four peptide representation schemes. We validated those epitopes conserved across strains against known T-cell epitopes. Once again, amino acid sequence representation was at least as effective as using properties. Assessment of novel epitope predictions awaits experimental verification.

Journal Article↗

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure↗

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control.

OBJECTIVE: Spinal meningiomas (SMs) are common primary spinal tumors for which surgery is considered the first-line treatment when safe and feasible. The ability to extrapolate the tumor grade from preoperative imaging may significantly inform early patient expectation-setting regarding recurrence. Building on radiomics studies in cranial meningiomas, the authors aimed to construct a benchmark radiomics model to preoperatively identify the histological grade of SMs. METHODS: Institutional surgical records from May 2012 to November 2025 were queried for pathology-confirmed meningiomas below the foramen magnum, with preoperative contrast-enhanced imaging available for segmentation. SMs were classified as low-grade (WHO grade 1) and high-grade (WHO grade 2 tumors and grade 1 tumors with atypia). Tumors were manually segmented, and features were extracted using the PyRadiomics software package. An ensemble model of k-nearest neighbors, random forest, and support vector machine classifiers was trained using nested cross-validation on a subset of 10 features to differentiate tumor grades. Clinical data for the cohort were also extracted, and disease control in an adjunctive clinical series was assessed. RESULTS: Seventy-four patients were included in radiomics analysis, with an area under the receiver operating characteristic curve of 0.879 and a mean F1 score of 0.748. The model's top 5 features were all texture features that differed significantly (p < 0.05) across low- and high-grade SMs. These included measures of tumor textural and contrast-enhancement heterogeneity, with overlap with features reported in radiomics models for histological grading of intracranial meningiomas. Fifty-five patients with a median radiographic follow-up of 22.2 (range 1.9-86.4) months remained for clinical analysis after exclusion of patients with less than 1 month of follow-up and syndromic meningiomas. Four recurrences occurred at a median of 20.8 (range 1.8-41.8) months. High-grade tumor pathology did not significantly impact progression-free survival (p = 0.682, log-rank test; Cox regression high vs low grade hazard ratio [HR] 0.62, 95% CI 0.06-6.11, p = 0.685). Subtotal resection was associated with poorer progression-free survival than gross-total resection (p = 0.004, log-rank test; Cox regression subtotal vs gross-total resection HR 10.62, 95% CI 1.46-77.05, p = 0.019). These findings remain contextualized within a relatively limited follow-up window and small recurrence event count, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs. CONCLUSIONS: A preoperative radiomics model can stratify high-grade SMs using open-source tools applied to single-institution data.

Humans↗

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans↗

Integrated multi-omics identification of m6A-SNP-related diagnostic biomarkers in amyotrophic lateral sclerosis.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) lacks reliable and minimally invasive biomarkers for early diagnosis. m6A-associated single-nucleotide polymorphisms (m6A-SNPs) may influence RNA methylation and gene expression, offering opportunities to identify clinically relevant diagnostic markers. METHODS: We integrated eQTLGen cis-eQTL data, RMVar m6A-SNP annotations, and ALS transcriptomic datasets to identify m6A-SNP-related genes. Random Forest and LASSO regression were combined to screen robust diagnostic markers. A nomogram was constructed and validated using independent cohorts. Immune infiltration, predicted m6A modification sites, and potential RBP-SNP interactions were assessed. Peripheral blood samples from ALS patients were used for exploratory validation of gene expression and global m6A levels. RESULTS: We identified 109 ALS-associated m6A-SNP-related genes with cis-eQTL signals and narrowed these to seven candidate diagnostic markers (TMED5, OXR1, BRI3, FEM1C, SUZ12, EIF2AK4, and TJAP1). The seven-gene model outperformed the individual markers in the training cohort and retained moderate discrimination in the independent validation cohort. ALS samples showed differences in inferred immune-cell composition, including monocytes, neutrophils, and T-cell subsets. The selected SNP loci were located near predicted m6A sites and annotated RBP-binding regions. Exploratory clinical validation showed significant upregulation of FEM1C and SUZ12 at both mRNA and protein levels, accompanied by reduced global m6A modification. CONCLUSIONS: Through multi-omics integration and exploratory clinical validation, this study identifies m6A-SNP-related candidate markers associated with ALS. The findings support further evaluation of m6A-related signatures for ALS discrimination and molecular characterization, while larger independent cohorts and additional calibration are required before clinical application.

Humans↗

Metabolomic Responses to Oral Glucose Tolerance Test and Hyperinsulinemic-euglycemic Clamp in CKD.

BACKGROUND: The oral glucose tolerance test (OGTT) captures integrated physiological responses involving intestinal glucose absorption, incretin signaling, and endogenous insulin secretion, whereas the hyperinsulinemic-euglycemic clamp (clamp) isolates insulin-mediated glucose uptake. Comparing plasma metabolomic responses to these two challenges may identify processes specific to intestinal nutrient delivery and how they vary in CKD. METHODS: Targeted plasma metabolomics was performed in 59 adults without diabetes (39 with CKD [eGFR <60 mL/min/1.73 m2] and 20 controls) from the Study of Glucose and Insulin in Renal Disease (SUGAR). Each participant underwent a 75-g OGTT and clamp approximately one week apart. Eighty-eight plasma metabolites were quantified at fasting and during each challenge. Metabolite levels were log-transformed and normalized using Systematic Error Removal Using Random Forest (SERRF). Metabolites were classified using adjusted regression slopes relating OGTT and clamp responses. RESULTS: The mean (SD) age and eGFR were 64 (13) years and 54 (26) mL/min/1.73 m2, respectively, and 41% were female. In the overall cohort, OGTT and clamp induced broad plasma metabolic changes, with 63 (72%) and 76 (86%) metabolites significantly altered from fasting, respectively. Seventy-three metabolites (83%) demonstrated a significant relationship between OGTT and clamp responses. Of these, 22 (25%) exhibited true concordance and 51 (58%) demonstrated similar directional changes but differed in magnitude. A total of 15 (17%) metabolites were discordant or non-corresponding, of which only three were discordant. The non-corresponding metabolites were enriched in amino acid metabolism. Eleven metabolites (13%) demonstrated differential responses between OGTT and clamp by CKD status, involving amino acid and glucose metabolism pathways. CONCLUSIONS: Metabolomic responses to OGTT and clamp were largely directionally concordant but differed in magnitude, with attenuation during OGTT. Discordant metabolites were rare, while non-corresponding metabolites were confined to amino acid pathways. CKD modified OGTT-clamp correspondence for metabolites involved in amino acid and glycolytic metabolism.

Journal Article↗

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance↗

A predictor based on the somatic genomic changes of the BRCA1/BRCA2 breast cancer tumors identifies the non-BRCA1/BRCA2 tumors with BRCA1 promoter hypermethylation.

The genetic changes underlying in the development and progression of familial breast cancer are poorly understood. To identify a somatic genetic signature of tumor progression for each familial group, BRCA1, BRCA2, and non-BRCA1/BRCA2 (BRCAX) tumors, by high-resolution comparative genomic hybridization, we have analyzed 77 tumors previously characterized for BRCA1 and BRCA2 germ line mutations. Based on a combination of the somatic genetic changes observed at the six most different chromosomal regions and the status of the estrogen receptor, we developed using random forests a molecular classifier, which assigns to a given tumor a probability to belong either to the BRCA1 or to the BRCA2 class. Because 76.5% (26 of 34) of the BRCAX cases were classified with our predictor to the BRCA1 class with a probability of >50%, we analyzed the BRCA1 promoter region for aberrant methylation in all the BRCAX cases. We found that 15 of the 34 BRCAX analyzed tumors had hypermethylation of the BRCA1 gene. When we considered the predictor, we observed that all the cases with this epigenetic event were assigned to the BRCA1 class with a probability of >50%. Interestingly, 84.6% of the cases (11 of 13) assigned to the BRCA1 class with a probability >80% had an aberrant methylation of the BRCA1 promoter. This fact suggests that somatic BRCA1 inactivation could modify the profile of tumor progression in most of the BRCAX cases.

BRCA1 Protein↗

Struct2net: integrating structure into protein-protein interaction prediction.

UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the structure-based interaction prediction technique threads these two sequences to all the protein complexes in the PDB and then chooses the best potential match. Based on this match, structural information is incorporated into logistic regression to evaluate the probability of these two proteins interacting. This paper also describes a random forest classifier which can effectively combine the structure-based prediction results and other functional annotations together to predict protein interactions. Experimental results indicate that the predictive power of the structure-based method is better than many other information sources. Also, combining the structure-based method with other information sources allows us to achieve a better performance than when structure information is not used. We also tested our method on a set of approximately 1000 yeast genes and, interestingly, the predicted interaction network is a scale-free network. Our method predicted some potential interactions involving yeast homologs of human disease-related proteins. SUPPLEMENTARY INFORMATION: http://theory.csail.mit.edu/struct2net

Algorithms↗

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient&#x2009;=&#x2009;-&#x2009;1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10&#xa0;years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung↗

Aggregation of individual trees and patches in forest succession models: capturing variability with height structured, random, spatial distributions.

Individual based, stochastic forest patch models have the potential to realistically describe forest dynamics. However, they are mathematically intransparent and need long computing times. We simplified such a forest patch model by aggregating the individual trees on many patches to height-structured tree populations with theoretical random dispersions over the whole simulated forest area. The resulting distribution-based model produced results similar to those of the patch model under a wide range of conditions. We concluded that the height- structured tree dispersion is an adequate population descriptor to capture the stochastic variability in a forest and that the new approach is generally applicable to any patch model. The simplified model required only 4.1% of the computing time needed by the patch model. Hence, this new model type is well-suited for applications where a large number of dynamic forest simulations is required.

Bias↗

Applying GPS to the study of primate ecology: a useful tool?

Data on the spatiotemporal distribution of resources can be collected and plotted using GPS (global positioning system) and GIS (geographical information system) technologies. By combining such data with information on foraging and ranging behavior of nonhuman primates, one can analyze the influence of resource distribution on social organization and group cohesion. We investigated the abilities of a three-channel GPS receiver to collect location data under varying canopy densities in both temperate and tropical forests. Eighty randomly selected points were sampled in a beech-maple forest in northeast Ohio, USA; 65 points also were sampled at several tropical forests in Costa Rica and Trinidad. At each point we attempted to obtain a GPS position fix; we also determined the speed of satellite acquisition and measured canopy density using a spherical densiometer. The ability to obtain a reading differed greatly between the two forest types (chi(2) = 53.79, P < 0.001). Ninety-seven percent of all attempts were successful in the temperate forest, whereas only a 34% acquisition rate was obtained in the tropical forests. Logistic regression showed that the probability of obtaining a reading in Neotropical forests was 75% but only when canopy cover was less than 20%. Thus, these minimal-channel GPS units may be of limited utility for behavioral ecologists working in closed-canopy Neotropical forests.

Animals↗

Metaanalysis of the accuracy of rapid prescreening relative to full screening of pap smears.

BACKGROUND: Efficient quality assurance and improvement measures are essential ingredients in a well organized cytology-based program for cervical carcinoma screening. Various pap smear review procedures, aiming for optimization of accuracy, are described throughout the literature. Evaluation and synthesis of those methods are needed. In a previous study, we pooled data on the diagnostic quality of rapid reviewing (RR) of cervical smears initially reported as normal or unsatisfactory. We now focus on rapid prescreening (RPS) of unreported smears. METHODS: Six published studies on the accuracy of RPS relative to subsequent full screening were pooled using metaanalytic methods. Individual and pooled sensitivity, specificity, and predictive values were assessed using forest plots. Random effect pooling methods were used for interstudy heterogeneity. Variation in sensitivity according to influencing factors was explored by metaregression. RESULTS: The pooled average sensitivity of RPS was 64.9% (95% confidence interval [CI] 50.7-79.1%) for all abnormalities, 72.6% (95% CI 60.6-85.2%) for low-grade lesions or more severe, and 85.7% (95% CI 77.8-93.6%) for high-grade lesions or more severe. The pooled specificity was estimated at 96.8% (CI 95.8-97.8%). The sensitivity increased significantly with duration of screening and decreased with workload. Almost 3% of all abnormal slides were detected only by RPS (2.8%; CI 0.0-5.8%). This is comparable to the proportion of false-negative smears detectable by RR. CONCLUSIONS: Rapid prescreening has a high yield for severe dysplasia and shows diagnostic properties that support its use as a quality control procedure in cytologic laboratories. We showed previously that RR is superior to full reviewing of a 10% random sample of negative slides (10% FR). Because the yield of additional abnormalities found by RR and RPS is comparable, we expect RPS to be more efficient than 10% FR as well.

False Negative Reactions↗

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans↗

Dynamic evolution of chaperone-mediated autophagy is associated with tumor microenvironment remodeling and prognostic stratification in lung adenocarcinoma: insights from single-cell transcriptomics, ensemble machine learning, and experimental validation.

BACKGROUND: Lung adenocarcinoma (LUAD) shows prognostic heterogeneity, and tumor-node-metastasis (TNM) staging is limited for individualized management. Chaperone-mediated autophagy (CMA) maintains proteostasis, but its role during adenocarcinoma in situ (AIS)-minimally invasive adenocarcinoma (MIA)-invasive adenocarcinoma (IAC) progression remains unclear. METHODS: Single-cell RNA sequencing (scRNA-seq) data from GSE189357 and bulk transcriptomes from The Cancer Genome Atlas (TCGA)-LUAD and Gene Expression Omnibus (GEO) cohorts were integrated. CMA activity, cell-cell communication, weighted gene co-expression network analysis (WGCNA), tumor-normal differential expression, machine-learning survival modeling, tumor microenvironment (TME) features, drug sensitivity, and EPC1 function were analyzed. RESULTS: CMA-high tumor epithelial cells increased from AIS (58.1%) to MIA (65.7%) but declined in IAC (44.4%; p < 0.001). CMA-low cells preferentially received fibroblast-derived extracellular matrix cues. A CMA-negatively correlated module identified 69 core genes. Random survival forest (RSF) performed best among 117 machine-learning combinations (mean concordance index > 0.873). High-risk patients had worse survival across cohorts, and the risk score was independently associated with overall survival (hazard ratio = 16.013, 95% confidence interval: 9.579-26.768, p < 0.001). High-risk tumors showed proliferative activation and M0 macrophage enrichment, whereas low-risk tumors showed stronger immune-related signaling. EPC1 overexpression suppressed malignant phenotypes in A549 cells. CONCLUSION: CMA dynamics are associated with stromal and immune remodeling during LUAD progression. A CMA-based model provides robust prognostic stratification and may offer a basis for future TME-guided studies.

Chaperone-mediated autophagy↗

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Integrating single-cell transcriptomics to construct an oncogene-driven prognostic model and elucidate metabolic-immune crosstalk in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related deaths, its progression and treatment heterogeneity are mainly influenced by driver gene and tumor micro-environment (TME) interactions. Nevertheless, the mechanisms of this process at the single-cell level remain unclear. This study integrated TCGA and multi-center single-cell transcriptome data to identify a 575 genes HCC-specific core set, developing a single-cell "oncogene scoring" system to quantify individual carcinogenic activity. This score is significantly elevated in malignant and proliferative T cells and is closely associated with metabolic reprogramming, aberrant cell&#x2012;cell communication, and immunosuppressive phenotypes. Based on these characteristics, we constructed a machine learning-based Random Survival Forest (RSF) prognostic model validated in multiple independent cohorts, which classifies patients into distinct risk subtypes. The high-risk group exhibits genomic instability, increased tumor stemness, and immune evasion, while the low-risk group was more sensitive to drugs such as sorafenib. This study highlights the potential pathways by which high oncogenic activity is associated with HCC progression, suggesting a profound link with single-cell metabolic&#x2012;immune crosstalk. The constructed RSF model offers a promising computational framework for risk stratification and provides hypothesis-generating insights that may inform future personalized treatment strategies for HCC patients.

Hepatocellular carcinoma↗

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 &#xb1; 0.0994, with a Log-rank testp-value of 1.6553&#xd7;10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 &#xb1; 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 &#xb1; 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 &#xb1; 0.1211) and discrete-time survival models such as DeepHit (0.7655 &#xb1; 0.1041) and Nnet-surv (0.7694 &#xb1; 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 &#xb1; 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 &#xb1; 0.0818) and Multimodal Co-Attention Transformer (0.8102 &#xb1; 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell↗