PubMed HealthSearch

SEARCH · PubMed Health

Results for “machine learning feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Unveiling the power of TIIC: A prognostic tool for esophageal adenocarcinoma.

BACKGROUND: Esophageal adenocarcinoma (EAC) remains a lethal malignancy with limited prognostic tools for guiding immunotherapy. Tumor-infiltrating immune cells (TIICs) play a critical role in EAC prognosis and treatment response. METHODS: We integrated single-cell RNA sequencing and bulk transcriptome data from TCGA and GEO databases. TIIC-specific RNAs were identified via tissue specificity index calculation combined with machine learning feature selection. Twenty machine learning algorithms were benchmarked to construct an optimal TIIC signature score (TIIC-Score) based on the comprehensive C-index. Immunotherapy response, genomic mutation, and copy number variation were analyzed. Summary-data-based Mendelian randomization (SMR) and two-sample Mendelian randomization (MR) were performed to explore genetic associations. Core prognostic TIIC-related genes were functionally validated in esophageal cancer cell lines through loss-of-function assays. RESULTS: The TIIC-Score demonstrated robust prognostic value for 1-, 2-, and 3-year overall survival across multiple cohorts, outperforming 22 published models. High TIIC-Score was associated with poor survival and increased chromosomal instability. Mutation profiling revealed high frequencies of TP53 (78.2%), TTN (48.7%), and SYNE1 (30.8%). MR analysis identified a significant association between gastro-oesophageal reflux and EAC risk at SNP rs8130507. Functionally, CCNI was upregulated in esophageal cancer cells, and its knockdown suppressed malignant phenotypes while promoting apoptosis, supporting its pro-tumorigenic role. CONCLUSION: The TIIC-Score provides a novel prognostic framework for EAC that effectively stratifies patient risk and may help identify individuals most likely to benefit from immunotherapy.

Esophageal adenocarcinoma

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)&#x2500;a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC &#x2265; 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median &#x3c1; &#x223c; 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a &#x3b2;-strand than an &#x3b1;-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 &#x3bc;L of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies

Development and validation of a serum peptidomic signature for early detection of asymptomatic ovarian cancer: A multi-center prospective study.

Early detection of asymptomatic ovarian cancer (asym-OC) remains a critical challenge, the failure of which underlies its high mortality. Performing serum peptidomic profiling of 843 participants in the cohort SOCFCP, we distill 1,081 initial features into a 7-marker panel for asym-OC detection via a biology-informed machine-learning (ML)-based feature selection strategy. Three markers significantly revert toward non-OC levels after surgery. Integrating the panel with age, CA125, and HE4, we develop and externally validate (n = 159) a LightGBM model, ProMS+. For early-stage OC detection, ProMS+ shows a specificity of 92.6% at 95.0% sensitivity, outperforming CA125 (44.7%), HE4 (11.2%), and Risk of Ovarian Malignancy Algorithm (ROMA) (24.0%), with an area under the curve (AUC) of 0.993. In a simulated high-risk population (n = 100,000; OC prevalence = 1%), ProMS+ yields a high AUC (0.983) and a higher positive predictive value than CA125, HE4, and Age + CA125 + HE4 combined model (0.201 vs. 0.027, 0.090, and 0.064). ProMS+ offers a promising, non-invasive, and interpretable approach for the early detection of asym-OC.

Humans

Stage-specific ROMO1 in rheumatoid arthritis: predictive immune insights into the MIF pathway and HLA-DR/IL2RA axis via integrated GWAS, transcriptomic, single-cell, and spatial profiling.

Emerging evidence links reactive oxygen species modulator 1 (ROMO1), a key mitochondrial ROS regulator, to rheumatoid arthritis (RA) pathogenesis. However, its exact mechanism remains elusive given the conflicting evidence about its specific function. We used a four-level integrative framework combining multi-omics data and literature&#x2011;supported mechanistic inference. At the genetic level, Mendelian randomization (MR) was performed to explore potential causal relationships between ROMO1, IL2RA, HLA-DR, MIF, and RA risk, followed by differential expression analysis and machine learning-based feature selection to identify key mROS genes. The temporal expression dynamics of ROMO1 were assessed in RA progression. At the cellular and tissue levels, we integrated single-cell RNA sequencing and spatial transcriptomics to map cell-type-specific expression and synovial localization of ROMO1-related immune cells and pathways. Finally, our multi-omics findings were contextualized with literature-supported mechanistic inference. (1) MR results were consistent with a potential protective effect of ROMO1 on RA (OR&#x2009;=&#x2009;0.52) and its potential regulation of risk factors IL2RA (OR&#x2009;=&#x2009;0.46) and HLA-DR (OR&#x2009;=&#x2009;0.40). Conversely, IL2RA (OR&#x2009;=&#x2009;1.42), HLA-DR (OR&#x2009;=&#x2009;1.88), and MIF (OR&#x2009;=&#x2009;1.17) were positively associated with RA risk. Additionally, ROMO1 was identified as a top candidate diagnostic predictor with stage-specific dynamics: downregulated in the early but upregulated in the late/remission stages. (2) Single-cell RNA sequencing showed ROMO1's cell-specific expression in CD14+&#x2009;HLA-DR+&#x2009;CD74+&#x2009;monocytes and CD4+&#x2009;IL2RA+&#x2009;T cells. Cell communication analysis further suggested that these cells may participate in MIF pathway regulation. Spatial transcriptomics subsequently identified that ROMO1-related cells localized to synovial pathological regions, with MIF pathway changes correlated with RA progression. (3) Finally, literature-supported mechanistic inference suggests that ROMO1 may modulate mROS levels to promote anti-inflammatory M2 macrophage polarization, which could theoretically contribute to reduced systemic inflammation and the alleviation of multi-organ decline in RA. This integrated multi-omics investigation, supported by literature-based mechanistic inference, suggests ROMO1 as a stage-dependent biomarker candidate and potential immune regulator in RA.

Humans

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-&#x3b1;/NF-&#x3ba;B, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95&#xa0;% CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in&#xa0;gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n&#x2009;=&#x2009;42, Healthy: n&#x2009;=&#x2009;54), Franzosa et al. (IBD: n&#x2009;=&#x2009;164, Healthy: n&#x2009;=&#x2009;56), and Yachida et al. (CRC: n&#x2009;=&#x2009;150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans

Quality over quantity: biopsy-anchored CT radiogenomics models outperform all-lesion training in a multi-tumour cohort despite a smaller sample size.

OBJECTIVE: Radiogenomics aims to non-invasively predict tumour genotypes from imaging, but most studies assume molecular homogeneity by assigning a single biopsy-derived label to all lesions within a patient. This approach risks substantial label noise given well-documented interlesional heterogeneity. We investigated whether anchoring training to biopsy-confirmed lesions improves radiogenomic model performance and generalisability. MATERIALS AND METHODS: We retrospectively analysed 1646 patients (11473 segmented lesions) with contrast-enhanced CT and EGFR mutation status from next-generation sequencing at the Netherlands Cancer Institute, alongside an external NSCLC radiogenomics cohort (n&#x2009;=&#x2009;158). All visible lesions were segmented, and the exact biopsy site was matched to its segmentation. Radiomic features were extracted, and machine learning models were trained with three lesion selection strategies: all lesions, non-biopsied lesions only, and biopsy-confirmed lesions only. To disentangle label quality from sample size, we created size-matched variants (one lesion per patient) for all-lesion and non-biopsied strategies. RESULTS: All models achieved significant discrimination of EGFR status on internal validation (AUC&#x2009;=&#x2009;0.62-0.68). However, performance of the all-lesion and non-biopsied models declined on external validation (AUC&#x2009;=&#x2009;0.55-0.63), while the biopsy-anchored model maintained stable performance (AUC&#x2009;=&#x2009;0.62), despite having only 1/10th of the training sample size. When training sets were size-matched, the biopsy-anchored approach significantly outperformed a model trained on all available lesions on external validation (p&#x2009;=&#x2009;0.037). CONCLUSIONS: Radiogenomic models trained on biopsy-confirmed lesions outperform conventional all-lesion strategies in external validation, despite using an order of magnitude fewer samples. Prioritising lesion-level label fidelity can mitigate heterogeneity-driven noise, enhancing robustness and clinical translation of imaging-based genomic prediction. KEY POINTS: Question Does assigning biopsy-derived molecular labels to all lesions introduce heterogeneity-driven label noise that reduces the generalisability of radiogenomic models? Findings Models trained exclusively on biopsy-confirmed lesions demonstrated superior external generalisability compared with all-lesion approaches, despite being trained on substantially fewer samples. Clinical relevance Biopsy-anchored radiogenomics improves the reliability of non-invasive mutation prediction by accounting for tumour heterogeneity, potentially supporting clinical decision-making when tissue sampling is limited or molecular results are discordant across lesions.

Humans

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics&#x2013;based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics&#x2013;based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans