PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Accelerate Your Science: Direct-to-Biology Strategies in Medicinal Chemistry.

Direct-to-biology (D2B) is a powerful strategy that accelerates early drug discovery. It enables compounds to be synthesized in miniaturized formats and evaluated directly as crude reaction mixtures. This bypasses the need for purification during the initial design-make-test cycle. Advances in robust synthetic methodologies, automation, reaction miniaturization, and biological screening have transformed D2B from a proof-of-concept approach into a versatile medicinal chemistry platform. This platform is applicable to fragment optimization, covalent ligands, macrocycles, proteolysis-targeting chimeras (PROTACs), molecular glues, and cellular phenotypic screening. This perspective focuses on the synthetic transformations, assay technologies, and platform implementations that drive modern D2B workflows. It emphasizes reaction robustness, assay compatibility, and practical implementation. Analysis of the current literature revealed that D2B is more governed by reaction reliability than synthetic diversity. Amide coupling and click chemistry dominate reported workflows, while more complex transformations remain underexplored. We discuss the complementary strengths and limitations of biochemical, biophysical, and cellular readouts, identify current bottlenecks in reaction scope and data management, and highlight emerging opportunities arising from reaction miniaturization, machine learning, automated experimentation, and advanced synthetic methodologies. Rather than replacing conventional medicinal chemistry, D2B fundamentally shifts experimental effort from purification toward early biological validation and is poised to become an integral component of future medicinal chemistry workflows.

Humans↗

Bioactive peptides for meat quality and preservation: Integrating peptidomics and computational screening.

Bioactive peptides generated from meat proteins, fermented meat products, and slaughter by-products have attracted increasing attention as functional molecules for improving meat quality and preservation. In meat systems, peptides can be produced through endogenous postmortem proteolysis, microbial fermentation, gastrointestinal digestion, or controlled enzymatic hydrolysis of underutilized animal by-products. These peptides are closely associated with key meat science endpoints, including postmortem tenderization, oxidative stability, color retention, flavor development, microbial inhibition, and the valorization of processing by-products. However, although high-resolution peptidomics has greatly expanded the identification of meat-derived peptide sequences, their translation into practical meat applications remains limited by matrix interactions, processing stability, sensory constraints, safety concerns, and insufficient validation in real meat systems. This review synthesizes recent advances in meat-related peptidomics and computational screening, including sequence-based prediction, machine learning, molecular docking, molecular dynamics, stability assessment, and safety-oriented filtering. Particular attention is given to how these approaches can prioritize peptides with antioxidant, antimicrobial, flavor-modulating, and preservation-related functions under meat-specific technological constraints. By integrating peptide generation pathways, mass spectrometry-based identification, in silico prioritization, and meat quality endpoints, this review proposes a stage-gated framework for translating meat-derived bioactive peptides from discovery to application. Future research should strengthen matrix-specific validation, standardized peptidomic reporting, and safety assessment to support the use of bioactive peptides in meat quality improvement, clean-label preservation, and circular utilization of meat industry by-products.

Animals↗

Decoding cancer with artificial intelligence: Transforming research, diagnosis, and therapy with future insights.

Cancer remains one of the leading global health burdens, with increasing complexity in genomic, imaging, and clinical datasets presenting significant challenges for effective management. Artificial intelligence (AI) has emerged as a powerful tool to address these challenges by enabling pattern recognition, knowledge integration, and data-driven decision-making. This review highlights recent advances in the application of AI across cancer research, diagnosis, and therapy. In research, AI accelerates drug discovery and repurposing, enhances genomic data interpretation, and facilitates biomarker identification through multi-omics integration. In diagnosis, AI has demonstrated high technical performance in radiology for lesion detection and image segmentation, in pathology for tumour grading and molecular prediction, and in liquid biopsy for non-invasive biomarker analysis. In therapy, AI supports precision medicine by predicting treatment responses, monitoring disease progression, and optimizing clinical trial design. Despite these advances, barriers such as data heterogeneity, algorithmic bias, interpretability, and regulatory challenges remain. Future directions, including explainable AI, federated learning, multimodal modelling, and digital twins, hold promise for translating AI-driven innovations into routine oncology practice. Significance Statement This review provides a timely synthesis of recent (2020-2025) advances in artificial intelligence across cancer research, diagnosis, and therapy, highlighting applications in drug discovery, genomics, multi-omics biomarker identification, and clinical decision-making. By integrating technological progress with translational and clinical relevance, this work serves as a valuable resource for bridging AI innovation with precision oncology practice. As a narrative review, the literature was identified through targeted PubMed, Scopus, and Google Scholar searches, combining terms for artificial intelligence, machine learning, and deep learning with cancer-related keywords, with priority given to peer-reviewed studies published between 2020 and 2025, seminal earlier works, and official regulatory or guideline documents. Within each domain, representative studies were selected to illustrate methodological diversity, clinical context, and current translational readiness rather than to provide exhaustive coverage of an extremely rapidly evolving field.

Artificial intelligence↗

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

Natural language processing-based model to predict radiation pneumonitis in patients with locally advanced non-small cell lung cancer undergoing chemoradiotherapy: a retrospective cohort study.

BACKGROUND: Radiation pneumonitis (RP) remains a significant treatment-related toxicity in patients with unresectable, locally advanced non-small cell lung cancer (NSCLC) undergoing chemoradiotherapy (CRT). Most existing predictive models rely on static baseline demographic or dosimetry variables and lack real-time clinical applicability. We developed a novel predictive framework that integrates longitudinal symptom data extracted from clinical notes using natural language processing (NLP) with clinical and dosimetry features to improve early RP prediction. METHODS: We retrospectively identified 227 patients with locally advanced NSCLC treated with definitive CRT at a high-volume cancer center in the United States. We included all patients older than 18 years who were diagnosed between Jan 1, 2006, and Dec 31, 2022 with histologically or cytologically confirmed unresectable Stage 2 or 3 NSCLC and treated with conformal radiotherapy to a minimum dose of &#x2265;45 Gy with or without chemotherapy. Of these, 31 RP events were identified through manual adjudication using radiologic criteria and chart review. NLP was used to extract the temporal relationship of 16 pre-specified symptoms with treatment from over 100,000 clinical notes spanning pre- and during-treatment intervals. We trained and validated machine learning models on combinations of baseline clinical data, radiation dosimetry, and NLP-derived symptom features. Model performance was evaluated using a nested cross-validation framework, with an outer cross-validation loop reserved for performance assessment and an inner cross-validation loop used for model training and integration, and summarized using area under the receiver operating characteristic curve (AUC) and partial AUC (pAUC) at high specificity thresholds. Clinical utility was evaluated using decision curve analysis (DCA). FINDINGS: The best-performing model incorporated longitudinal NLP features and achieved a median AUC of 0.759 (90% confidence interval 0.753-0.766), significantly outperforming baseline models using only dosimetry (AUC 0.613) or clinical variables (AUC 0.635). NLP-based features such as cough trajectory, shortness of breath, and wheezing were among the most important predictors. Inclusion of NLP-derived symptom data improved early identification of high-risk patients, particularly in the clinically relevant high-specificity range (pAUC 0.021 vs. 0.010 for dosimetry alone). DCA showed that the calibrated MLP model provided greater net benefit than default strategies of treating all or no patients across clinically relevant threshold possibilities. INTERPRETATION: In this early work, NLP-based extraction of longitudinal symptoms from routine clinical documentation meaningfully enhances RP prediction in patients undergoing CRT for NSCLC. This approach leverages existing electronic health record infrastructure to deliver real-time, scalable, and interpretable risk estimates, offering a pathway toward potential early intervention and personalized toxicity management. The model and DCA requires external and prospective validation before clinical deployment; as such, future work should focus on this validation and integration into clinical decision support systems. FUNDING: AstraZeneca.

Chemoradiotherapy↗

Emerging genes implicated in human congenital heart disease: a 2023-2025 scoping review.

BACKGROUND: Congenital heart disease (CHD) is the most common major congenital anomaly and a leading cause of infant morbidity and mortality. The rapid expansion of genomic technologies has accelerated the discovery of rare genetic variants implicated in CHD pathogenesis. However, most individuals with CHD still lack an identifiable molecular etiology. The purpose of this scoping review is to systematically characterize genes reported in the recent literature as candidate CHD-associated genes and contextualize these findings within the stages of cardiac morphogenesis. METHODS: PubMed was searched using predefined terms related to CHD and genetic variants, supplemented by a prospectively maintained internal database. We included human studies published between January 2023 and December 2025 that identified pathogenic, likely pathogenic, or uncertain monogenic variants in at least one patient with CHD. Animal-only studies, chromosomal abnormalities, copy number variants, multigenic associations, transcriptomic/proteomic analyses, reviews, and maternal-only genetic studies were excluded. Gene-disease validity classifications were assigned using the Clinical Genome Resource (ClinGen) CHD Gene Curation Expert Panel framework. RESULTS: Of 2,834 screened articles, 391 studies met inclusion criteria, identifying 912 unique genes reported as candidate CHD-associated genes. Frequently reported genes included PTPN11, NOTCH1, GATA4, JAG1, MYH6, GATA6, and LZTR1. Identified genes spanned all major stages of cardiogenesis, including developmental priming, cardiac progenitor specification, left-right axis formation, neural crest migration, outflow tract development, septation, and postnatal structural remodeling. Studies increasingly implicated ciliary dysfunction, transcriptional regulation, ribosomal biology, and multigenic inheritance in CHD pathogenesis. Emerging methodologies included stem cell-derived cardiac models, machine learning-based gene prioritization, and epigenetic analyses. CONCLUSIONS: Recent literature substantially expands the catalog of candidate genes that may be associated with CHD and highlights the biologic complexity underlying cardiac morphogenesis. Integration of genomic, developmental, and functional approaches will be essential to improve mechanistic understanding, refine genetic counseling, and support future precision medicine strategies for CHD.

Cardiac development↗

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our&#xa0;prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by&#xa0;the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article↗

Learning optimized features for hierarchical models of invariant object recognition.

There is an ongoing debate over the capabilities of hierarchical neural feedforward architectures for performing real-world invariant object recognition. Although a variety of hierarchical models exists, appropriate supervised and unsupervised learning methods are still an issue of intense research. We propose a feedforward model for recognition that shares components like weight sharing, pooling stages, and competitive nonlinearities with earlier approaches but focuses on new methods for learning optimal feature-detecting cells in intermediate stages of the hierarchical network. We show that principles of sparse coding, which were previously mostly applied to the initial feature detection stages, can also be employed to obtain optimized intermediate complex features. We suggest a new approach to optimize the learning of sparse features under the constraints of a weight-sharing or convolutional architecture that uses pooling operations to achieve gradual invariance in the feature hierarchy. The approach explicitly enforces symmetry constraints like translation invariance on the feature set. This leads to a dimension reduction in the search space of optimal features and allows determining more efficiently the basis representatives, which achieve a sparse decomposition of the input. We analyze the quality of the learned feature representation by investigating the recognition performance of the resulting hierarchical network on object and face databases. We show that a hierarchy with features learned on a single object data set can also be applied to face recognition without parameter changes and is competitive with other recent machine learning recognition approaches. To investigate the effect of the interplay between sparse coding and processing nonlinearities, we also consider alternative feedforward pooling nonlinearities such as presynaptic maximum selection and sum-of-squares integration. The comparison shows that a combination of strong competitive nonlinearities with sparse coding offers the best recognition performance in the difficult scenario of segmentation-free recognition in cluttered surround. We demonstrate that for both learning and recognition, a precise segmentation of the objects is not necessary.

Learning↗

Proteomic signatures and predictive modeling of cadmium-associated anxiety in middle-aged and elderly populations: an environmental exposure association study.

BACKGROUND: Emerging evidence implicates environmental contaminants such as cadmium (Cd) as modifiable risk factors for anxiety. Despite growing recognition of heavy metal toxicity in neuropsychiatric disorders, the molecular mechanisms linking environmental exposure to anxiety pathogenesis remain poorly understood. METHODS: Based on the established cohort of individuals with cognitive impairment in cadmium-contaminated areas, this cross-sectional association study enrolled 50 middle-aged and elderly hospitalized patients from these regions, adhering to the STROBE guidelines. Blood concentrations of cadmium (Cd), lead (Pb), and mercury (Hg) were analyzed in relation to anxiety severity assessed via the Hamilton Anxiety Rating Scale (HAMA). Plasma proteomic profiling was performed using data-independent acquisition (DIA) quantitative technology with an LC-MS/MS platform (timsTOF Pro, Bruker Daltonics), systematically characterizing 2,531 proteins across all samples. Machine learning techniques, specifically XGBoost and LASSO, were employed to identify biomarkers that were subsequently validated through mediation analysis and animal experiments, allowing for the screening of key protein signatures. Finally, clinical variables were integrated to construct a comprehensive model, which was then thoroughly evaluated. RESULTS: Anxious individuals exhibited significantly higher blood Cd levels than controls (&#x3b2;&#x2009;=&#x2009;0.50, 95% CI: 0.07-0.93, p&#x2009;<&#x2009;0.01), with anxiety positively correlating with depression (r&#x2009;=&#x2009;0.62, p&#x2009;=&#x2009;0.003) and inversely with ApoE3 genotype prevalence. Proteomics identified 120 differentially expressed proteins in anxious patients, enriched in oxidative phosphorylation and neurodegenerative pathways. CCDC126 emerged as a cadmium-associated biomarker, validated in rat models exposed to Cd. Combining CCDC126, blood Cd, Pb, and hypertension, a clinical prediction model achieved robust discrimination (AUC&#x2009;=&#x2009;0.80, validation cohort). CONCLUSIONS: This first integrative environmental-proteomic study highlights cadmium's synergistic role in anxiety pathophysiology and psychiatric comorbidity. The predictive model offers translatable potential for early risk stratification, while CCDC126 provides mechanistic insights for targeted interventions in populations exposed to environmental pollutants.

Cadmium↗

CCT2 defines a highly cisplatin-resistant and poor-prognosis subtype of lung adenocarcinoma.

Cisplatin-based chemotherapy is a standard treatment for lung adenocarcinoma (LUAD), yet acquired cisplatin resistance remains a marked cause of treatment failure. The molecular mechanisms driving cisplatin resistance in LUAD have not been fully elucidated. The present study integrated bulk transcriptomic data, genomic mutation profiles and single-cell RNA sequencing data to systematically investigate cisplatin resistance in LUAD. Resistance-associated genes were identified through differential expression, survival analysis and database integration. Unsupervised clustering was used to define cisplatin resistance-associated subtypes. Functional characteristics were explored using pathway enrichment, immune infiltration, tumor mutation burden and weighted gene co-expression network analysis. A machine learning framework incorporating 101 algorithms was applied to identify key genes and construct a prognostic model. Single-cell analyses and in vitro experiments were performed to validate the biological role of the core gene. Molecular docking and molecular dynamics simulations were conducted to identify potential therapeutic compounds. A total of two molecular subtypes with distinct cisplatin resistance levels and prognostic outcomes were identified. The high-resistance subtype exhibited enhanced cell cycle activity, DNA repair signaling and immune heterogeneity. Machine learning analysis revealed a five-gene signature, with chaperonin-containing TCP1 subunit 2 (CCT2) emerging as a key regulator of cisplatin resistance. Single-cell analyses showed that CCT2 was predominantly enriched in resistant epithelial cell subpopulations. Functional experiments demonstrated that CCT2 knockdown significantly inhibited cell proliferation and enhanced cisplatin sensitivity in LUAD cell lines. A number of candidate compounds targeting CCT2 exhibited stable binding in silico. The present findings identified CCT2 as a key mediator of cisplatin resistance in LUAD and provided potential therapeutic strategies to overcome chemotherapy resistance.

chaperonin-containing TCP-1 subunit 2↗

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation↗

Machine learning-guided risk stratification in elderly AML based on genomic, immunophenotypic and therapeutic profiles.

BACKGROUND: Elderly patients with acute myeloid leukemia (AML) exhibit considerable biological and clinical heterogeneity, hindering precise prognosis. Existing prognostic systems inadequately capture the complexity of elderly AML due to their reliance on data from younger cohorts and omission of key factors like immunophenotypic markers and therapeutic profiles. This study aimed to develop and internally validate a machine learning-based prognostic model specifically tailored to elderly AML patients. METHODS: A total of 156 patients were analyzed using a two-stage modeling strategy. Clinical and genomic variables were modeled first, followed by independent analysis of immunophenotypic features. Feature selection was performed using multilayer perceptron (MLP) and random forest (RF), while multivariate Cox regression was used for final model construction. Internal validation was conducted using 1000 bootstrap iterations to assess model stability and performance. RESULTS: The model demonstrated strong predictive performance, with a concordance index (C-index) of 0.702. Time-dependent area under the curve (AUC) and calibration plots confirmed accurate prediction of 1-, 3-, and 5-year overall survival. Decision curve analysis indicated favorable net benefit across a range of threshold probabilities. Key independent prognostic factors identified included TP53 mutations, high CD13 expression, and IDH2 mutations. CONCLUSION: This model provides a robust and interpretable tool for individualized risk stratification in elderly AML. By integrating genomic, immunophenotypic, and therapeutic variables, it may help optimize treatment decisions and improve outcomes for this vulnerable population. Future efforts should focus on external validation and integration of dynamic biomarkers.

Humans↗

Artificial Intelligence for Colorectal Surgeons-Part II: Research Applications, Challenges in Adoption, and Practical Resources.

BACKGROUND: This is part II of a 2-part series examining artificial intelligence in colorectal surgery. Part I established foundational concepts and clinical applications. Implementation, however, requires understanding research methodologies, available resources, and the specific challenges currently limiting widespread adoption. These topics are the focus of part II. OBJECTIVE: To examine artificial intelligence's transformation of surgical research, provide practical implementation resources, address adoption challenges, and explore future directions in colorectal surgery. METHODS: Comprehensive literature review focusing on artificial intelligence research methodology, implementation barriers, educational resources, and emerging technologies relevant to colorectal surgeons. RESULTS: Artificial intelligence streamlines clinical trial design through predictive modeling and natural language processing, reducing enrollment challenges that contribute to failed or inadequate trial accrual. Machine learning enables heterogeneity analysis within clinical trials, identifying treatment-responsive subgroups. Foundation models unlock analysis of unstructured electronic health record data at scale. Professional societies and universities offer specialized artificial intelligence education programs, with open-access data sets facilitating research participation. However, implementation faces multifaceted challenges: technical infrastructure demands, with real-time processing requiring dedicated graphics processing unit clusters; regulatory frameworks struggling with continuously evolving algorithms; undefined liability distribution for artificial intelligence-assisted decisions; algorithmic bias risking health care disparities; and the "black box" problem limiting clinical trust. Economic barriers include substantial initial costs without clear reimbursement pathways. Future directions include multimodal artificial intelligence integrating imaging, genomics, and histopathology; cognitive robotic systems with real-time decision support; digital twin technology for patient-specific surgical simulation; and global surgical artificial intelligence networks enabling distributed learning across institutions. CONCLUSIONS: Although artificial intelligence offers transformative potential for colorectal surgery research and practice, successful implementation requires addressing technical, regulatory, ethical, and economic challenges. The surgeon's evolving role demands both traditional expertise and computational fluency. Future advances in multimodal integration, autonomous systems, and global collaboration will fundamentally reshape surgical practice but will require thoughtful implementation prioritizing patient benefit and clinical value.

Humans↗

NR3C1 Modulates Wnt Signalling to Influence the Invasiveness and Immune Features of Nonfunctioning Invasive Pituitary Adenomas.

Pituitary adenomas (PAs) are common intracranial tumours, and invasiveness in nonfunctioning invasive pituitary adenomas (NIPAs) predicts poor prognosis. The molecular mechanisms driving this phenotype remain unclear. This study explored the role of nuclear receptor subfamily 3 group C member 1 (NR3C1) in NIPA invasiveness and its regulation of Wnt signalling. mRNA expression profiles of 32 PA samples were generated by RNA-seq, and proteomic data from 19 samples were obtained by mass spectrometry. Immune-related differentially expressed genes (DEGs) were retrieved from GeneCards. Weighted gene coexpression network analysis identified modules and hub genes linked to invasiveness, while machine learning methods (support vector machine, LASSO, random forest) prioritised key genes. Gene set enrichment analysis (GSEA) assessed pathways associated with candidate gene expression. NR3C1 expression and function were validated by immunohistochemistry, Western blotting and invasion assays. Integration of transcriptomic, proteomic and immune-related datasets yielded 11 overlapping genes, with NR3C1 emerging as the top candidate. NR3C1 was significantly upregulated in NIPAs and demonstrated good discriminatory power by ROC analysis. GSEA associated high NR3C1 expression with Wnt pathway activation. Functional experiments confirmed that NR3C1 overexpression enhances the invasive capacity of PA cells. NR3C1 promotes the invasive phenotype of NIPAs by activating Wnt signalling. These findings suggest NR3C1 as a potential biomarker and therapeutic target for invasive pituitary adenomas.

Humans↗

Mitochondria related gene signature serves as prognosis prediction and risk stratification of cholangiocarcinoma.

BACKGROUND: Cholangiocarcinoma (CHOL) is a highly aggressive biliary malignancy with poor clinical outcomes and limited effective prognostic biomarkers. Mitochondrial dysfunction participates in multiple oncological processes of CHOL, yet the prognostic roles of mitochondria&#x2011;related genes (MRGs) remain poorly understood. This study aimed to characterize MRGs expression in CHOL and develop a molecular prognostic model for predicting patient survival and guiding clinical management. METHODS: RNA sequencing (RNA-seq) and clinical data of CHOL were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) (GSE89748) databases. Differentially expressed MRGs were identified, and 10 machine learning algorithms were used to construct prognostic models. The optimal model (highest average C-index) was selected to establish a mitochondria-related risk score (MRRS), which was validated internally and externally. A nomogram integrating clinical factors and MRRS was developed, and biological mechanisms were explored via functional and immune analyses. RESULTS: A 3-MRG (MAP3K1, MRPL18, PYGB) prognostic signature was constructed, stratifying patients into high- and low-risk groups with significantly different overall survival. The model showed high predictive accuracy, with an area under the curve (AUC) up to 0.845, and MRRS was an independent prognostic factor. The signature was associated with mitochondrial pathways, and the high-risk group had distinct immune infiltration and mutation profiles. CONCLUSIONS: A validated MRG prognostic model effectively stratifies CHOL patients and has potential clinical value for prognosis prediction. Further validation in larger cohorts is needed to confirm its applicability.

Cholangiocarcinoma (CHOL)↗

Rapid glycomic analysis of serum EVs reveals altered N-glycosylation patterns in ASD.

Objective laboratory diagnostics for autism spectrum disorder (ASD) are lacking, necessitating rapid clinical screening tools. Because serum extracellular vesicle (EV) N-glycosylation captures critical neurodevelopmental signatures, we developed a fast, biologically interpretable diagnostic strategy. EVs from ASD patients with language impairment and neurotypical controls were isolated using a rapid extra-polyethylene glycol precipitation/filtration (EPF) workflow, benchmarked against ultracentrifugation. Following MALDI-TOF/MS profiling, machine learning was re-evaluated using repeated nested cross-validation to reduce optimistic bias and potential information leakage. Among five classifiers, Random Forest (RF) showed the best overall balance across discrimination, calibration, and classification metrics. RF-based SHAP analysis provided transparent interpretation, highlighting key discriminative glycans, including H4N3S1F1, H5N5S1F1, and H3N5F1. To elucidate molecular mechanisms, we integrated public EV transcriptomic data. This revealed significant dysregulation of N-glycosylation machinery genes (e.g., MAN1A1, NEU1, OSTC, RPN2), whose expression directionally aligned with observed glycan shifts in synaptic pathways. Collectively, this rapid serum EV N-glycomic workflow, combined with leakage-controlled RF-based interpretation, provides a promising foundation for non-invasive ASD biomarker discovery and future multicenter validation.

Humans↗

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution↗

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11&#xa0;476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans↗