PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine-learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization

Rhizosphere Dialogue: Microorganisms Mediated by Root Exudates Alleviate Drought Stress in Grasses.

Drought stress threatens the ecological functions and economic value of grasses, posing a major challenge to their sustainable production. Plants co-evolve with rhizosphere microbial communities, sometimes described as the plant's second genome, that can contribute to drought adaptation. Drought alters root architecture, hormonal and redox regulation and belowground carbon allocation, thereby modifying the quantity and composition of root exudation and reshaping the rhizosphere environment. This review uses the rhizosphere dialogue as an integrative framework to link these plant responses with microbial recruitment and subsequent feedback to the host. We summarise three linked stages of this dialogue: drought-induced changes in root exudation; microbial recruitment and colonisation through chemotaxis, attachment, biofilm formation, and root colonisation; and microbiome-mediated feedback that improves plant water relations, hormonal and redox homoeostasis, nutrient acquisition, and root function. We highlight microbial extracellular polymeric substances, 1-aminocyclopropane-1-carboxylate deaminase, and microbial volatile organic compounds as key mediators of drought alleviation. We then discuss how this framework may inform rational synthetic microbial community (SynCom) design, microbiome-informed breeding, artificial intelligence and machine-learning assisted strain prioritisation, rhizosphere legacy effects, and real-time monitoring. Future work should distinguish active exudate-mediated recruitment from drought-driven environmental filtering and integrate multi-omics, plant genetics, functional validation, and multi-location field trials to determine whether rhizosphere dialogue can become a predictive framework for climate-resilient grass production.

drought stress

The prevalence and clinical significance of clonal monocytosis.

The terms clonal monocytosis of undetermined significance (CMUS) and clonal cytopenia and monocytosis of undetermined significance (CCMUS) were introduced by the International Consensus Classification of Myeloid Neoplasms to describe cases of clonal hematopoiesis (CH) and concurrent monocytosis that did not meet the diagnostic criteria of chronic myelomonocytic leukemia. To date, their practical relevance as clinicopathological entities at a population level has not been assessed. Here, we assess the prevalence, significance, and natural history of CMUS and CCMUS among 431 531 UK Biobank participants through analysis of clinical, genomic, and health outcome data. We find that CMUS with an absolute monocytosis and CCMUS are high-risk entities strongly associated with incident myeloid neoplasia (MN), cardiovascular disease, and renal disease. Noting the overall higher monocyte counts in men and the low rate of progression of DNMT3A-CMUS, we reveal that amending the definition of CMUS/CCMUS to incorporate sex-specific monocyte thresholds and the exclusion of isolated DNMT3A mutations from the definition significantly strengthens the association with incident MN. Finally, given their association with poor outcomes, we develop MoSAIC, a machine-learning classifier, to infer the presence of SRSF2 mutations (associated with high MN risk) among individuals with monocytosis, based on complete blood count indices alone. We corroborate our findings in an independent cohort of 625 328 Danish primary care patients. Our findings underscore the clinical relevance of CMUS and CCMUS as distinct high-risk states within the spectrum of CH and establish an evidence base to refine their diagnostic definition.

Humans

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides

Transcriptome-based high-frequency recurrence index predicts frequent recurrence in non-muscle-invasive bladder cancer after Bacillus Calmette-Gu&#xe9;rin therapy.

BACKGROUND: High-frequency recurrence (HfR,&#x2009;&#x2265;&#x2009;2 recurrences) in non-muscle-invasive bladder cancer (NMIBC) poses a significant clinical burden. Current risk models, such as the European Organization for Research and Treatment of Cancer (EORTC), the European Association of Urology (EAU), and the UROMOL classification, offer limited predictive accuracy for identifying patients at risk for frequent recurrence despite appropriate treatment. METHODS: A 75-gene high-frequency recurrence index (HfRI) was constructed by selecting recurrence-associated genes using differential expression and Cox regression analyses. The HfRI was computed as a weighted sum of normalized gene expression values. The model was trained on a discovery cohort and validated in multiple cohorts (n&#x2009;=&#x2009;1379) using machine-learning approaches. Clinical relevance was assessed using recurrence-free survival (RFS) and Cox models, and predictive performance was compared with that of the EORTC, EAU, and UROMOL classifications using the area under the curve (AUC) and the concordance index (c-index). RESULTS: The HfRI robustly stratified patients into high-risk and low-risk groups across six independent NMIBC cohorts. Patients classified as HfRI-high had a significantly greater likelihood of experiencing&#x2009;&#x2265;&#x2009;2 recurrences (&#x3c7;2, p&#x2009;=&#x2009;0.001) and showed markedly reduced RFS (log-rank test, p&#x2009;<&#x2009;0.001). The adverse prognostic effect of the HfRI persisted even among patients treated with BCG therapy (log-rank test, p&#x2009;=&#x2009;0.02). Multivariate analysis revealed that the HfRI was an independent predictor of HfR (HR&#x2009;=&#x2009;2.82, 95% CI&#x2009;=&#x2009;1.89-4.20, p&#x2009;<&#x2009;0.001). Compared with established clinical risk classifiers, the HfRI demonstrated superior predictive performance (AUC&#x2009;=&#x2009;0.736, c-index&#x2009;=&#x2009;0.673) in terms of the EORTC (AUC&#x2009;=&#x2009;0.594), EAU (AUC&#x2009;=&#x2009;0.557) risk groups, and UROMOL2021 (AUC&#x2009;=&#x2009;0.596) classification. Pathway analysis revealed that HfRI-high tumors were characterized by upregulation of cell cycle progression and DNA replication pathways, accompanied by suppression of immune signaling pathways. These biological features provide a mechanistic explanation for the reduced responsiveness to intravesical BCG therapy, underscoring the role of HfRI not only as a predictor of recurrence risk but also as a biomarker capable of identifying patients unlikely to benefit from standard BCG treatment. CONCLUSIONS: HfRI represents a robust, transcriptome-based tool for predicting frequent recurrence in NMIBC patients. The HfRI supports earlier identification of patients at risk of high-frequency recurrence, thereby supporting personalized treatment strategies.

Humans

Biomarkers related to m6A and succinic acid metabolism in papillary thyroid carcinoma.

BACKGROUND: Studies have shown that m6A modification is related to the occurrence and development of papillary thyroid carcinoma (PTC). The disorder of succinic acid metabolism is associated with the occurrence and development of various tumors. However, there are few studies based on m6A and succinate metabolism-related genes (SMRGs) in PTC. METHODS: The TCGA-Thyroid carcinoma (THCA), GSE33630, 1159 SMRGs, and 23 m6A regulatory factors were collected from the online databases. Subsequently, the differentially expressed genes (DEGs) were selected between PTC (Tumor) and Normal samples. The overlapping genes among the DEGs, m6A, and SMRGs were applied to screen the biomarkers. Using the 3 machine-learning algorithms, the biomarkers were determined based on the overlapping genes. Next, the biomarkers were evaluated by the ROC curve and expression analysis in TCGA-THCA and GSE33630. Then, the overall survival (OS) differences were compared between the high-and low-expression biomarkers. Finally, immune infiltration analysis, molecular regulatory network, and drug prediction were performed based on the biomarkers. RESULTS: In TCGA-THCA, there were 2800 DEGs between and Normal samples, and then 7 overlapping genes were obtained. Importantly, ADK, TNFRSF10B, CYP7B1, FGFR2, and CPQ were determined as biomarkers with excellent diagnostic efficiency (AUC&#x2009;>&#x2009;0.7). In PTC samples, ADK and TNFRSF10B were high-expressed while CYP7B1, FGFR2, and CPQ were low-expressed. Especially, the high-expression groups of ADK had a better prognosis, while the high-expression groups of CYP7B1, FGFR2, and CPQ had a worse prognosis. Afterward, immune infiltration analysis found that 16 immune cells had infiltration differences between the Tumor and Normal samples. Finally, transcription factor SP1 could regulate CYP7B1 and TNFRSF10B. Moreover, Navitoclax was a potential drug for PTC patients. CONCLUSION: Overall, we described 5 biomarkers associated with adverse prognosis of PTC, including ADK, TNFRSF10B, CYP7B1, FGFR2, and CPQ. All these biomarkers were involved in succinate metabolism and m6A modification of RNA. This set of biomarkers should be explored further for their diagnostic value in PTC. Investigations into the mechanistic role of alteration of succinate metabolism and m6A modification of RNA pathways in the pathophysiology of PTC are warranted.

Humans

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive

Establishment of a prognostic model based on ER stress-related cell death genes and proposing a novel combination therapy in acute myeloid leukemia.

BACKGROUND: Acute myeloid leukemia (AML) is a highly heterogeneous malignancy, presenting significant challenges in accurately predicting patient prognosis. Dysregulation of endoplasmic reticulum (ER) stress and resistance to programmed cell death (PCD) are hallmarks of AML cells. However, the prognostic significance of the interplay between ER stress and cell death pathways in AML remains largely unexplored. METHODS: We analyzed RNA sequencing and clinical data from 887 AML patients across 4 cohorts to develop an ER stress-related cell death index (ERCDI) using 10 machine-learning algorithms with 117 unique combinations. Survival and time-dependent Receiver Operating Characteristic Curve (ROC) analyses were performed to assess the model's efficacy. Clinical characteristics, the tumor immune microenvironment, and drug sensitivity differences between the high- and low-risk groups were also analyzed. The CMap database was used to identify potential therapeutic drugs. In vitro and in vivo experiments, including CCK-8, colony formation, flow cytometry, Transwell assays, and xenograft mouse models, were conducted to evaluate the effects of the target genes and candidate drugs. RESULTS: The ERCDI demonstrated strong prognostic and predictive performance for prognosis in AML patients. Furthermore, the ERCDI effectively predicted immunotherapy and chemotherapy outcomes and was associated with the immune features of the different risk groups. DNA damage-inducible transcript 4 protein (DDIT4), a key gene associated with ERCDI, is related to poor prognosis in AML patients with high expression. Additionally, the knockdown of DDIT4 significantly inhibited AML cell proliferation, induced cell apoptosis, and promoted cell cycle arrest. Chaetocin was subsequently identified as a candidate compound for AML treatment. Subsequent experiments suggested that combining chaetocin and venetoclax is a potentially promising therapeutic strategy for AML. CONCLUSION: The ERCDI provides personalized risk assessment and treatment recommendations for individual AML patients. The combined use of chaetocin and venetoclax can potentially be repurposed for AML therapy.

Humans

Uncovering the genetic architecture of ME/CFS: a precision approach reveals impact of rare monogenic variation.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a disabling and heterogeneous disorder lacking validated biomarkers or targeted therapies. Clinical variability and elusive pathophysiology hinder progress toward effective diagnostics and treatment. Core symptoms include persistent fatigue, post-exertional malaise, unrefreshing sleep, cognitive dysfunction, and pain. We tested whether an individualized, &#x201c;n-of-1&#x201d; genomic and transcriptomic framework combined with comprehensive, participant-informed phenotyping could reveal molecular signatures unique to each patient. METHODS: Clinical-grade whole-genome sequencing was conducted in 31 affected individuals from 25 families, with RNA-seq performed on a subset (16 affected, 7 unaffected) using blood samples. Machine-learning assisted variant triage, transcript-aware damage prediction, and expert review identified pathogenic or likely pathogenic variants in 8 of 25 probands (32%) and 12 of 31 affected individuals (39%). RESULTS: Findings revealed marked genetic heterogeneity, including large-effect rare and more common variants. Implicated pathways included ATP generation, oxidative phosphorylation, fatty acid oxidation; regulation of glycolysis, amino acid and lipid turnover; ion and solute homeostasis; synaptic signaling, excitability, oxygen transport, and muscle integrity, resilience, and post-exertional recovery; previously implicated processes. Plausible modifiers influencing disease onset, severity, and relapsing&#x2013;remitting patterns and possibly explaining intrafamilial variability and inconsistent findings across studies, were also identified. Despite gene-level diversity, downstream effects converged on impaired energy production, reduced stress resilience, and vulnerability to post-exertional metabolic failure; disruptions consistent with core ME/CFS symptoms of exertional intolerance, cognitive fog, and fatigue. CONCLUSIONS: Our findings support the hypothesis that at least a subset of ME/CFS cases represent distinct molecular disorders that converge on shared physiological pathways. Validation in larger, more diverse cohorts will be essential to test this hypothesis and establish generalizability, but increase size alone is unlikely to resolve causation in a disorder defined by rarity, heterogeneity, and molecular complexity. We suggest that progress will require experimental designs that integrate individual-level genomic data with deep, participant-informed deep phenotyping, capturing the combined effects of rare and common variants and environmental modifiers on disease expression and progression. We believe that an individualized precision medicine framework will uncover molecular drivers and modifiers of ME/CFS previously obscured by heterogeneity, enabling biologically informed stratification, improved trial design, biomarker discovery, and targeted interventions in this historically neglected condition.

Humans

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans

Proteomic and machine learning analysis predicts treatment response signatures in Myasthenia Gravis.

BACKGROUND: Myasthenia gravis (MG) is a prototypical antibody-mediated autoimmune disease with variable treatment responses with a need for biomarkers to guide therapeutic decision making. Proteomic profiling, coupled with machine learning, offers a hypothesis-free approach to identify multi-protein signatures associated with treatment response. METHODS: We analyzed sera collected at entry (baseline) from participants in a phase 3 trial randomized trial comparing thymectomy plus prednisone versus prednisone alone, along with matched controls using liquid chromatography-mass spectrometry. We derived disease-specific proteomic signatures and evaluated associations between baseline proteins and 6-month clinical outcomes using multiple machine-learning approaches with internal validation. RESULTS: Baseline serum proteomes distinguished MG from controls, with pathway enrichment implicating complement activation, immunoglobulin production, and T-cell receptor signaling. Distinct protein panels predicted 6-month clinical improvement within each treatment arm. In the thymectomy-plus-prednisone group, models captured non-linear relationships of predictive proteins in contrast with the predominant additive patterns observed in the prednisone-alone group. Predictive proteins were enriched for T-cell signaling and leukocyte trafficking functions, providing insight into treatment-specific biology. CONCLUSIONS: Baseline serum proteomics captures core disease characteristics of MG and predicts short-term clinical response in a treatment-specific manner. While our results require validation in independent cohorts, these findings could enable biomarker-guided selection of thymectomy, refine risk stratification, and furnish mechanistic readouts for future MG trials and clinical care. We aim to conduct future studies using -omic approaches to validate these baseline predictive biomarkers and pathways of treatment response in patients with MG.

Adult

Machine learning and multi-omics clustering to map cellular rewiring and immune evasion in ccRCC.

Immune checkpoint blockade (ICB) efficacy in clear cell renal cell carcinoma (ccRCC) is limited by tumor microenvironment (TME) heterogeneity. Because traditional bulk-derived models lack spatial resolution, we developed an integrated framework connecting macroscopic survival risks to microscopic TME structures. We applied ten algorithms to establish multi-omics subtypes and evaluated 101 machine-learning combinations across three independent cohorts to generate a Consensus Machine Learning-driven Signature (CMLS). The signature's spatial and cellular origins were decoded using spatial transcriptomics (ST) and a 140,000-cell scRNA-seq atlas. Expression of key genes was experimentally validated via RT-qPCR in 17 paired ccRCC clinical tissues. We identified two molecular subtypes with distinct clinical and epigenetic profiles. SuperPC optimization yielded a 24-gene CMLS serving as an independent prognostic factor. scRNA-seq and ST deconvolution revealed these signals predominantly originate from cancer-associated fibroblasts (CAFs) and malignant epithelial cells, which collaborate to drive spatial immune exclusion. RT-qPCR confirmed significant overexpression of five core CMLS genes in ccRCC versus adjacent normal tissues. Low CMLS scores correlated with enhanced ICB responsiveness, whereas high-CMLS tumors demonstrated specific vulnerability to dasatinib and dabrafenib. The CMLS translates spatial immune-exclusion dynamics into a quantifiable metric, outperforming tumor mutational burden in predicting ICB benefits, providing a robust tool for patient stratification in ccRCC.

Humans

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans

Genetic targets related to aging for the treatment of coronary artery disease.

BACKGROUND: Coronary Artery Disease (CAD) is the most common cardiovascular disease worldwide, threatening human health, quality of life and longevity. Aging is a dominant risk factor for CAD. This study aims to investigate the potential mechanisms of aging-related genes and CAD, and to make molecular drug predictions that will contribute to the diagnosis and treatment. METHODS: We downloaded the gene expression profile of circulating leukocytes in CAD patients (GSE12288) from Gene Expression Omnibus database, obtained differentially expressed aging genes through "limma" package and GenaCards database, and tested their biological functions. Further screening of aging related characteristic genes (ARCGs) using least absolute shrinkage and selection operator and random forest, generating nomogram charts and ROC curves for evaluating diagnostic efficacy. Immune cells were estimated by ssGSEA, and then combine ARCGs with immune cells and clinical indicators based on Pearson correlation analysis. Unsupervised cluster analysis was used to construct molecular clusters based on ARCGs and to assess functional characteristics between clusters. The DSigDB database was employed to explore the potential targeted drugs of ARCGs, and the molecular docking was carried out through Autodock Vina. Finally, single-cell data (GSE159677) of arterial intima was used to further explore the expression of aging signature genes in different cell subpopulations. RESULTS: We identified 8 ARCGs associated with CAD, in which HIF1A and FGFR3 were up while NOX4, TCF7L2, HK3, CDK18, TFAP4, and ITPK1 were down in CAD patients. Based on this, CAD patients can be divided into two molecular clusters, among which cluster A mainly involves functional pathways such as ECM receptor interaction and focal adhesion; cluster B mainly involves functional pathways such as amimo sugar and nucleotide sugar metabolism and pyrimidine metabolism. In addition, the molecular docking results showed that retinoic acid and resveratrol had good binding affinity with targets genes. Further single-cell analysis results showed that NOX4, TCF7L2, ITPK1, and HIF1A were specifically expressed in different types of cells in atherosclerotic tissues. CONCLUSION: Our study identified several ARCGs that may be involved in the pathogenesis and progression of CAD. Further, retinoic acid and resveratrol were potential candidate molecule drugs for inhibiting these targets.

Humans