PubMed HealthSearch

Biomedical subjects

Yong Chen

Publications and source records attributed to Yong Chen.

6 recordsLinked to original sources

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model

Frequent mutations in the BIRC3 gene promote metastatic potential of nasopharyngeal carcinoma cells through the TRAF2-NF-κB pathway.

Nasopharyngeal carcinoma (NPC) is a head and neck cancer characterized by highly locoregionally invasive behavior attributable to the latent infection with Epstein-Barr virus (EBV) and genomic instability. It is well established that EBV-encoded oncogenic molecules actively contribute to the malignant behavior of NPC cells. However, the mechanism by which aberrant genomic alterations enable NPC cells to become aggressive remains largely unknown. In the present study, whole-exome sequencing (WES) revealed that the gene encoding the baculoviral IAP repeat-containing 3 (BIRC3) protein was frequently mutated in circulating tumor cells (CTCs) but not in paired primary tumor cells from patients with metastatic NPC. A minigene assay indicated that the c.637 A > G mutation disrupted normal mRNA splicing, resulting in the partial deletion of Exons 2 and 3 and altered stability of BIRC3 mRNA. In vitro experiments demonstrated that ectopic expression of the BIRC3c.637A>G mutant enhanced NPC cell invasive properties, including proliferation, resistance to apoptosis, migration, and invasion. Furthermore, overexpression of wild-type BIRC3 promoted invasive characteristics in NPC cells through the TRAF2-NF-κB signaling axis. In summary, BIRC3 acts as a regulator of the malignant features of NPC cells. Frequent BIRC3 mutations in CTCs, such as the c.637 A > G mutation, further enhance the metastatic potential of disseminated NPC cells by inducing aberrant alternative splicing. These findings suggest the therapeutic feasibility of targeting the BIRC3/TRAF2/NF-κB axis in the treatment of NPC.

Humans

An Integrative Morphological and Genomic Analysis With a Refined Fluorescence In Situ Hybridization (FISH) Threshold and Novel Kinase Fusions in a Large Asian Cohort of Spitzoid Neoplasms.

Differentiating atypical Spitz tumors (ASTs) from true Spitz melanomas (SMs) and conventional melanomas with spitzoid features (MSFs) remains a formidable diagnostic challenge. Because current molecular epidemiological data are overwhelmingly derived from Caucasian cohorts, the genomic landscape of Asian populations remains largely unexplored. To elucidate the molecular progression landscape and refine the diagnostic criteria, we performed a comprehensive multimodal analysis-integrating histomorphology, immunohistochemistry, multiprobe fluorescence in situ hybridization (FISH), and targeted RNA/DNA-based next-generation sequencing (NGS)-on a cohort of 140 spitzoid neoplasms. This cohort, comprising 126 ASTs, 8 SMs, and 6 MSFs, represents the largest Asian cohort to date. Malignant phenotype strongly correlated with lesional asymmetry, deep atypical mitoses, a sheet-like growth pattern, diffuse preferentially expressed antigen of melanoma positivity, and significant loss of p16 expression (64.3% in SM/MSF vs 9.5% in ASTs; P < .0001). Building upon the established melanoma FISH criteria, we optimized a prognostic threshold of &#x2265;2 FISH abnormalities specifically tailored for spitzoid neoplasms. We demonstrated that isolated single chromosomal aberrations (particularly MYB loss) are relatively stable events that are frequent in indolent ASTs, whereas our refined &#x2265;2 threshold yielded 100% sensitivity and 92.5% specificity for predicting regional lymph node metastasis/local recurrence. Molecularly, NGS identified mutually exclusive initiating driver alterations (comprising kinase fusions and HRAS mutations) in 89.9% of true Spitz neoplasms, a remarkably high prevalence suggesting a distinct genetic background in Asian populations. We also characterized 5 entirely novel kinase fusions (ZNF24::ROS1, PCBP1::ROS1, NUMA1::RET, CBWD1::ALK, and TPR::NTRK1). Furthermore, NGS definitively segregated true Spitz neoplasms from morphological mimics (MSF), which lacked fusions and were driven by canonical genomic alterations of the conventional melanoma pathway. Integrating these genomic landscapes validated a stepwise progression model. Although isolated kinase fusions drove indolent ASTs, malignant SM invariably harbored concurrent pathogenic secondary alterations, demonstrating a profound reliance on CDKN2A/B, TP53, and CDK4 aberrations. Ultimately, we propose an integrated diagnostic algorithm combining morphological evaluation, the refined FISH threshold, and comprehensive NGS profiling, providing a precise, evidence-based framework for pathway classification and clinical management of spitzoid neoplasms.

fluorescence in situ hybridization

Comparative genomic characterization and antimicrobial resistance of bacteremia-causing Enterococcus faecium and Enterococcus faecalis in a Chinese hospital.

Enterococci are common commensals of the human gut and important opportunistic pathogens, with Enterococcus faecium and Enterococcus faecalis being the most clinically prevalent species. A significant epidemiological shift has emerged with an increasing clinical burden of E. faecium. To compare genomic evolution of E. faecium and E. faecalis, we performed whole-genome sequencing on 93 E. faecium and 32 E. faecalis isolates causing bloodstream infections at a single hospital (2022-2024). Analysis of patient demographics revealed that E. faecium infections originated from fewer sources than E. faecalis, with a higher proportion deriving from intra-abdominal infections. Multilocus sequence typing identified ST78 and ST789 as the predominant sequence types for E. faecium, whereas ST16 and ST179 were most common for E. faecalis. E. faecium carried more antimicrobial resistance genes and putative virulence marker (PVM)-type virulence genes than E. faecalis, with vancomycin resistance predominantly mediated by vanHAX (33/93, 35.5%) and a single E. faecalis isolate also carrying vanHAX (1/32, 3.1%); the structurally incomplete vanHMX gene cluster was detected in 11 E. faecium isolates. Pan-genome analysis indicated a larger core genome in E. faecalis compared to E. faecium, consistent with greater plasmid replicon diversity in the latter. Intra-host comparisons showed that two E. faecalis pairs from the same patient were clonally related, with one isolate acquiring a vanHAX plasmid conferring vancomycin resistance. In contrast, E. faecium isolates exhibited marked genomic diversity even among clonally related pairs. These findings suggest that E. faecium possesses greater genomic plasticity and adaptive potential to the clinical environment.IMPORTANCEThis study provides a detailed comparison of clinical and genomic features between Enterococcus faecium and Enterococcus faecalis from the same hospital setting. We show that E. faecium isolates, mainly ST78/ST789, carry more antimicrobial resistance genes and a higher number of putative virulence marker (PVM) genes than E. faecalis, reflecting their hospital-adapted nature. E. faecium also exhibits a smaller core genome and greater diversity of plasmid replicon types, indicating higher genomic plasticity and capacity for horizontal gene transfer. By contrast, E. faecalis retains a larger core genome and a set of classical virulence factors, and its within-host isolates are clonally related. These distinct genomic profiles help to understand how the two species adapt to clinical environments and may inform more targeted infection control strategies and resistance surveillance.

Enterococcus faecium

Protein isolation markedly enhances in vitro digestibility, nutritional quality, and bioactivity of fungal mycelial proteins.

Fungal mycelial proteins are promising sustainable protein sources, yet their nutritional utilization is often limited by structural constraints. This study systematically evaluated the effects of protein isolation on the proteomic composition, gastrointestinal digestion behavior, amino acid utilization, and bioactivity of Pleurotus citrinopileatus mycelial proteins. Quantitative proteomics identified 3591 proteins, of which 3374 were shared between mycelial flour (PCMF) and protein isolate (PCMPI), indicating that PCMPI primarily represents the soluble proteome fraction. In vitro digestion revealed that PCMPI exhibited significantly higher digestibility (93.98%) than PCMF (42.98%) (p&#xa0;<&#xa0;0.05), reaching levels comparable to whey protein isolate. Enhanced enzymatic accessibility in PCMPI promoted rapid peptide generation during the gastric phase and efficient amino acid release during the intestinal phase, resulting in higher peptide (634.76&#xa0;mg/g) and free amino acid levels (341.69&#xa0;mg/g) at the digestion endpoint. Consequently, PCMPI achieved a balanced amino acid profile with a PDCAAS of 1.0. Moreover, its digestion products exhibited stronger antioxidant activity (IC&#x2085;&#x2080;&#xa0;=&#xa0;8.36&#xa0;mg/mL) and ACE inhibitory activity (IC&#x2085;&#x2080;&#xa0;=&#xa0;15.65&#xa0;mg/mL) compared with PCMF. Mechanistically, protein isolation disrupted the cell wall matrix, shifting digestion from a structure-limited to an accessibility-driven regime. Collectively, these findings demonstrate that protein isolation markedly enhances the digestibility, nutritional quality, and functional potential of mycelial proteins, supporting their application as high-value sustainable protein ingredients.

Digestion

dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference charts.

MOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS.

Algorithms