Identification and genome characterization of circovirus in masked palm civets (Paguma larvata) from Guangdong Province, China.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Jing Chen.
Explore the source record for details and available documents.
Animal hearts display diverse anatomical structures during adaptive evolution. Here, we present a multiomics atlas of adult hearts from 27 species across chordates, arthropods, and mollusks. Joint analysis indicates that Bilateria hearts share a core gene repertoire, taking a stepwise "add-on" approach as a universal evolutionary strategy. The "proto-heart" is populated by key cell types, including cardiomyocytes, fibroblasts, endothelial cells, and neural cells, which maintained core signatures while evolving with shifts in living environments and corresponding adaptations in the cardiovascular system. Additionally, we reveal an evolutionarily conserved cardiomyocyte state dynamic potentially linked to cardiac development and stress responses. Finally, we identify a common molecular program underpinning chamber evolution from a ventricular foundation. This work establishes a resource for understanding the intrinsic mechanisms of heart evolution.
BACKGROUND: In ∼10% of asthma patients, symptoms remain uncontrolled despite maximal treatment, representing an unmet clinical need. The causal variants, genes and pathways underlying genetic risk factors have not been fully elucidated, and it is unclear whether there are unique genetic risk factors for this asthma subtype. METHODS: We used electronic healthcare records linked to UK Biobank to identify asthma patients with high treatment burden and/or worse outcomes. We performed a genome-wide association study (GWAS) with this case population and healthy controls. We sought replication for associated (p≤5×10-6) signals in four independent studies (12 152 cases and 32 316 controls). Replicated signals were fine-mapped and linked to genes and pathways. RESULTS: In total, 7681 participants met our case definition and showed enrichment for adult-onset asthma, female gender and higher body mass index compared to asthma individuals not meeting case criteria. GWAS with 7681 cases and 38 405 controls revealed 21 reproducible association signals that had previously been associated with asthma, but had a larger effect size in our study. Variant-to-gene mapping highlighted 85 candidate genes, five of which were considered high confidence (BACH2, D2HGDH, IL1RL1, RPS26, SMAD3). CONCLUSION: We present the first use of electronic healthcare records in UK Biobank to identify a subtype of asthma enriched for patients with high treatment burden and/or worse outcomes. Our findings support the role of known asthma genes, highlighting genetic risk variants with stronger effect in these groups of patients. The prioritised genes provide potential therapeutic opportunities for this difficult-to-treat patient population.
BACKGROUND: Despite multiple clinical trials, disease-modifying treatments for COPD are currently limited. Since many drugs target proteins, identifying causality between proteins and lung function informs understanding of COPD pathophysiology and may suggest novel targets. We used Mendelian randomisation (MR) to prioritise proteins as potentially causal for imparied lung function. For prioritised proteins, we explored their potential suitability as drug targets by predicting their effects on a range of clinical outcomes. METHODS: We used genome-wide association study (GWAS) data on 2923 proteins (n=48 195, UK Biobank) to identify single genetic variants (protein quantitative trait loci (cis-pQTLs)) associated with protein levels (p≤5×10-9, variant ≤100 kb of a transcription start site). We performed cis-pQTL-MR analyses of four spirometric traits (n=149 166, 36 independent cohorts). Sensitivity analyses included colocalisation and reverse direction MR. We report associations between cis-pQTLs for prioritised proteins and multiple clinical respiratory outcomes, and use phenome-wide analysis to explore potential adverse effects or drug repurposing opportunities. FINDINGS: 1841 proteins had a suitable cis-pQTL. We implicated 16 proteins as potentially causal for lung function (p<1.71×10-5): seven proteins have not been implicated by previous lung function GWAS or MR (CCND2, DTD1, PILRA, PTPRK, TDRKH, GRHPR, NUDT5), and we provide corroborative evidence for 10 proteins. We add to the literature identifying surfactant protein D (SFTPD) as a candidate, yet predict that integrin subunit alpha V (ITGAV) inhibition could impair some lung function measures, mimicking adverse results from a recent trial. INTERPRETATION: Our approach identifies proteins (some novel) that are potentially therapeutic targets for respiratory disease, and which warrant follow-up for utility and safety.
BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.
BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.
BACKGROUNDWe constructed multi-trait polygenic risk scores (PRSs) predicting chronic obstructive pulmonary disease (COPD) and exacerbations, validated their performance in diverse cohorts, and identified PRS-related proteins for potential therapeutic targeting.METHODSPRSmix+, a multi-trait PRS framework, is used to train a composite PRS (PRSmulti) in COPDGene non-Hispanic White participants (n = 6,647). Associations of PRSmulti with COPD status (GOLD 2-4 vs. GOLD 0 or ICD) and exacerbation frequency were tested in COPDGene African American (n = 2,466), ECLIPSE (n = 1,858), Mass General Brigham Biobank (n = 15,152), and All of Us (n = 118,566). Protein prediction models were applied to GWAS summary statistics from traits contributing to PRSmulti and were validated with proteomic data in COPDGene (n = 5,173) and UK Biobank (n = 5,012).RESULTSPRSmix+ selected 7 traits for PRSmulti. In multivariable models, PRSmulti was associated with COPD status (meta-analysis random effects [RE] OR 1.58 [95% CI: 1.28-1.94]) and exacerbation frequency (meta-analysis RE β 0.21 [95% CI: 0.11-0.31]), with higher effect sizes observed in smoking-enriched cohorts. PRSmulti outperformed traditional single-trait PRS in all tested cohorts. Using protein prediction models, we identified 73 proteins associated with the PRSs that were also validated with measured protein levels in COPDGene and UK Biobank. Of these proteins, 25 were linked to approved or investigational drugs. Notable targets include RAGE/sRAGE, IL1RL1, and SCARF2, all implicated in COPD pathogenesis and exacerbations.CONCLUSIONSMulti-trait PRS improves prediction of COPD and exacerbation risk. Integration with proteomic data identifies druggable protein targets, offering a promising avenue for precision medicine in COPD management.TRIAL REGISTRATIONCOPDGene: ClinicalTrials.gov NCT00608764; ECLIPSE: ClinicalTrials.gov NCT00292552.
Historical and archaeological records indicate that the Maritime and Land Silk Roads played a pivotal role in facilitating Trans-Eurasian migrations and cultural exchanges. However, the extent to which population movements or the spread of ideas shape Chinese Hui populations remains debated. We present the largest genomic resource to date, including 2,280 Hui individuals sequenced or genotyped from 30 diverse regions, to examine the genetic origins, population structure, and biological adaptations of this underrepresented group in global human genome research. We identified a detailed population structure characterized by five distinct genetic lineages of the Hui, influenced by geography and varying gene flow. The admixture history and demographic events suggest that the northwestern and northern Hui lineages emerged from demic diffusion during the Tang and Yuan Dynasties via the Land Silk Road. In contrast, the southern and island Hui lineages reflect cultural diffusion along the Maritime Silk Road, while the mixed southern-northern lineage likely developed through a combination of demic and cultural diffusion. Our findings support a hybrid model for Hui formation, indicating that both demographic processes and sociocultural transmissions contributed to their population history. We identified east-west highly differentiated variants and pre- and post-admixture adaptations in Hui genomes, demonstrating that admixture-driven adaptive or neutral variants impacted susceptibility to cardiovascular diseases and immune- and diet-related traits. These adaptive signatures include post-admixture signals of SLC24A5 and ECHDC1 in the Hui, as well as pre-admixture signals of the HLA region, BCL2A1, and KCNH8 in the East Asian source. Overall, our study suggests that Han-related genetic components helped the Hui population rapidly adapt to new local environments. Additionally, the frequency spectrum of clinically essential variants differed significantly between Hui and Han individuals, emphasizing the importance of including underrepresented populations in genomic research to promote health equity.
Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.
To investigate how aging hallmarks exert roles in the age-related disease of coronary artery disease (CAD). R software and the GEO2R online tool identified differentially expressed genes (DEGs) and differentially expressed microRNAs (DEMis) in CAD microarray datasets from the Gene Expression Omnibus. Genes common to target genes of DEMis, DEGs, and an aging gene list from Human Aging Genomic Resources were then identified and analyzed for protein-protein interactions and functional and pathway enrichment. An miR-mRNA network was constructed using Cytoscape. Receiver operating characteristic curve analysis assessed the diagnostic utility of DEMis in CAD. The expression of two DEMis from a CAD cohort was employed to validate the findings. An aging hallmark gene set, comprising 18 genes, was delineated, with the hub gene TP53 established through protein-protein interaction and microRNA-mRNA networks. Within the microRNA-mRNA network, two DEMis (hsa-miR-423-5p and hsa-miR-564) potentially regulated TP53, rendering them potential CAD biomarkers, as indicated by their area under the curves (AUC) surpassing 0.6. Validation experiments corroborated an AUC of 0.7002 for hsa-miR-423-5p and 0.7261 for hsa-miR-564, highlighting its protective association with CAD. Combining hsa-miR-423-5p, hsa-miR-564, total cholesterol (TC), high-density lipoprotein-cholesterol (HDL-C), low-density lipoprotein-cholesterol (LDL-C), white blood cells (WBC) achieved an area under the receiver operating characteristics curve of 0.783. A CAD-associated gene set was identified, with TP53 as the central hub. Hsa-miR-564 emerged as a potential protective factor against CAD.
BACKGROUND: Although natural killer (NK) cells play a crucial role in antitumor immunity, the metabolic changes driving their dysfunction in lung adenocarcinoma remain poorly understood. This study investigates how these metabolic modifications impact NK cell function within the lung adenocarcinoma microenvironment. METHODS: A total of 13 pairs of lung adenocarcinoma samples were obtained from The Cancer Genome Atlas. Differential gene expression, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment, and single-cell metabolic quantification analyses were used to characterize the transcriptomic, pathway, and metabolic signatures of NK cells. The developmental trajectory was reconstructed via pseudotime analysis. The fatty acid-binding protein 4 (FABP4) and spondin2 (SPON2) expression was examined using immunofluorescence (IF) and immunohistochemistry (IHC) in patients with lung adenocarcinoma. In NK cells with FABP4 downregulation, FABP4 function was analyzed using antibody-independent cell-mediated cytotoxicity assays, flow cytometry (FCM), and liquid chromatography-mass spectrometry. RESULTS: The number of NK cells was significantly decreased in the lung adenocarcinoma microenvironment. FABP4 and SPON2 expression was significantly lower in NK cells within tumor tissues than in the adjacent tissues. FABP4 expression was significantly lower in tumor tissues than in the adjacent tissues, whereas no significant difference in SPON2 expression was observed. The cytotoxic function of NK cells with decreased FABP4 levels was impaired. Non-targeted lipid metabolism analysis indicated that differentially expressed lipids in NK cells with low FABP4 levels were functionally enriched in the glycerophospholipid metabolism pathway compared to those in normal NK cells. CONCLUSIONS: The study findings present new evidence showing that low FABP4 and SPON2 gene expression may impair NK cell maturity by affecting lipid metabolism in lung adenocarcinoma. These results provide a new perspective on restoring immune function in patients with lung cancer.
The systematic identification and functional characterization of noncanonical translation products, such as novel peptides, will facilitate the understanding of the human genome and provide new insights into cell biology. Here, we constructed a high-coverage peptide sequencing reference library with 11,668,944 open reading frames and employed an ultrafiltration tandem mass spectrometry assay to identify novel peptides. Through these methods, we discovered 8945 previously unannotated peptides from normal gastric tissues, gastric cancer tissues and cell lines, nearly half of which were derived from noncoding RNAs. Moreover, our CRISPR screening revealed that 1161 peptides are involved in tumor cell proliferation. The presence and physiological function of a subset of these peptides, selected based on screening scores, amino acid length, and various indicators, were verified through Flag-knockin and multiple other methods. To further characterize the potential regulatory mechanisms involved, we constructed a framework based on artificial intelligence structure prediction and peptide‒protein interaction network analysis for the top 100 candidates and revealed that these cancer-related peptides have diverse subcellular locations and participate in organelle-specific processes. Further investigation verified the interacting partners of pep1-nc-OLMALINC, pep5-nc-TRHDE-AS1, pep-nc-ZNF436-AS1 and pep2-nc-AC027045.3, and the functions of these peptides in mitochondrial complex assembly, energy metabolism, and cholesterol metabolism, respectively. We showed that pep5-nc-TRHDE-AS1 and pep2-nc-AC027045.3 had substantial impacts on tumor growth in xenograft models. Furthermore, the dysregulation of these four peptides is closely correlated with clinical prognosis. Taken together, our study provides a comprehensive characterization of the noncanonical proteome, and highlights critical roles of these previously unannotated peptides in cancer biology.
Multiplexed bimolecular profiling of tissue microenvironment, or spatial omics, can provide deep insight into cellular compositions and interactions in healthy and diseased tissues. Proteome-scale tissue mapping, which aims to unbiasedly visualize all the proteins in a whole tissue section or region of interest, has attracted significant interest because it holds great potential to directly reveal diagnostic biomarkers and therapeutic targets. While many approaches are available, however, proteome mapping still exhibits significant technical challenges in both protein coverage and analytical throughput. Since many of these existing challenges are associated with mass spectrometry-based protein identification and quantification, we performed a detailed benchmarking study of three protein quantification methods for spatial proteome mapping, including label-free, TMT-MS2, and TMT-MS3. Our study indicates label-free method provided the deepest coverages of ∼3500 proteins at a spatial resolution of 50 μm and the highest quantification dynamic range, while TMT-MS2 method holds great benefit in mapping throughput at >125 pixels per day. The evaluation also indicates both label-free and TMT-MS2 provides robust protein quantifications in identifying differentially abundant proteins and spatially covariable clusters. In the study of pancreatic islet microenvironment, we demonstrated deep proteome mapping not only enables the identification of protein markers specific to different cell types, but more importantly, it also reveals unknown or hidden protein patterns by spatial coexpression analysis.
Multiplexed bimolecular profiling of tissue microenvironment, or spatial omics, can provide deep insight into cellular compositions and interactions in healthy and diseased tissues. Proteome-scale tissue mapping, which aims to unbiasedly visualize all the proteins in a whole tissue section or region of interest, has attracted significant interest because it holds great potential to directly reveal diagnostic biomarkers and therapeutic targets. While many approaches are available, however, proteome mapping still exhibits significant technical challenges in both protein coverage and analytical throughput. Since many of these existing challenges are associated with mass spectrometry-based protein identification and quantification, we performed a detailed benchmarking study of three protein quantification methods for spatial proteome mapping, including label-free, TMT-MS2, and TMT-MS3. Our study indicates label-free method provided the deepest coverages of ~3500 proteins at a spatial resolution of 50 µm and the highest quantification dynamic range, while TMT-MS2 method holds great benefit in mapping throughput at >125 pixels per day. The evaluation also indicates both label-free and TMT-MS2 provide robust protein quantifications in identifying differentially abundant proteins and spatially co-variable clusters. In the study of pancreatic islet microenvironment, we demonstrated deep proteome mapping not only enables the identification of protein markers specific to different cell types, but more importantly, it also reveals unknown or hidden protein patterns by spatial co-expression analysis.