PubMed HealthSearch

Biomedical subjects

Yue Yang

Publications and source records attributed to Yue Yang.

6 recordsLinked to original sources

Predicting risk of ischemic stroke: A transformer model using genomic data.

BACKGROUND AND OBJECTIVE: Ischemic stroke is a leading cause of mortality and long-term disability worldwide. Genetic factors contribute to IS susceptibility, yet conventional polygenic risk score approaches are primarily based on additive effects and may not fully capture non-linear relationships or positional context and interactions among genetic variants. This study aimed to develop and evaluate a transformer-based genomic model incorporating position-wise genotype embedding for IS risk prediction. METHODS: We conducted a genome-wide association study using the UK Biobank dataset to identify IS-associated loci. Gene prioritisation was subsequently performed using tissue-specific expression quantitative trait locus-based Mendelian randomisation and colocalization analyses in whole blood and brain cortex. We then developed a transformer-based model that encoded genotype and SNP-position information using a position-wise embedding layer. Model performance was evaluated across three UK Biobank control definitions and externally assessed in the independent All of Us cohort. Performance metrics included the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score. RESULTS: Across the three UK Biobank control definitions, the proposed method achieved the numerically highest discrimination among the evaluated models, with AUROCs of 0.8109, 0.7843, and 0.7468 using MRF-negative, combined, and MRF-positive controls, respectively. In the external All of Us cohort, the proposed method achieved an AUROC of 0.7251 and retained the highest AUROC among the evaluated models. In a separate incident-stroke survival analysis, medium- and high-score groups had hazard ratios of 1.13 and 1.21, respectively, relative to the low-score group. A total of 18 IS-associated loci were identified. Among the tissue-specific MR results, EDEM2 in the brain cortex remained significant after Bonferroni correction, while DCHS2 showed a nominal association. CONCLUSIONS: The proposed transformer-based framework provides a genomic modelling approach that achieved the highest discrimination among the evaluated models in this study and retained comparative performance in an independent external cohort. In further applications, integrating this genomic framework with conventional clinical, lifestyle, and environmental risk factors may support more comprehensive and personalised IS risk assessment. Prospective, population-representative, and multi-ancestry validation will be important to establish its potential role in future prevention-oriented risk management.

Genomics and bioinformatics

Mega-enhancers compartmentalize transcriptionally active long genes in the brain.

Exceptionally long genes and cis-regulatory enhancers are selectively activated in mammalian brain neurons, and these loci are mutation hotspots in neurological disorders. However, the organization of these large genomic elements at the level of chromosome folding, beyond local enhancer-promoter interactions, remains poorly understood. Here we report the discovery of a genomic subcompartment in the mouse cerebellum formed by near-megabase-long enhancers and their associated long genes encoding synaptic or signalling proteins. Genomic regions within this subcompartment are enriched in the outer half of the nucleus, whereas other transcriptionally active structures are enriched in the nuclear interior. Using an in vivo CRISPR genetic mini screen, we uncover a specific role for the transcription factor Etv1 in coupling the compartmentalization of neuronal long genes with their expression. Together, our study defines mechanisms that organize transcriptionally active genes across chromosomes in the mammalian brain.

Animals

Combinatorial genome engineering of pseudorabies virus Bartha by developing a reverse genetic system based on three overlapping genomic segments.

INTRODUCTION: The 138-kilobase genome of pseudorabies virus vaccine strain Bartha K61 harbors many nonessential genes for replication and exhibits remarkable capacity for incorporating foreign genes for therapeutic applications. However, the large size of the Bartha genome complicates its efficient engineering. OBJECTIVES: Development of a reverse genetic system for pseudorabies virus Bartha based on three overlapping genomic segments to facilitate multiplex genome engineering. METHODS: The 138-kb genome of Bartha was split into three overlapping segments (42 kb, 43 kb, and 53 kb), each cloned in a bacterial artificial chromosome (BAC) to facilitate genome engineering. The infectious virus was reconstituted by transfecting the 3 genomic fragments released from the BACs into Vero cells in which a complete virus genome was assembled using 2-kb overlaps between adjacent pieces. RESULTS: Employing the reverse genetic system, we individually deleted 15 candidate nonessential genes and confirmed that 10 were dispensable for viral growth in cell culture. Deletion of 7 nonessential genes had no impact on viral growth, whereas UL47 deletion reduced viral growth rate and deletions of UL44, UL47, or US3 resulted in smaller viral plaques. A total of 45 viral genomes with double deletions of nonessential genes were constructed, among which 22 were successfully rescued into infectious virions. Fifteen double-deletion mutant viruses had a viral titer comparable with the wild-type Bartha, while the remaining 7 showed a lower titer. Additionally, expressions of the mNeonGreen reporter gene at nonessential gene loci were evaluated. Cells infected with recombinant viruses carrying mNeonGreen at 8 loci showed strong green fluorescence, whereas those with mNeonGreen at 2 loci exhibited very weak fluorescence. CONCLUSION: The reverse genetic system developed in this study enables rapid and combinatorial engineering of viruses with the large DNA genome, and will accelerate development of large DNA virus-based therapeutics including live-attenuated vaccines, vector vaccines, and oncolytic herpesviruses.

Herpesvirus 1, Suid

RNA-DNA hybrid binding domain broadens the editing window of base editors.

Adenine base editors (ABEs) and cytosine base editors (CBEs) are prominent tools for precise genome editing but are hindered by limited editing activity at positions proximal to the protospacer adjacent motif (PAM). This study investigates the potential of enhancing base editors editing activity by fusing them with RNA-DNA hybrid binding domains (RHBDs). Specifically, fusing ABE8e with the RHBD of Homo sapiens RNaseH1 (RHBD1) significantly increased A-to-G editing efficiency in the PAM-proximal region (A9-A15) by up to 3.5-fold, while reducing off-target cytosine editing. Additionally, RHBD1 is compatible with ABEmax, BE4max, and dual base editor (eA&C-BEmax), enhancing their editing activity at the PAM-proximal bases. Notably, RHBD1-fused BE4max led to a 3.1-fold improvement in C-to-T editing efficiency at PAM-proximal region (C9-C12). Furthermore, we demonstrated that RHBD1-fused ABE8e could effectively edit disease-related single nucleotide variations (SNVs) in human cells and validated its efficacy in adult mouse liver. These findings highlight the significance of the RHBD in expanding editing window and the applicability of base editors for gene therapy and disease modeling.

Gene Editing

Genome-wide association study reveals that TaODORANT1 negatively contributes to thousand grain weight by affecting starch synthesis in wheat.

Thousand grain weight (TGW) is one of the most important factors that control grain weight and crop yield. To date, dozens of wheat genes related to TGW have been isolated; however, the underlying molecular mechanisms governing grain development in wheat (Triticum aestivum) remain largely unknown. Benefiting from whole-genome resequencing and genome-wide association study, we identified an R2R3-type myeloblastosis (MYB) transcription factor, TaODORANT1, which was tightly associated with TGW. TaODORANT1 was specifically and highly expressed during the wheat grain developing stage. Knockout of TaODORANT1 led to an increase in TGW and starch content, as well as affected the expression of starch synthesis-related genes. Loss of function of TaODORANT1 altered the molecular structure and physiochemical properties of grain starch. Haplotype analysis showed that favorable Hap IV of TaODORANT1-A and favorable Hap I of TaODORANT1-B were significantly associated with the production of larger grains and higher TGW, respectively. Moreover, TaODORANT1 was a crucial targeted gene continuously selected in wheat domestication and breeding, and its orthologous genes might have retained similar functions in response to grain development. Our results highlight the importance of TaODORANT1 in affecting TGW, presenting potential targets for improving yield in wheat.

Triticum

Mega-Enhancer Bodies Organize Neuronal Long Genes in the Cerebellum.

Dynamic regulation of gene expression plays a key role in establishing the diverse neuronal cell types in the brain. Recent findings in genome biology suggest that three-dimensional (3D) genome organization has important, but mechanistically poorly understood functions in gene transcription. Beyond local genomic interactions between promoters and enhancers, we find that cerebellar granule neurons undergoing differentiation in vivo exhibit striking increases in long-distance genomic interactions between transcriptionally active genomic loci, which are separated by tens of megabases within a chromosome or located on different chromosomes. Among these interactions, we identify a nuclear subcompartment enriched for near-megabase long enhancers and their associated neuronal long genes encoding synaptic or signaling proteins. Neuronal long genes are differentially recruited to this enhancer-dense subcompartment to help shape the transcriptional identities of granule neuron subtypes in the cerebellum. SPRITE analyses of higher-order genomic interactions, together with IGM-based 3D genome modeling and imaging approaches, reveal that the enhancer-dense subcompartment forms prominent nuclear structures, which we term mega-enhancer bodies. These novel nuclear bodies reside in the nuclear periphery, away from other transcriptionally active structures, including nuclear speckles located in the nuclear interior. Together, our findings define additional layers of higher-order 3D genome organization closely linked to neuronal maturation and identity in the brain.

Journal Article