PubMed HealthSearch

Biomedical subjects

Yuzhen Ye

Publications and source records attributed to Yuzhen Ye.

2 recordsLinked to original sources

A novel transformer model of protein domains for viral taxonomy classification.

MOTIVATION: Viruses with carefully curated taxonomic assignments (such as those in the ICTV taxonomy) still represent only a small fraction of viruses identified through sequencing data from virome or microbiome projects. It is therefore critical to develop methods that can assign viruses at multiple taxonomic ranks, so that a virus deemed novel at a given rank may still be placed into a higher-level taxon. Sequence-similarity-based approaches can classify viruses that share substantial genomic similarity with known viruses (e.g. those belonging to the same species or genus); however, their performance drops significantly when applied to more divergent viruses. Recent deep learning models, such as ViTax, which utilize DNA language models, aim to address these limitations, but their performance also degrades when applied to novel viruses lacking genus-level similarity to known references. Proteins are more conserved than genomic sequences, and the multiple proteins encoded by a virus can be leveraged to reveal evolutionary relationships among viruses. RESULTS: We propose a new tool, D2T (Domain-to-Taxonomy), that leverages recent advances in protein language models to improve viral taxonomic assignment. D2T represents a virus as a sequence of protein domain tokens and learns a transformer-based model for taxonomic classification. Experiments on multiple closed-set and open-set datasets show that D2T excels at assigning higher-level taxonomic labels (family and above). Furthermore, by combining D2T with Kraken2, which performs well at the genus level, the hybrid method (K+D2T) achieves accurate viral taxonomic classification across multiple taxonomic ranks. AVAILABILITY AND IMPLEMENTATION: D2T is available as a GitHub repository at https://github.com/mgtools/D2T.

Viruses

Rapid Generation of Reverse Genetics Systems for Coronavirus Research and High-Throughput Antiviral Screening Using Gibson DNA Assembly.

Coronaviruses (CoVs) pose a significant threat to human health, as demonstrated by the COVID-19 pandemic. The large size of the CoV genome (around 30 kb) represents a major obstacle to the development of reverse genetics systems, which are invaluable for basic research and antiviral drug screening. In this study, we established a rapid and convenient method for generating reverse genetic systems for various CoVs using a bacterial artificial chromosome (BAC) vector and Gibson DNA assembly. Using this system, we constructed infectious cDNA clones of coronaviruses from three genera: human coronavirus 229E (HCoV-229E) of the genus Alphacoronavirus, mouse hepatitis virus A59 (MHV-59) of Betacoronavirus, and porcine deltacoronavirus (PDCoV-Haiti) of Deltacoronavirus. Since beta coronaviruses including severe acute respiratory syndrome coronavirus (SARS-CoV), severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), and Middle East respiratory syndrome coronavirus (MERS-CoV) represent major human pathogens, we modified the infectious clone of the beta coronavirus MHV-A59 by replacing its NS5a gene with a fluorescent reporter gene to create a system suitable for high-throughput drug screening. Thus, this study provides a practical and cost-effective approach to developing reverse genetics platforms for CoV research and antiviral drug screening.

Reverse Genetics