PubMed HealthSearch

Biomedical subjects

Yanni Sun

Publications and source records attributed to Yanni Sun.

5 recordsLinked to original sources

Library strategies differentially shape microbial, functional, and host signals in clinical metagenomic sequencing.

Metagenomic next-generation sequencing (mNGS) is increasingly used in infectious disease diagnostics, yet how library preparation shapes the microbial, functional, and host signals recovered from clinical samples remains poorly defined. Here, we performed a within-sample parallel comparison of three mNGS library preparation strategies-DNA-based libraries (DNAlib), RNA-based libraries (RNAlib), and total nucleic acid-based libraries (TNAlib)-across a diverse range of clinical specimens spanning five sample types. Using a curated clinical infectome as a benchmark, we show that library strategies are not interchangeable but capture distinct biological dimensions of the same specimen. RNAlib provided the most comprehensive standalone recovery of the clinical infectome, with improved detection of RNA viruses and cellular pathogens, enhanced resolution of resistance and virulence signals, and preservation of infection-associated host immune signatures. DNAlib showed stronger baseline recovery of DNA viruses and broader host genome coverage, whereas the TNAlib workflow evaluated here largely behaved as an intermediate strategy rather than a consistent improvement over dedicated DNA- or RNA-based workflows. Together, these results establish that the library preparation protocol is a major determinant of how clinical mNGS data should be interpreted and provide a framework for selecting sequencing strategies according to specific diagnostic and biological questions.IMPORTANCEMetagenomic sequencing is increasingly used in infectious disease research and clinical diagnostics, but different library preparation strategies may recover fundamentally different biological signals from the same sample. These signals include not only pathogens but also background microbes, microbial functional activity, and host immune-response patterns. Here, we systematically compared DNA-, RNA-, and total nucleic acid-based metagenomic sequencing libraries using the same clinical samples processed in parallel. We found that the three strategies did not provide equivalent information. RNA-based sequencing generated the most informative single-library view of infection, particularly for RNA viruses, cellular pathogens, functional microbial signals, and host immune-response patterns. DNA-based sequencing was more effective for DNA virus and host genome recovery, whereas the total nucleic acid sequencing workflow evaluated here generally behaved as an intermediate strategy. These findings show that library preparation can substantially influence the interpretation of metagenomic data.

functional characterization

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics

ViralQC: a tool for assessing completeness and contamination of predicted viral contigs.

MOTIVATION: Viruses represent the most abundant biological entities on Earth, playing vital roles in diverse ecosystems. Cataloging viruses across various environments is essential for understanding their properties and functions. Metagenomic sequencing has emerged as the most comprehensive method for virus discovery. However, distinguishing viral sequences from the vast background of microbial organisms in metagenomic data remains a significant challenge. Existing tools experience varying degrees of false positive rates due to noise in sequencing and assembly, and the integration of proviruses into microbial genomes. This highlights the urgent need for an accurate and efficient method to evaluate the quality of viral contigs. RESULTS: To address these challenges, we introduce ViralQC, a tool designed to assess the quality of viral contigs or bins. ViralQC identifies microbial contamination within putative viral sequences using an ensemble framework powered by DNA and protein foundation models and estimates completeness by analyzing protein organization. We evaluated ViralQC on multiple datasets and compared its performance against the state-of-the-art tool, CheckV. Leveraging both DNA and protein foundation models, ViralQC achieves higher sensitivity on contamination detection for contigs longer than 10 kbp while maintaining comparable accuracy. Additionally, ViralQC delivers more accurate estimation on contigs with completeness > 50%. AVAILABILITY: The source code of ViralQC is available via: https://github.com/ChengPENG-wolf/ViralQC.

Software

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses

GiantHunter: accurate detection of giant virus in metagenomic data using reinforcement-learning and Monte Carlo tree search.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) are notable for their large genomes and extensive gene repertoires, which contribute to their widespread environmental presence and critical roles in processes such as host metabolic reprogramming and nutrient cycling. Metagenomic sequencing has emerged as a powerful tool for uncovering novel NCLDVs in environmental samples. However, identifying NCLDV sequences in metagenomic data remains challenging due to their high genomic diversity, limited reference genomes, and shared regions with other microbes. Existing alignment-based and machine learning methods struggle with achieving optimal trade-offs between sensitivity and precision. RESULTS: In this work, we present GiantHunter, a reinforcement learning-based tool for identifying NCLDVs from metagenomic data. By employing a Monte Carlo tree search strategy, GiantHunter dynamically selects representative non-NCLDV sequences as the negative training data, enabling the model to establish a robust decision boundary. Benchmarking on rigorously designed experiments shows that GiantHunter achieves high precision while maintaining competitive sensitivity, improving the F1-score by 10% and reducing computational cost by 90% compared to the second-best method. To demonstrate its real-world utility, we applied GiantHunter to 60 metagenomic datasets collected from six cities along the Yangtze River, located both upstream and downstream of the Three Gorges Dam. The results reveal significant differences in NCLDV diversity correlated with proximity to the dam, likely influenced by reduced flow velocity caused by the dam. These findings highlight GiantHunter's potential to advance our understanding of NCLDVs and their ecological roles in diverse environments. AVAILABILITY AND IMPLEMENTATION: The source code of GiantHunter is available via: https://github.com/FuchuanQu/GiantHunter.

Metagenomics