PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “WGS sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans↗

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum↗

Nanopore-based, long-range Parvovirus B19 amplicon sequencing for near-whole genome characterization.

BACKGROUND: Whole-Genome Sequencing (WGS) enables monitoring of genomic variation and evaluation of diagnostic PCR assays. However, WGS data for Parvovirus B19 (B19V) remains limited despite its relevance for clinical care and transfusion safety. To increase the availability of high-quality B19V genomic data, a near-WGS protocol was developed and validated. METHOD: The protocol combines long-range PCR to generate a 4.6-kb amplicon, covering &#x223c;82% of the B19V genome, with Oxford Nanopore sequencing. Validation was performed using six reference samples and nineteen B19V-positive donor plasma samples. RESULTS: After quality control, samples achieved a median sequencing depth of 152x. Sequences generated from the six reference samples showed 100% concordance with previously published data. Genomic analysis of donor samples explained atypical amplification profiles observed during routine PCR screening. CONCLUSION: The newly developed protocol provides a scalable method for B19V genome characterization, enabling assessment of oligonucleotide-binding regions for PCR assay monitoring and facilitating the generation of genomic data for future epidemiological investigations.

PCR assay monitoring↗

WxS-QC-a quality control pipeline for human germline short-variant Whole-Genome and Whole-Exome cohorts for population-scale analyses.

SUMMARY: Whole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. In this paper, we present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human germline short-variant WGS and WES cohorts for population-scale analyses. Our pipeline is suitable for both rare-variant discovery and common-variant association studies. It is based on deeply refactored gnomAD v3 and v4 quality control pipelines, contains several methods we have developed de novo, and is aligned with current best practices in WGS/WES germline cohort QC. We provide all methods in a single codebase, aligned to work together and controlled via a single YAML config, with automatic export of resulting graphs and summary tables, excellent performance and scalability, and comprehensive documentation. The pipeline can run in any UNIX-like environment and can efficiently process cohorts of up to 200&#x2009;000 whole-exome samples, with the potential to handle bigger datasets. AVAILABILITY AND IMPLEMENTATION: The pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. The detailed description of the pipeline is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md. We also provide an open dataset with all required metadata, which is available at https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.

Humans↗

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n&#x202f;=&#x202f;53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans↗

Whole-genome surveillance supports hazard profiling of Escherichia coli lineages in recycled water treatment systems.

UNLABELLED: The use of treated wastewater is increasingly important for sustainable water management under a changing climate, yet conventional monitoring based on Escherichia coli enumeration provides limited insight into strain diversity and associated public health hazards. Here, we applied longitudinal whole-genome sequencing (WGS) to 180 E. coli isolates collected across the treatment continuum of a recycled water facility, from influent to final effluent. Genomic analysis revealed extensive strain-level heterogeneity, comprising 88 sequence types across eight phylogroups, with greater diversity in influent than in treated effluent. Phylogenetic comparisons with contextual Australian genomes indicated clustering with strains associated with companion animals, wild birds, humans, and livestock, suggesting multiple potential source reservoirs rather than a single dominant origin, although source contributions were not definitive. Despite a >90% reduction in total E. coli loads, isolates recovered from upstream and downstream stages exhibited broadly comparable virulence factor and antimicrobial resistance gene (ARG) profiles, suggesting that, within the cultured isolate collection, reductions in abundance exceeded shifts in genomic composition. To assess operational relevance, we prototyped a genomics-informed hazard framework integrating virulence determinants, ARGs, plasmid-associated mobility, and reuse-specific exposure context. Using this framework, 92.8% of isolates were classified as low hazard, and 7.2% as moderate hazard, with no isolates meeting criteria for high or critical hazard classifications. These findings demonstrate that genomic profiling of indicator organisms can reveal population structure and hazard heterogeneity not captured by conventional enumeration alone, and can provide a practical basis for incorporating genomic information into hazard-informed monitoring of recycled water systems. IMPORTANCE: Routine recycled water monitoring relies largely on culture-based E. coli counts, which indicate regulatory compliance but provide limited insight into strain diversity, persistence, and genomic characteristics relevant to public health. Using longitudinal whole-genome sequencing, we show that genetically distinct E. coli lineages, including isolates carrying combinations of virulence and antimicrobial resistance determinants, can persist through advanced treatment despite substantial reductions in overall E. coli loads. While most isolates were classified as low genomic hazard and no high- or critical-hazard isolates were detected, these findings demonstrate that conventional enumeration alone cannot distinguish between genetically diverse lineages with differing hazard potential in highly treated systems. By integrating genomic data into a hazard classification framework, this study demonstrates an applied approach to contextualize E. coli detections and distinguish low-risk background populations from isolates with elevated genomic hazard profiles. This work supports the use of genomic profiling of indicator organisms to improve surveillance, inform treatment performance assessment, and enable more risk-based management of recycled water systems.

Escherichia coli↗

Evaluation of amplicon-based nanopore sequencing for foot-and-mouth disease viruses in clinical and environmental samples.

Foot-and-mouth disease (FMD) causes severe global economic loss, necessitating rapid viral characterization. Nanopore sequencing provides a simple, real-time workflow suitable for on-site outbreak response, addressing the limitations of conventional methods. In this study, we optimized a previously published amplicon-based protocol and used this method to characterize a diverse range of samples (vesicular fluid, epithelium, serum, nasal/oral swabs, and environmental samples) collected during FMD outbreaks in 2025 in the Republic of Korea. Of the 129 samples collected, we successfully recovered complete genomes from 37 samples and VP1 sequences from 85 samples. Amplifying the S-fragment in isolation and separately barcoding each pool of PCR amplicons markedly improved sequence recovery. Furthermore, sequencing success depended on viral load and sample type. Based on comparisons with real-time RT-PCR results, whole-genome sequence (WGS) recovery exceeded 77.3% at cycle threshold (Ct) values &#x2264;25 across all clinical samples. In the Ct > 30 category, serum samples yielded the highest WGS recovery rates (44.4%). This rate was markedly higher than the success rates observed for epithelium (20.0%) and nasal swabs (9.1%), whereas oral swabs and environmental samples failed to yield any sequences (0%). However, VP1 recovery from environmental samples reached 80% at Ct &#x2264; 30 (8/10), providing an approach to enable non-invasive monitoring. These findings demonstrate that amplicon-based nanopore sequencing is a practical method for the rapid generation of genomic data during FMD outbreaks.IMPORTANCEAlthough rapid detection and genomic data analysis are crucial for effective foot-and-mouth disease (FMD) control, the collection of these data can be challenging for certain sample types and impacted by reduced viral loads that result from nationwide FMD vaccination. This study provides a practical solution through large-scale evaluation of an optimized amplicon-based nanopore sequencing protocol to enhance the sequencing success rates for both clinical and environmental samples. Using a modified protocol to enhance genome recovery, we demonstrated that sequence data could be retrieved from diverse sample types (even with high real-time RT-PCR cycle threshold values). We identified serum as the most suitable sample, with environmental sample sequencing allowing for non-invasive monitoring during outbreaks. These results support the use of nanopore sequencing for rapid genomic analysis, particularly in outbreak responses, such as rapid surveillance, emergency vaccine selection, and epidemiological monitoring.

Foot-and-Mouth Disease↗

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7&#xa0;Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (&#x223c;33.25 million SNPs and&#xa0;&#x223c;&#xa0;1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

Development and evaluation of an ARTIC-based amplicon sequencing assay for whole-genome characterization of respiratory syncytial virus.

Respiratory syncytial virus (RSV), a ~15.2 kb negative-sense RNA virus, causes acute respiratory infections in infants and older adults. Its two subtypes, RSV-A and RSV-B, evolve rapidly, making ongoing monitoring of circulating strains essential. The Georgia Public Health Laboratory (GPHL) developed and evaluated an amplicon-based whole-genome sequencing (WGS) assay for RSV surveillance. A total of 214 de-identified remnant clinical specimens (102 RSV-A and 112 RSV-B) with RT-PCR Cq values <31 were included. RSV genomes were amplified using ARTIC-style and custom primer sets, with the ARTIC set showing superior performance. Libraries were prepared using a modified Illumina COVIDSeq protocol, sequenced on NextSeq 1000/2000 instruments, and analyzed using the GPHL-RSV-PIPE bioinformatics pipeline. Among genomes meeting validation criteria, sequencing depth was slightly higher for RSV-A (median 53,433&#xd7;; mean 51,076&#xd7;) than RSV-B (median 49,699&#xd7;; mean 46,945&#xd7;), whereas genomic coverage was slightly lower for RSV-A (median 97.5%; mean 96.6%) than RSV-B (median 98.3%; mean 97.6%). Predominant lineages were A.D.3.1 and A.D.5.2 for RSV-A and B.D.E.1 for RSV-B. For RSV-A, the assay showed 92.8% accuracy, 96.2% sensitivity, 87.2% specificity, 92.6% positive predictive value, and 93.2% negative predictive value. Intra- and inter-run precision assessed using 16 and 53-57 genomes, respectively, showed nearly 100% consensus genome identity with 0-5 nucleotide differences. Specificity testing of 31 non-RSV specimens produced no false-positive detections. Limits of detection were 4.4 TCID50/mL for RSV-A and 18.6 TCID50/mL for RSV-B. These results demonstrate that the ARTIC-based RSV WGS assay enables near real-time surveillance and strengthens data-driven public health responses to future outbreaks.IMPORTANCERSV, with two major subtypes, RSV-A and RSV-B, causes acute respiratory infections that can be severe in infants under 6 months and older adults. Current RSV surveillance at the GPHL relies on the Thermo Fisher TaqMan Gene Expression Capillary assay, which detects and subtypes RSV but lacks resolution for lineage classification and identification of emerging variants. To address this critical gap, GPHL developed and evaluated an amplicon-based WGS assay using 214 de-identified RSV clinical specimens. Genomes were amplified using ARTIC-style and custom-primer sets, with ARTIC primers showing superior performance. The assay demonstrated strong sequencing depth, genomic coverage, specificity, repeatability, reproducibility, and low limits of detection. RSV lineages were accurately determined based on genetic variation. These results establish that the ARTIC-based WGS assay enables near real-time genomic surveillance, supporting monitoring of circulating RSV strains and informing data-driven public health responses.

bioinformatics pipeline↗

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis↗

Genomic Insights into Mammaliicoccus sciuri from Subclinical Bovine Mastitis to Unveil Key Resistance, Virulence, Biofilm and Adaptation Traits.

The Mammaliicoccus sciuri (M. sciuri), is recognized as a reservoir of antimicrobial resistance (AMR) genes, poses challenges in the Indian dairy sector where antibiotic use is poorly regulated. This study aimed to genomically characterize M. sciuri (formerly Staphylococcus sciuri) isolates recovered from subclinical mastitis (SCM) cattle milk. A total of 128 composite (quarter-wise pooled) milk samples were collected from 199 households (HH) across 16 epiunits /villages in four blocks of Chikkaballapur district, Karnataka, India. Of these, 36 milk samples (28.13%, 36/128; 95% CI: 21.06&#x2013;36.46%) were diagnosed with SCM using the California Mastitis Test (CMT) and bacteriological culture yielded 113 isolates (88.28%; 113/128; 95% CI: 81.56&#x2013;92.77%) were phenotypically identified as Staph spp. Through molecular technique PCR targeting the gap gene, two isolates (1.77%; 2/113; 95% CI: 0.49&#x2013;6.22%) from Hosuru and Gattamaranahalli epiunits were confirmed as M. sciuri and both isolates were mecA-positives indicating methicillin resistance. Whole genome sequencing (WGS) identified 36&#x2013;37 resistance genes (mecA and blaZ), conferring resistance to &#x3b2;-lactams, macrolides, fluoroquinolones and aminoglycosides. Horizontal gene transfer (HGT) was evidenced by diverse mobile genetic elements (MGEs) such as SCCmec variants, insertion sequences, transposons (IS3, IS6, IS256, and IS1182) and plasmids (Rep1, Rep13, RepUS5 and RepUS43). Virulence profiling uncovered biofilm-associated genes (ica, bap) and heavy metal resistance operons (ars, cop, znu) suggesting mechanisms for environmental persistence and co-selection of resistance traits. Phylogenetic analysis of 99 global isolates revealed host-and geography-specific clustering with Indian isolates occupying distinct evolutionary niches. These findings highlights its possible role as an AMR reservoir and also in bovine mastitis.

Animals↗

Backtracking Cell Phylogenies in the Human Brain with Somatic Mosaic Variants.

Somatic mosaic variants, and especially somatic single nucleotide variants (sSNVs), occur in progenitor cells in the developing human&#xa0;brain frequently enough to provide permanent, unique, and cumulative markers of cell divisions and clones. Here, we describe an experimental workflow to perform lineage studies in the human brain using somatic variants. The workflow consists in two major steps: (1) sSNV calling through&#xa0;whole-genome sequencing&#xa0;(WGS) of bulk (non-single-cell) DNA extracted from human fresh-frozen tissue&#xa0;biopsies, and (2) sSNV validation and&#xa0;cell phylogeny deciphering through&#xa0;single nuclei whole-genome amplification (WGA) followed by&#xa0;targeted sequencing&#xa0;of sSNV loci.

Humans↗

Global inequities in hepatitis B and C genomic surveillance revealed through an interactive data integration dashboard.

OBJECTIVES: To assess global disparities in hepatitis B virus (HBV) and hepatitis C virus (HCV) genomic surveillance and to develop an integrated platform that links genomic data with epidemiological burden. STUDY DESIGN: Retrospective observational analysis. METHODS: We reviewed existing viral genomic repositories to identify structural and analytical limitations. Subsequently, we integrated 10&#xa0;996 HBV and 3533 HCV whole-genome sequences (WGS) from public databases with Global Burden of Disease (GBD) estimates to quantify inequities in genomic surveillance across countries and genotypes. Using these data, we developed the open-access Hepatitis Dashboard, incorporating >14&#xa0;000 sequences from 141 countries with GBD metrics to evaluate representativeness and sequencing coverage relative to disease burden. RESULTS: Marked inequities in hepatitis genomic surveillance were identified. Despite increasing HBV- and HCV-associated mortality, virus sequence availability remains geographically and genotypically skewed-dominated by China and the United States, with substantial underrepresentation of HBV genotype E and HCV genotypes 5 and 8. Many high-endemic countries in Africa and the Western Pacific remain severely undersampled. We detected circulating antiviral drug-resistance mutations and developed a burden-adjusted sequencing coverage metric, revealing that several high-burden countries, including China, Nigeria and India, are among the least represented in global genomic datasets. Projections to 2030 indicate that neither HBV nor HCV are currently on track to meet WHO elimination targets. CONCLUSIONS: The Hepatitis Dashboard provides an integrated, continuously updated resource that links genomic and epidemiological data to quantify and visualise global surveillance gaps. This analysis highlights a critical disconnect between sequencing efforts and public health needs, which may limit the effectiveness of surveillance-informed strategies to support progress toward WHO 2030 elimination goals. By enabling burden-adjusted prioritisation and longitudinal tracking of genomic coverage, the platform supports evidence-based sampling strategies, equitable resource allocation, and monitoring of global progress toward hepatitis elimination.

Humans↗

Phylogenetics and genomic variation of Hepatocystis isolated from shotgun sequencing of wild primate hosts.

Hepatocystis are apicomplexan parasites nested within the Plasmodium genus that infect primates and other vertebrates, yet few isolates have been genetically characterized. Using taxonomic classification and mapping characteristics, we searched for Hepatocystis infections within publicly available, blood-derived whole genome sequence (WGS) data from 326 wild non-human primates (NHPs) in 17 genera. We identified 37 Hepatocystis infections in Papio cynocephalus (yellow baboons) and four species of Chlorocebus monkeys (grivets, green monkeys, vervet monkeys, and malbroucks) sampled from locations in west, east, and south Africa. Hepatocystis cytb sequences from Papio and Chlorocebus hosts each clustered within host species among previously reported isolates from other NHP taxa. Utilizing the low-coverage sequence data (0.11-0.76X per sample) recovered across the nuclear Hepatocystis genome, we identified 349,893 polymorphic sites. Principle components analysis based on genotype likelihoods across all samples showed evidence for population structure by primate host species. Across the genome, windows of high SNP density revealed candidate hypervariable loci including Hepatocystis-specific gene families possibly involved in immune evasion and genes that may be involved in adaptation to their insect vector and hepatocyte invasion. Overall, this work demonstrates how WGS data from wild NHPs can be leveraged to study the evolution of apicomplexan parasites and potentially test for association between host genetic variation and parasite infection.

Animals↗

4CMenB vaccine coverage of invasive serogroup B meningococci collected in Belgium between 2016 and 2022.

Neisseria meningitidis infections can cause life-threatening meningitis and septicemia. In Europe, serogroup B (MenB) is the leading cause of invasive meningococcal disease (IMD), particularly in young children. Genomic surveillance of circulating MenB strains through whole genome sequencing (WGS) provides a powerful tool to assess the potential impact of vaccination strategies, including the 4CMenB vaccine, which is available for infants from 2 months of age. Here, we present a retrospective WGS-based analysis of clinical MenB IMD cases (n&#x2009;=&#x2009;311) recovered in Belgium from 2016 to 2022 by the Belgian National Reference Center. High-quality WGS data were obtained for 281 of these strains, demonstrating high genetic diversity of the antigen targets included in the 4-component meningococcal serogroup B vaccine 4CMenB (fHbp, PorA, NHBA and NadA) and at the 4CMenB Antigen Sequence Types (BAST) level. Novel antigen combinations, not yet assigned a BAST ID, were detected in 23.5% of isolates. Vaccine coverage was predicted using the Genetic Meningococcal Antigen Typing System (gMATS) and the Meningococcal Deduced Vaccine Antigen Reactivity (MenDeVAR) index. Of the 281 strains, 79.5% (lower limit-upper limit: 68.0-91.5%) were predicted to be covered by the vaccine by gMATS, and 80.7% (lower limit-upper limit: 66.5-95.4%) by MenDeVAR. No evidence of variation in vaccine coverage was found throughout the study period nor between different age groups, demonstrating the broad applicability of 4CMenB. This study highlights the benefits of a pathogen surveillance program and the need for experimental characterization of continuously evolving antigenic subvariants of Neisseria meningitidis.

Humans↗

Clinical and genomic features of mitis group streptococcal bacteremia in patients with febrile neutropenia.

BACKGROUND: Viridans group streptococci (VGS) can cause the life-threatening viridans streptococcal shock syndrome (VSSS) in patients with febrile neutropenia (FN). The Mitis group, a major subgroup of VGS, is frequently implicated in these severe infections, but its specific clinical and genomic characteristics remain incompletely characterized, particularly in patients with FN. This study aimed to systematically describe these features in this population. METHODS: In this single-center retrospective study, we compared the clinical data and whole-genome sequencing (WGS) results of Mitis group streptococcal isolates from patients with and without FN. Virulence-associated and antimicrobial resistance genes were initially screened using a reference-based approach, followed by assembly-based reanalysis and manual sequence validation. RESULTS: Compared with the non-FN cohort (n&#x2009;=&#x2009;34), the FN cohort (n&#x2009;=&#x2009;61) was significantly younger, had a higher prevalence of hematologic malignancy, and more frequently presented with primary bacteremia. VSSS occurred exclusively in the FN group (11.5%) and was associated with high mortality (14-day mortality, 42.9%), which did not correlate with in vitro antimicrobial susceptibility. Genomic analyses revealed marked diversity among isolates. Initial screening suggested variable detection of several virulence-associated loci, including pavA, slrA, and rfb-related loci; however, subsequent assembly-based analyses indicated that many apparent absences were attributable to extreme allelic divergence rather than true gene loss. No single virulence determinant clearly segregated with clinical severity. CONCLUSIONS: Mitis group bacteremia in patients with FN appears to be characterized by distinct clinical features and marked genomic diversity. Our findings suggest that the development of severe disease, including VSSS, may not be explained by microbial factors alone and potentially reflects complex host-pathogen interactions. CLINICAL TRIAL: Not applicable.

Humans↗

Detection and characterization of antiviral-resistant viruses during the influenza season of 2024-25.

UNLABELLED: During the high severity season of 2024-25, CDC with public health partners sequenced and analyzed genomes of >10,000 influenza viruses for antiviral resistance markers. Available sequence-flagged and representative viruses were tested with antivirals using in vitro assays. In the US, three oseltamivir-resistant A(H3N2) viruses had treatment-emergent neuraminidase (NA) mutations, either E119V or R292K. Oseltamivir-resistant A(H1N1)pdm09 viruses with NA-H275Y were detected in 15 states, albeit at a low frequency (0.53%). They belonged to several phylogenetic groups, with hemagglutinin (HA) subclade D.3.1 combined with either NA subclade D.1 or D.2 being most common. Based on shared sequence data, nearly all H275Y viruses from Australia, Canada, and Chile also belonged to these HA and NA subclades. Conversely, most H275Y viruses (68/81) from China belonged to HA subclade C.1.9 and NA subclade D and shared the permissive mutation R257K. Influenza polymerase acidic (PA) mutations conferring 4- to 92-fold decreased baloxavir susceptibility were detected in nine influenza A viruses. Viruses with PA-I38T showed mild attenuation of replicative fitness in three cell lines. Based on available data, NA-H275Y and PA-I38T viruses were collected from patients with no exposure to antivirals. Baseline susceptibility to all US-approved influenza antivirals remained largely unchanged compared to previous seasons. All swine-origin viruses detected in the US had adamantane resistance-conferring marker, M2-S31N, but remained susceptible to other approved antivirals. Monitoring antiviral susceptibility has substantially improved with increased sequencing capacities and bioinformatic support at public health laboratories. Information gained through influenza surveillance has been used to guide recommendations on antiviral use. IMPORTANCE: Circulation of influenza viruses with reduced susceptibility to antivirals can diminish the usefulness of medications prescribed for influenza. This study informs on the prevalence of drug-resistant influenza viruses in the US during the high severity season of 2024-25. It provides information on susceptibility profile to all approved antiviral medications and on replicative fitness of representative drug-resistant viruses. Most drug-resistant viruses were collected from patients who were not exposed to antivirals indicating their ability to transmit from human to human. Whole-genome sequence (WGS)-based analysis is the cornerstone for surveillance, and numerous laboratories have been utilizing this approach. However, CDC laboratory is the only laboratory in the US conducting phenotypic testing of circulating viruses needed to confirm the outcomes of sequence-based analysis and to identify new molecular markers of resistance. Data gathered through virologic surveillance give much-needed information on drug susceptibility of influenza viruses which are used to guide recommendations on antiviral use.

Antiviral Agents↗

Multiplex PCR assay for the rapid detection of Klebsiella pneumoniae pathotypes.

Introduction. Klebsiella pneumoniae (Kp) is a major cause of nosocomial infections, with its evolving pathotypes including multidrug-resistant, hypervirulent (hvKp) and convergent strains posing significant diagnostic and treatment challenges due to combined antimicrobial resistance and virulence.Gap Statement. While there is a pressing requirement for thorough detection of Kp pathotypes, current assays in resource-limited environments are unable to effectively focus on essential carbapenemase and hypervirulence genes with the necessary reliability and precision.Aim. To develop and validate a multiplex PCR (m-PCR) assay capable of simultaneously detecting Kp isolates including those carrying partial or full virulence markers, alongside antimicrobial resistance.Methodology. In this study, an m-PCR assay was designed and optimized for the simultaneous detection of key biomarkers associated with hypervirulent (rmpA, rmpA2, iucA, peg344 and iroB), carbapenem-resistant (bla NDM, bla OXA-48-like and bla KPC) and convergent Kp pathotypes in clinical isolates. The assay was evaluated on clinical isolates and validated against whole-genome sequencing (WGS) data for accuracy, specificity and sensitivity.Results. The developed m-PCR assay exhibited 100% specificity when compared to WGS data, successfully detecting all target genes without cross-amplification in ATCC control strains. The assay demonstrated high sensitivity, efficiently amplifying bacterial genomes from minimal DNA input as low as 1&#x2009;ng &#xb5;l-1. Additionally, validation through sequencing confirmed the accuracy of detected amplicons.Conclusion. This m-PCR assay offers a rapid, sensitive and specific diagnostic tool for differentiating Kp pathotypes in clinical settings, aiding in timely intervention and improved infection control measures.

Klebsiella pneumoniae↗