PubMed HealthSearch

PubMed · 40971804

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Abstract

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rok Kogoj, Mauro Petrillo, Samo Zakotnik, Alen Suljič, Miša Korva, Gabriele Leoni. 2025-10-02. Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.. https://doi.org/10.1093/bioinformatics%2Fbtaf516

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Epistasis and the changing fitness landscapes of SARS-CoV-2.

Since its emergence in late 2019, millions of SARS-CoV-2 genomes have been generated as part of global efforts to monitor the evolution and spread of the virus. This unprecedented volume of data provides a unique opportunity to study viral evolution at unparalleled resolution. In particular, individual genomic sites can be observed to have mutated independently thousands of times. These mutation counts have been used to estimate site-specific mutation rates and fitness effects for most mutations across the viral genome. Here, we use these data to investigate how the landscape of mutational fitness costs has changed over the course of the pandemic. SARS-CoV-2 evolution over the past 6 years has been characterized by the emergence of distinct variants separated by long branches corresponding to evolutionary saltations involving up to 50 mutations. We compare inferred fitness landscapes of the Spike protein across these variants and find that shifts in the estimated effects of non-synonymous mutations are linked to genetic differences between them. Sites with altered fitness costs are enriched near positions where the genetic backgrounds differ. To explain the observed changes, we introduce a model with pairwise epistatic interactions between mutations and residues that differ between variants. This model is able to explain about half of the variance in the shifts of fitness effects and suggests that each mismatch between variants substantially alters mutation effects at typically 1 to 3 additional positions.

SARS-CoV-2

Longitudinal characterization of mixed-genotype SARS-CoV-2 infections in a military cohort reveals compartmentalized viral populations.

UNLABELLED: Mixed-genotype severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections are a concern due to the potential generation of novel recombinants that give rise to new variants. To better understand intra-host viral dynamics, we analyzed specimens from 24 participants from the U.S. Military Health System's Epidemiology, Immunology, and Clinical Characteristics of Emerging Infectious Diseases with Pandemic Potential COVID-19 cohort with suspected mixed-genotype SARS-CoV-2 infections. From an initial 24 suspected cases, we confirmed 17 as genuine coinfections and graded them by evidence: 7 were "strong"; 4 were "moderate"; 6 were "weak"; and 7 were deemed unlikely to be true mixed-genotype infections. Access to swabs from multiple body sites across the course of infection allowed us to observe compartmentalization and shifts in variant dominance that would have been missed by a single-timepoint analysis, as well as one recombinant Omicron BA.1/BA.2 genome. By using an evidence-based bioinformatic framework to assess sequencing data from well-characterized clinical cases, we distinguished genuine coinfections from bioinformatic artifacts. Our findings emphasize the importance of both extensive specimen collection and careful bioinformatic approaches in ascertaining dual genotype infections. IMPORTANCE: Novel recombinants of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) arise from coinfections with different lineages, but mixed infections are not screened for despite risk to public health, and most surveillance relies on single swabs. We analyzed a longitudinal data set with specimens from multiple body sites, providing an opportunity to assess intra-host dynamics. To distinguish true coinfection from bioinformatic artifacts with confidence, we applied a framework that grades evidence for mixed genotypes by incorporating lineage and clade with manually validated variant calls. This allowed investigation beyond abundance levels of mixed genotypes within a single specimen, including observations of compartmentalization and a recombinant virus. This work enables further study of evolutionary, immunological, and clinical implications of mixed SARS-CoV-2 genotypes. Detecting dual-genotype infections and discriminating between true dual-genotype infection vs potential bioinformatics-based artifacts support public health and military readiness. These efforts provide evidence to bolster decision-making in molecular epidemiological studies to track transmission and for the choice of effective countermeasures.

SARS-CoV-2

Targeted ORF8-N Sanger Sequencing as a SARS-CoV-2 Surveillance Contingency During Supply Shortages.

BACKGROUND: Global shortages of next-generation sequencing (NGS) reagents threatened SARS-CoV-2 genomic surveillance in low- and middle-income countries during the COVID-19 pandemic. METHODS: During the 2021 NGS reagent shortages, we implemented targeted ORF8-N Sanger sequencing for SARS-CoV-2 variant surveillance in Brazilian public health laboratories. RESULTS: In silico analysis of whole-genome sequencing (WGS)-derived SARS-CoV-2 genomes from the Federal District, Brazil, showed that the ORF8-N target discriminated the major 2021 lineages (Gamma and Delta) and enabled analysis of ˃300 samples despite constrained NGS access. CONCLUSIONS: Targeted ORF8-N Sanger sequencing was a useful temporary contingency during NGS reagent shortages but offered lower phylogenetic resolution than WGS.

SARS-CoV-2