PubMed Health⌕ Search

Biomedical subjects

Matthew Bashton

Publications and source records attributed to Matthew Bashton.

5 recordsLinked to original sources

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans↗

Effectiveness of rapid SARS-CoV-2 genome sequencing in supporting infection control for hospital-onset COVID-19 infection: Multicentre, prospective study.

BACKGROUND: Viral sequencing of SARS-CoV-2 has been used for outbreak investigation, but there is limited evidence supporting routine use for infection prevention and control (IPC) within hospital settings. METHODS: We conducted a prospective non-randomised trial of sequencing at 14 acute UK hospital trusts. Sites each had a 4-week baseline data collection period, followed by intervention periods comprising 8 weeks of 'rapid' (<48 hr) and 4 weeks of 'longer-turnaround' (5-10 days) sequencing using a sequence reporting tool (SRT). Data were collected on all hospital-onset COVID-19 infections (HOCIs; detected &#x2265;48 hr from admission). The impact of the sequencing intervention on IPC knowledge and actions, and on the incidence of probable/definite hospital-acquired infections (HAIs), was evaluated. RESULTS: A total of 2170 HOCI cases were recorded from October 2020 to April 2021, corresponding to a period of extreme strain on the health service, with sequence reports returned for 650/1320 (49.2%) during intervention phases. We did not detect a statistically significant change in weekly incidence of HAIs in longer-turnaround (incidence rate ratio 1.60, 95% CI 0.85-3.01; p=0.14) or rapid (0.85, 0.48-1.50; p=0.54) intervention phases compared to baseline phase. However, IPC practice was changed in 7.8 and 7.4% of all HOCI cases in rapid and longer-turnaround phases, respectively, and 17.2 and 11.6% of cases where the report was returned. In a 'per-protocol' sensitivity analysis, there was an impact on IPC actions in 20.7% of HOCI cases when the SRT report was returned within 5 days. Capacity to respond effectively to insights from sequencing was breached in most sites by the volume of cases and limited resources. CONCLUSIONS: While we did not demonstrate a direct impact of sequencing on the incidence of nosocomial transmission, our results suggest that sequencing can inform IPC response to HOCIs, particularly when returned within 5 days. FUNDING: COG-UK is supported by funding from the Medical Research Council (MRC) part of UK Research & Innovation (UKRI), the National Institute of Health Research (NIHR) (grant code: MC_PC_19027), and Genome Research Limited, operating as the Wellcome Sanger Institute. CLINICAL TRIAL NUMBER: NCT04405934.

Humans↗

Supra-domains: evolutionary units larger than single protein domains.

Domains are the evolutionary units that comprise proteins, and most proteins are built from more than one domain. Domains can be shuffled by recombination to create proteins with new arrangements of domains. Using structural domain assignments, we examined the combinations of domains in the proteins of 131 completely sequenced organisms. We found two-domain and three-domain combinations that recur in different protein contexts with different partner domains. The domains within these combinations have a particular functional and spatial relationship. These units are larger than individual domains and we term them "supra-domains". Amongst the supra-domains, we identified some 1400 (1203 two-domain and 166 three-domain) combinations that are statistically significantly over-represented relative to the occurrence and versatility of the individual component domains. Over one-third of all structurally assigned multi-domain proteins contain these over-represented supra-domains. This means that investigation of the structural and functional relationships of the domains forming these popular combinations would be particularly useful for an understanding of multi-domain protein function and evolution as well as for genome annotation. These and other supra-domains were analysed for their versatility, duplication, their distribution across the three kingdoms of life and their functional classes. By examining the three-dimensional structures of several examples of supra-domains in different biological processes, we identify two basic types of spatial relationships between the component domains: the combined function of the two domains is such that either the geometry of the two domains is crucial and there is a tight constraint on the interface, or the precise orientation of the domains is less important and they are spatially separate. Frequently, the role of the supra-domain becomes clear only once the three-dimensional structure is known. Since this is the case for only a quarter of the supra-domains, we provide a list of the most important unknown supra-domains as potential targets for structural genomics projects.

Animals↗

Structure, function and evolution of multidomain proteins.

Proteins are composed of evolutionary units called domains; the majority of proteins consist of at least two domains. These domains and nature of their interactions determine the function of the protein. The roles that combinations of domains play in the formation of the protein repertoire have been found by analysis of domain assignments to genome sequences. Additional findings on the geometry of domains have been gained from examination of three-dimensional protein structures. Future work will require a domain-centric functional classification scheme and efforts to determine structures of domain combinations.

Computer Simulation↗

The geometry of domain combination in proteins.

Most proteins in genomes are the result of the recombination of two or more domains. It has been found that if proteins are formed by a combination of domains from superfamilies A and B, then the domains may occur in the sequential order AB or BA but only in about 2% of cases do they occur in both sequential orders. The classical Rossmann domains of known structure are combined with catalytic domains from seven different superfamilies. In addition, there are eight cases where structures with both AB and BA domain combinations are known. For these two sets of structures, we analysed: (i) the relative orientation of the domains; (ii) the type of domain connection; (iii) the structure of the interdomain links; and (iv) domain function. The results of this analysis indicate that in most cases domain order is conserved because recombination of the domains has only occurred once during the course of evolution. Functional reasons become important when the domain connections are short. In seven out of the eight known cases where domains are combined in the AB and BA sequential orders they have different geometrical relationships that give them different functional properties.

Animals↗