PubMed HealthSearch

SEARCH · PubMed Health

Results for “strain-level classification”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2 recordsLinked to original sources

Using Mapping-Profiles to Refine Strain-Level Metagenomic Classification.

Metagenomic classification at the strain level remains challenging due to high sequence similarity among closely related genomes, which leads to ambiguous read mappings and frequent false-positive strain detections. Reducing such errors improves the reliability of strain-level analyses, which is critical for applications such as pathogen detection. We introduce StrainRefine, a post-mapping refinement method that analyzes read-reference mapping profiles to resolve ambiguous assignments among highly similar genomes. The method represents candidate reference genomes using binary profiles that capture read-support patterns and measures similarity between references based on profile overlap. The method clusters references based on similar mapping profiles, filters weakly supported genomes, and reassigns reads to representative references, reducing redundant reporting of near-identical strains. StrainRefine substantially reduces false-positive strain detections while preserving recall and improving agreement between predicted and true abundance profiles. On large-scale metagenomic datasets, it achieves a substantially improved precision-recall balance compared with existing mapping-based approaches, with the standalone method obtaining the highest read-level classification accuracy on the most complex evaluated dataset. Unlike many strain-level tools designed for individual species, StrainRefine operates without prior assumptions about sample composition or curated species-specific reference collections, while still achieving comparable performance in single-species settings on species-specific reference databases. These results highlight mapping-profile similarity as an effective signal for improving strain-level metagenomic classification.

false-positive reduction

Whole-genome surveillance supports hazard profiling of Escherichia coli lineages in recycled water treatment systems.

UNLABELLED: The use of treated wastewater is increasingly important for sustainable water management under a changing climate, yet conventional monitoring based on Escherichia coli enumeration provides limited insight into strain diversity and associated public health hazards. Here, we applied longitudinal whole-genome sequencing (WGS) to 180 E. coli isolates collected across the treatment continuum of a recycled water facility, from influent to final effluent. Genomic analysis revealed extensive strain-level heterogeneity, comprising 88 sequence types across eight phylogroups, with greater diversity in influent than in treated effluent. Phylogenetic comparisons with contextual Australian genomes indicated clustering with strains associated with companion animals, wild birds, humans, and livestock, suggesting multiple potential source reservoirs rather than a single dominant origin, although source contributions were not definitive. Despite a >90% reduction in total E. coli loads, isolates recovered from upstream and downstream stages exhibited broadly comparable virulence factor and antimicrobial resistance gene (ARG) profiles, suggesting that, within the cultured isolate collection, reductions in abundance exceeded shifts in genomic composition. To assess operational relevance, we prototyped a genomics-informed hazard framework integrating virulence determinants, ARGs, plasmid-associated mobility, and reuse-specific exposure context. Using this framework, 92.8% of isolates were classified as low hazard, and 7.2% as moderate hazard, with no isolates meeting criteria for high or critical hazard classifications. These findings demonstrate that genomic profiling of indicator organisms can reveal population structure and hazard heterogeneity not captured by conventional enumeration alone, and can provide a practical basis for incorporating genomic information into hazard-informed monitoring of recycled water systems. IMPORTANCE: Routine recycled water monitoring relies largely on culture-based E. coli counts, which indicate regulatory compliance but provide limited insight into strain diversity, persistence, and genomic characteristics relevant to public health. Using longitudinal whole-genome sequencing, we show that genetically distinct E. coli lineages, including isolates carrying combinations of virulence and antimicrobial resistance determinants, can persist through advanced treatment despite substantial reductions in overall E. coli loads. While most isolates were classified as low genomic hazard and no high- or critical-hazard isolates were detected, these findings demonstrate that conventional enumeration alone cannot distinguish between genetically diverse lineages with differing hazard potential in highly treated systems. By integrating genomic data into a hazard classification framework, this study demonstrates an applied approach to contextualize E. coli detections and distinguish low-risk background populations from isolates with elevated genomic hazard profiles. This work supports the use of genomic profiling of indicator organisms to improve surveillance, inform treatment performance assessment, and enable more risk-based management of recycled water systems.

Escherichia coli