PubMed HealthSearch

SEARCH · PubMed Health

Results for “DNA Contamination”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A rare and atypical case of long-distance indirect DNA transfer: Contamination from an investigator never present at the scene.

To maximize the usefulness of DNA obtained from biological samples in forensic genetics, it is crucial to avoid DNA contamination throughout all procedures, from sample collection at crime scenes to STR profile generation in DNA laboratories. This study reports a rare and atypical case of DNA contamination in a forensic setting. During the analysis of biological evidence from a cold case preserved for 18 years, the STR profile obtained from the surface of a plastic bag matched that of an investigator, identified through the DNA elimination database. Case reconstruction confirmed that the investigator-who was located 80 km from the DNA laboratory and had never entered the crime scene or the sample storage room-was not a suspect and that the obtained STR profile originated from contamination. The most plausible explanation for the contamination was indirect transfer: investigator's DNA had adhered to a colleague's clothing and was subsequently dislodged and deposited onto the surface of the plastic bag as the colleague approached the sample pretreatment area. This study integrates trace DNA profiling of challenged samples with rapid contamination investigation and proposes prevention and control measures. This case underscores that, although DNA is widely regarded as the "gold standard" in forensic genetics, its interpretation must be considered within the context of the entire case. Conclusions should not be drawn based solely on a single DNA result.

Humans

DNA sequencing for microbial surveillance in cystic fibrosis airways: advances, challenges, and clinical translation.

SUMMARYDNA sequencing has revolutionized microbial surveillance in cystic fibrosis (CF), transforming pathogen identification from culture-dependent to total microbial community identification using molecular-based approaches. Techniques such as 16S rRNA gene sequencing have uncovered the complexity of the CF airway microbiome, while shotgun metagenomics, metatranscriptomics, and viromics now provide strain-level, functional, and viral insights beyond bacterial identification. Despite these advances, key technical and logistical challenges remain, including the processing of high-viscosity sputum samples, overwhelming host DNA contamination, managing large data sets, and the integration of complex bioinformatic outputs into clinical workflows. Emerging innovations such as host DNA depletion protocols, targeted enrichment panels, and adaptive sampling on Oxford Nanopore platforms are helping to overcome these barriers, improving microbial recovery and sequencing efficiency. As cystic fibrosis transmembrane conductance regulator (CFTR) modulator therapies are changing the lives of people with cystic fibrosis (pwCF), sequencing offers an unprecedented opportunity to track potential microbial adaptation in response. This review investigates current advances, limitations, and translational opportunities in DNA sequencing for CF airway microbiome surveillance, highlighting how these technologies can help reshape research and clinical microbiology in the post-modulator era.

Cystic Fibrosis

Short-read genome skimming enables molecular barcoding of old myxomycete collections.

This study evaluates the effectiveness of Illumina-based genome skimming for barcoding myxomycete herbarium collections ranging from 29 to 91 years in age. We successfully retrieved partial sequences of the standard marker gene (nucSSU) in all cases, as well as additional markers (mtSSU, EF1a, and COI) for certain collections. Altogether, 28 genes were recognized in the studied material. In a 33-year-old specimen of Lindbladia tubulina, the assembly reached an N50 of 4.19 kb, enabling the recovery of extended functional loci. The input genomic DNA quantity emerges as the primary determinant of sequencing success. Samples with high DNA yields provide representative amounts of contigs coming confirmedly (matching sequences in the NCBI nucleotide database) or potentially (no-hit fraction) from myxomycetes, regardless of specimen age. In addition to target DNA, we revealed distinct signals of both anthropogenic contamination (human DNA and skin microflora) and natural substrate inhabitants, including oribatid mites and bacteria from dead wood, soil, and grass litter. Thus, even in old collections, metagenomic data still carry information regarding the substrate upon which the myxomycete developed. The results demonstrate that short-read genome skimming may help to integrate historical type material of myxomycetes into contemporary phylogenetic research. This method overcomes the length-dependent limitations of traditional Sanger sequencing, thus providing a roadmap for the future of museomics in myxomycetology.

Amoebozoa

Intra-amniotic infection: diagnosis, nomenclature, clinical significance, management, and microbiologic tools used for the diagnosis.

SUMMARYIntra-amniotic infection is the main cause of spontaneous preterm birth and adverse maternal-fetal outcomes; therefore, rapid, robust, and accurate diagnosis remains a clinical priority. Conventional microbiological techniques, especially culture-based methods, are limited by long turnaround times and the inability to detect fastidious or unculturable organisms. This review summarizes the diagnosis, nomenclature, clinical significance, management, and laboratory approaches for diagnosing intra-amniotic infection. Targeted nucleic acid amplification methods, including species-specific polymerase chain reaction and broad-range 16S rRNA gene sequencing, have improved the detection of bacterial DNA and enabled the identification of organisms that evade routine culture in intra-amniotic infection. More recently, whole-genome sequencing and metagenomic next-generation sequencing have provided culture-independent strategies for comprehensive pathogen profiling, allowing simultaneous detection of bacteria, viruses, and fungi, as well as characterization of antimicrobial resistance determinants and virulence-associated genes. However, challenges remain, particularly in low-biomass samples such as amniotic fluid, where contamination, host DNA background, and data interpretation can compromise specificity. This review critically evaluates the advantages and limitations of each molecular modality and discusses pre-analytical, analytical, and bioinformatic considerations essential for reliable implementation. Integration of molecular diagnostics into clinical workflows holds promise for improving etiological diagnosis and guiding targeted therapy in intra-amniotic infection, thereby improving maternal and fetal outcomes.

Humans

Towards a Robust cell-free DNA Isolation Protocol for NGS Applications in a Clinical Molecular Diagnostics Setting.

Cell-free DNA (cfDNA), released from apoptotic and necrotic cells into body fluids, is a non-invasive source of genetic information for disease prediction, diagnosis, and monitoring. However, its low abundance makes cfDNA highly susceptible to various pre-analytical influences, potentially increasing high molecular weight (HMW) or genomic DNA (gDNA) compromising downstream cfDNA analyses. This study evaluated the impact of different cfDNA-stabilizing blood collection tubes (BCT; Cell-Free DNA BCT, Streck; S-Monovette cfDNA Exact, Sarstedt) stored at room temperature for 1, 5, or 10 days, prior to plasma isolation using different isolation methods (magnetic bead-based or silica column-based) on cfDNA stability and yield. DNA quantity and quality were assessed by fluorometric quantification, automated fragment analysis, and gene-specific quantitative PCR. Streck-based workflows maintained stable cfDNA yields and characteristic mononucleosomal fragmentation profiles across all storage times. In contrast, Sarstedt tubes showed reduced cfDNA concentrations after 5 days and a pronounced increase at 10 Days, accompanied by high-molecular weight DNA patterns consistent with white-blood cells (WBC) lysis. These trends were largely independent of the extraction method. Overall, the results demonstrate that blood collection tube chemistry critically influences cfDNA integrity during delayed processing. Streck tubes, particularly when combined with silica column-based isolation method, provided the most robust and reproducible workflow for routine molecular diagnostics, whereas Sarstedt tubes produced physiologically implausible results after extended storage.

blood collection tubes

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

ViralQC: a tool for assessing completeness and contamination of predicted viral contigs.

MOTIVATION: Viruses represent the most abundant biological entities on Earth, playing vital roles in diverse ecosystems. Cataloging viruses across various environments is essential for understanding their properties and functions. Metagenomic sequencing has emerged as the most comprehensive method for virus discovery. However, distinguishing viral sequences from the vast background of microbial organisms in metagenomic data remains a significant challenge. Existing tools experience varying degrees of false positive rates due to noise in sequencing and assembly, and the integration of proviruses into microbial genomes. This highlights the urgent need for an accurate and efficient method to evaluate the quality of viral contigs. RESULTS: To address these challenges, we introduce ViralQC, a tool designed to assess the quality of viral contigs or bins. ViralQC identifies microbial contamination within putative viral sequences using an ensemble framework powered by DNA and protein foundation models and estimates completeness by analyzing protein organization. We evaluated ViralQC on multiple datasets and compared its performance against the state-of-the-art tool, CheckV. Leveraging both DNA and protein foundation models, ViralQC achieves higher sensitivity on contamination detection for contigs longer than 10 kbp while maintaining comparable accuracy. Additionally, ViralQC delivers more accurate estimation on contigs with completeness > 50%. AVAILABILITY: The source code of ViralQC is available via: https://github.com/ChengPENG-wolf/ViralQC.

Software

Cleanifier: contamination removal from microbial sequences using spaced seeds of a human pangenome index.

MOTIVATION: The first step when working with DNA data of human-derived microbiomes is to remove human contamination for two reasons. First, many countries have strict privacy and data protection guidelines for human sequence data, so microbiome data containing partly human data cannot be easily further processed or published. Second, human contamination may cause problems in downstream analysis, such as metagenomic binning or genome assembly. For large-scale metagenomics projects, fast and accurate removal of human contamination is therefore critical. RESULTS: We introduce Cleanifier, a fast and memory frugal alignment-free tool for detecting and removing human contamination based on gapped k-mers, or spaced seeds. Cleanifier uses a pangenome index of known human gapped k-mers, and the creation and use of alternative references is also possible. Reads are classified and filtered according to their gapped k-mer content. Cleanifier supports two filtering modes: one that queries all gapped k-mers and one that queries only a sample of them. A comparison of Cleanifier with other state-of-the-art tools shows that the sampling mode makes Cleanifier the fastest method with comparable accuracy. When using a probabilistic Cuckoo filter to store the complete k-mer set, Cleanifier has similar memory requirements to methods that use a sampled minimizer index. At the same time, Cleanifier is more flexible, because it can use different sampling methods on the same index. AVAILABILITY AND IMPLEMENTATION: Cleanifier is available via gitlab (https://gitlab.com/rahmannlab/cleanifier), PyPi (https://pypi.org/project/cleanifier/), and Bioconda (https://anaconda.org/bioconda/cleanifier). The pre-computed human pangenome index is available at Zenodo (https://doi.org/10.5281/zenodo.15639519).

Humans

AdDeam: a fast and scalable tool for estimating and clustering reference-level damage profiles.

MOTIVATION: DNA damage patterns, such as increased frequencies of C→T and G→A substitutions at fragment ends, are widely used in ancient DNA studies to assess authenticity and detect contamination. In metagenomic studies, fragments can be mapped against multiple references or de novo assembled contigs to identify those likely to be ancient. Generating and comparing damage profiles, however, can be both tedious and time-consuming. Although tools exist for estimating damage in single reference genomes and metagenomic datasets, none efficiently cluster damage patterns. RESULTS: To address this methodological gap, we developed AdDeam, a tool that combines rapid damage estimation with clustering for streamlined analyses and easy identification of potential contaminants or outliers. Our tool takes aligned ancient DNA (aDNA) fragments from various samples or contigs as input, computes damage patterns, clusters them, and outputs representative damage profiles per cluster, a probability of each sample pertaining to a cluster, as well as a Principal Component Analysis of the damage patterns for each sample for fast visualisation. We evaluated AdDeam on both simulated and empirical datasets. AdDeam effectively distinguishes different damage levels, such as uracil-DNA glycosylase-treated samples, sample-specific damages from specimens of different time periods, and can also distinguish between contigs containing modern or ancient fragments, providing a clear framework for aDNA authentication and facilitating large-scale analyses. AVAILABILITY AND IMPLEMENTATION: AdDeam is publicly available at https://github.com/LouisPwr/AdDeam and can also be installed via Bioconda. It is implemented in Python and C++. All analysis scripts and datasets are available at https://github.com/LouisPwr/AdDeamAnalysis and on Zenodo under: 10.5281/zenodo.15052427.

Software

Triphenyl Phosphate Alters Methyltransferase Expression and Induces Genome-Wide Aberrant DNA Methylation in Zebrafish Larvae.

Emerging environmental contaminants, organophosphate flame retardants (OPFRs), pose significant threats to ecosystems and human health. Despite numerous studies reporting the toxic effects of OPFRs, research on their epigenetic alterations remains limited. In this study, we investigated the effects of exposure to 2-ethylhexyl diphenyl phosphate (EHDPP), tricresyl phosphate (TMPP), and triphenyl phosphate (TPHP) on DNA methylation patterns during zebrafish embryonic development. We assessed general toxicity and morphological changes, measured global DNA methylation and hydroxymethylation levels, and evaluated DNA methyltransferase (DNMT) enzyme activity, as well as mRNA expression of DNMTs and ten-eleven translocation (TET) methylcytosine dioxygenase genes. Additionally, we analyzed genome-wide methylation patterns in zebrafish larvae using reduced-representation bisulfite sequencing. Our morphological assessment revealed no general toxicity, but a statistically significant yet subtle decrease in body length following exposure to TMPP and EHDPP, along with a reduction in head height after TPHP exposure, was observed. Eye diameter and head width were unaffected by any of the OPFRs. There were no significant changes in global DNA methylation levels in any exposure group, and TMPP showed no clear effect on DNMT expression. However, EHDPP significantly decreased only DNMT1 expression, while TPHP exposure reduced the expression of several DNMT orthologues and TETs in zebrafish larvae, leading to genome-wide aberrant DNA methylation. Differential methylation occurred primarily in introns (43%) and intergenic regions (37%), with 9% and 10% occurring in exons and promoter regions, respectively. Pathway enrichment analysis of differentially methylated region-associated genes indicated that TPHP exposure enhanced several biological and molecular functions corresponding to metabolism and neurological development. KEGG enrichment analysis further revealed TPHP-mediated potential effects on several signaling pathways including TGFβ, cytokine, and insulin signaling. This study identifies specific changes in DNA methylation in zebrafish larvae after TPHP exposure and brings novel insights into the epigenetic mode of action of TPHP.

Animals

WHOLE GENOME TARGETED ENRICHMENT AND SEQUENCING OF HUMAN-INFECTING CRYPTOSPORIDIUM spp.

Cryptosporidium spp. are protozoan parasites that cause severe illness in vulnerable human populations. Obtaining pure Cryptosporidium DNA from clinical and environmental samples is challenging because the oocysts shed in contaminated feces are limited in quantity, difficult to purify efficiently, may derive from multiple species, and yield limited DNA (<40 fg/oocyst). Here, we develop and validate a set of 100,000 RNA baits (CryptoCap_100k) based on six human-infecting Cryptosporidium spp. (C. cuniculus, C. hominis, C. meleagridis, C. parvum, C. tyzzeri, and C. viatorum) to enrich Cryptosporidium spp. DNA from a wide array of samples. We demonstrate that CryptoCap_100k increases the percentage of reads mapping to target Cryptosporidium references in a wide variety of scenarios, increasing the depth and breadth of genome coverage, facilitating increased accuracy of detecting and analyzing species within a given sample, while simultaneously decreasing costs, thereby opening new opportunities to understand the complex biology of these important pathogens.

Journal Article

Metax enables accurate cross-domain taxonomic profiling of metagenomes.

Taxonomic profiling is fundamental to microbiome research, yet achieving high species-level accuracy remains challenging for complex communities that span bacteria, viruses, eukaryotes, and archaea, and these limitations are exacerbated in low-biomass, host-dominated samples. We introduce Metax, a cross-domain taxonomic profiler that integrates coverage-based probabilistic modeling with an expectation-maximization framework to distinguish true microbial signals from artifacts. Across >600 samples from host-associated, environmental, wastewater, and low-biomass clinical settings, including benchmarks with limited reference representation, Metax improved profiling accuracy, achieving on average 55% higher F1 scores and 45% lower Bray-Curtis dissimilarity than other methods. Moreover, this broad evaluation demonstrated that Metax resolved bacterial and viral signatures of peri-implantitis in oral microbiomes and revealed signals suggestive of reagent-borne contaminants and reference misassemblies in plasma-cell-free DNA. By leveraging genome-wide coverage evidence, Metax enables robust cross-domain profiling across diverse sample types and sequencing depths, including settings where reference databases are highly incomplete.

abundance estimation

Rapidly decellularized adipose tissue induces soft tissue vascularization in potential anatomical spaces.

Decellularized tissues provide biological cues owing to the wealth of structural and regulatory factors that promote angiogenesis, adipogenesis, and myogenesis and facilitate neurite outgrowth. Here, we demonstrated the advantages of decellularized adipose tissue (adipoECM) over defined collagen-based biomaterials for host tissue integration. Three batches of human adipose tissue were decellularized using a rapid decellularization protocol and analyzed using mass spectrometry. To assess the biological activity of the decellularized materials, adipoECM and a reference standard of care biomaterial (Integra&#xae;DRT, also containing collagen I and glycosaminoglycans) were implanted subcutaneously, but far from the wound bed (in anatomical potential spaces) of immunocompetent BALB/c mice. The mice were euthanized in the acute (1 day) and chronic (day 60) inflammatory reaction phases, followed by biomaterial excision and Masson&#x2019;s trichrome immunohistofluorescence imaging of the paraffin-embedded specimens. Each batch of processed tissue passed a quality control check, showing a low level of donor genomic DNA, lack of nuclei, lipids, endotoxins, and bacterial contamination. Mass spectrometry revealed that all batches of decellularized tissue mainly contained collagen I and, to a lesser degree, collagen III, collagen IV, collagen V, laminin, fibrillin, fibronectin, tenascin, and elastin. No acute inflammatory reaction was observed in either material one day post-transplantation. At 60&#x2009;days post-implantation, different cell types were detected in adipoECM specimens, whereas Integra&#xae;DRT remained acellular. Additional immunohistochemical staining of adipoECM revealed CD31-positive cells in the blood vessels. Mesenchymal (CD90 positive) and myeloid (CD14 positive) cells were also detected. Primary cell types involved in soft tissue healing and remodeling were found in the adipoECM-treated group. The ingrowth of blood vessels and mesenchymal cells confirmed the effective integration of adipoECM with host tissues. Our results demonstrate that decellularized adipose tissue implanted away from the wound bed possesses contextual biological activities that promote efficient integration with host tissues.

Adipose Tissue

Comparative performance of sponge versus flocked swabs for culture-based and metagenomic detection of microbial contamination in the healthcare environment.

BACKGROUND: Identifying optimal methods for sampling surfaces in the healthcare environment is critical for future research requiring the identification of multidrug-resistant organisms (MDROs) on surfaces. METHODS: We compared 2 swabbing methods, use of a flocked swab versus a sponge-stick, for recovery of MDROs by both culture and recovery of bacterial DNA via quantitative 16S polymerase chain reaction (PCR). This comparison was conducted by assessing swab performance in a longitudinal survey of MDRO contamination in hospital rooms. Additionally, a laboratory-prepared surface was also used to compare the recovery of each swab type with a matching surface area. RESULTS: Sponge-sticks were superior to flocked swabs for culture-based recovery of MDROs, with a sensitivity of 80% compared to 58%. Similarly, sponge-sticks demonstrated greater recovery of Staphylococcus aureus from laboratory-prepared surfaces, although the performance of flocked swabs improved when premoistened. In contrast, recovery of bacterial DNA via quantitative 16S PCR was greater with flocked swabs by an average of 3 log copies per specimen. CONCLUSIONS: The optimal swabbing method of environmental surfaces differs by method of analysis. Sponge-sticks were superior to flocked swabs for culture-based detection of bacteria but inferior for recovery of bacterial DNA.

Humans

Mitochondrial Impostors: Prevalence and Impacts of NUMTs on Genetic and Evolutionary Studies in Carnivora.

Nuclear mitochondrial pseudogenes are mitochondria-derived DNA sequences integrated into the nuclear genome, which can introduce errors in species identification, phylogenetic inference, and population genetics. Although nuclear mitochondrial pseudogene contamination has been reported in some Carnivora species, a systematic investigation into the prevalence and impacts of nuclear mitochondrial pseudogenes across an order is still lacking. In this study, 22,102 mitochondrial DNA sequences of 80 Carnivora species from 14 families and 54 genera were retrieved from the public National Center for Biotechnology Information database and further analyzed. Using alignment-based methods, 158 problematic sequences/sequence groups were identified and categorized into four types: nuclear mitochondrial pseudogenes, species misidentification or mislabeling, sequence errors, and anomalous sites. Among families, Felidae exhibited the highest rate of nuclear mitochondrial pseudogene contamination, particularly in species of the genus Panthera. In contrast, no nuclear mitochondrial pseudogene contamination was detected in members of Ursidae and Ailuridae. Phylogenetic analysis revealed multiple independent origins of nuclear mitochondrial pseudogene, with some tracing back to the common ancestor of Carnivora. To mitigate nuclear mitochondrial pseudogene-related errors, rigorous sequence verification strategies, such as sequence alignment and phylogenetic validation, should be implemented. In conclusion, our findings highlight the necessity of nuclear mitochondrial pseudogene awareness in genetic and evolutionary studies of Carnivora and other taxa.

Animals

Cold-adapted RNA polymerase from Pseudomonas phage Njord improves synthesis of therapeutic mRNA.

An RNA polymerase identified in the genome of Pseudomonas phage Njord offers a promising tool for the synthesis of mRNA and other therapeutic nucleic acids. Originating from a marine microbial ecosystem, Njord RNAP transcribes RNA at high yield even under low temperature conditions. Key properties of the enzyme relevant to mRNA synthesis are presented including transcriptional fidelity, promoter specificity, incorporation of modified nucleotides, and the impurity profile of the RNA. Specific attention is given to the formation of contaminating double-stranded RNA (dsRNA) species. Analysis of transcription reactions shows that DNA-templated promoter-independent transcription is a major source of detectable dsRNA impurities and that Njord RNAP displays a minimal level of this activity. Consistent with the known inflammatory role of dsRNA in synthetic mRNA, transcriptomic analysis of cell culture and a live animal study demonstrates that mRNA synthesized with Njord RNAP elicits only a minimal immune response. This natural enzyme enables efficient mRNA synthesis at ambient temperature and produces transcripts essentially free of dsRNA, offering significant potential to streamline mRNA manufacturing processes.

DNA-Directed RNA Polymerases

The diagnostic potential of combined quantitative polymerase chain reaction and next-generation sequencing using the same primers for periprosthetic joint infection.

Next-generation sequencing (NGS) enables the detection of specific pathogens unidentifiable by conventional cultures, but its application in orthopedics remains inconsistent due to background contamination and irreproducible findings. This study evaluated the diagnostic performance of a novel workflow combining broad-range 16S rRNA gene quantitative PCR (qPCR) screening with downstream NGS, focusing on bacterial biomass thresholds. The qPCR assay demonstrated excellent intrarater reliability, with an intraclass correlation coefficient (ICC) of 0.961 (95% confidence interval, 0.881 to 0.997). Based on serially diluted positive controls, a quantitative threshold of 10&#x2075; CFU/mL was established as the minimum concentration required for the consistent detection of fastidious taxa, such as Escherichia coli. When evaluated against conventional cultures using 95 sonicate fluid and 276 pre/intraoperative tissue samples, the qPCR assay achieved a sensitivity of 80% and a specificity of 72%. Subsequent NGS sequencing of 26 clinical samples and 9 controls showed concordance in 4 of 6 culture-positive infected cases with NGS taxonomy, whereas the remaining discrepancies were likely attributable to culture-based phenotypic misidentification. Notably, among the qPCR-positive cases, three were culture-negative, including two hip prosthesis loosening cases exhibiting polymicrobial profiles, and one post-traumatic osteoarthritis case harboring low-level Staphylococcus. Crucially, this post-traumatic patient developed delayed periprosthetic joint infection (PJI) 2 years post-surgery, with cultures identifying Staphylococcus previously detected by the initial NGS analysis. Integrating qPCR screening with targeted NGS effectively refines pathogen identification, filters environmental artifacts, and overcomes the diagnostic limitations of culture-negative infections in orthopedic practice.IMPORTANCENext-generation sequencing (NGS) enables the detection of specific pathogens in clinical samples that are not identifiable by conventional methods. However, NGS applications in orthopedics have not been quantitatively evaluated, and findings have been inconsistent owing to contaminants and the presence of non-credible causative organisms. These factors primarily stem from the failure to evaluate low-biomass samples and the absence of proper controls, such as negative controls or mock community DNA samples. This study demonstrates that interpreting results from low-biomass samples requires careful consideration because NGS relies on relative bacterial abundances; distinguishing likely pathogens from contaminants is particularly challenging when bacterial loads are low. We demonstrated that combining NGS with quantitative PCR (qPCR) and applying a Cq cutoff can reduce false positives.

Humans

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article