PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome quality”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A high-quality genomic catalog of the human oral microbiome broadens its phylogeny and clinical insights.

The oral microbiome is increasingly linked to human health. To further examine this microbial community, we present the human reference oral microbiome (HROM), with 72,641 high-quality genomes from 3,426 species, including 2,019 previously unidentified species, improving metagenomic sequence read classification over existing catalogs. Notably, HROM unveils 1,137 previously uncharacterized candidate phyla radiation (CPR) species, establishing Patescibacteria as the most prevalent phylum in the oral microbiota and distinct from environmental Patescibacteria. Additionally, an oral CPR subclade is associated with periodontitis, complementing Porphyromonas gingivalis in predicting disease. Finally, comparing HROM with reference genomes of the gut microbiome reveals taxonomic and functional divergence between these microbiomes. HROM contains 42 ectopic oral species, and their relative abundance in gut microbiota is predictive of intestinal, cardiovascular, and liver diseases. Thus, HROM offers an expanded view of the oral microbiome and highlights the clinical importance of further examining the links between oral microbes and systemic disorders.

Humans

Cumulus cells enhance oocyte genomic quality control by promoting DNA damage-induced meiotic arrest.

Cumulus cells are known to maintain oocyte arrest at prophase I through gap junction-mediated cAMP signalling, but their role after meiotic resumption remains unclear. Here, we show that cumulus cells enhance oocyte genomic quality control by sensitizing mouse oocytes to DNA damage-induced meiotic arrest. Time-lapse imaging of SiR-tubulin-labelled spindles revealed that oocytes from cumulus-oocyte complexes (COCs) matured faster than denuded oocytes (DOs). Upon mild DNA damage induced by low-dose etoposide, COC oocytes arrested at metaphase I, whereas DOs completed maturation despite similar levels of DNA lesions. This arrest required spindle assembly checkpoint (SAC) activity, as reversine rescued polar body extrusion and BubR1 and Mad2 were elevated in COCs but not DOs. Disruption of gap junctions or inhibition of mTOR signalling abolished the checkpoint response. Notably, cumulus cells did not enhance oocyte response to minor spindle perturbations. These findings reveal a previously unrecognized role of cumulus cells in mediating DNA damage-induced SAC activation, providing post-GVBD genomic surveillance beyond prophase I arrest.

Animals

High-Quality Genome Assembly, Metabolome, Pangenome, and Metabolic Models of Megasphaera hexanoica KCCM 43214T.

Megasphaera hexanoica KCCM 43214T, isolated from cow rumen, is capable of producing medium-chain carboxylic acids such as hexanoate and octanoate. In this study, we present a high-quality genome assembly, along with intracellular metabolomic profiling and pangenomic analysis. Illumina sequencing generated 2.3 Gbp from 15,293,634 reads with a GC content of 49.5%, while PacBio HiFi sequencing produced 331.5 Mbp across 45,266 reads, with an average read length of 7,323 bp and a HiFi read N50 of 8,214 bp. Hybrid assembly of short and long reads resulted in a single 2.88 Mbp contig, containing 2,835 protein-coding genes. Genome-scale metabolic models were constructed to evaluate its metabolic capabilities under specific growth conditions. Intracellular metabolomic analysis of cells grown in medium containing fructose and lactate revealed key metabolic activities associated with chain elongation. Pangenomic analysis across nine annotated genomes identified 6,721 orthologous genes using OrthoMCL, emphasizing the genetic and functional diversity within the Megasphaera genus. This dataset offers valuable insights into the metabolism and biotechnological potential of M. hexanoica KCCM 43214T.

Metabolome

High quality genome assemblies of African cattle breeds using PacBio HiFi sequencing.

Africa has a uniquely rich cattle diversity of ~150 breeds comprising the Bos taurus indicus sub-species, Bos taurus taurus, and their crosses. These represent ~23% of the global cattle population. However, high quality, representative assemblies are limited for African cattle and especially for indicine breeds. Here we built high quality de novo assemblies for five important African indigenous cattle breeds using PacBio HiFi sequencing: Lagune (Bos taurus taurus), Gudali, Iringa Red and Singida White (Bos taurus indicus), and Mpwapwa (Bos taurus taurus x Bos taurus indicus). These new assemblies are the most contiguous and complete African cattle assemblies produced so far, with genome sizes of 3.25-3.36 Gb, contiguity N50s ranging from 83.59 Mb to 97.87 Mb and scaffold N50s from 100.30 Mb to 113.37 Mb. BUSCO genome completeness scores were also higher than 99.68%, indicative of highly contiguous assemblies. These improved and highly contiguous genome assemblies are consequently a valuable resource for future African and global livestock genomic studies.

Animals

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and ~ 2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia

Comparative evaluation of three high-molecular-weight DNA extraction kits for Oxford Nanopore sequencing of Clostridioides difficile and Clostridium perfringens.

UNLABELLED: Clostridioides difficile and Clostridium perfringens are Gram-positive, spore-forming anaerobic pathogens affecting humans and animals, for which genomic data have been mainly generated using short-read or hybrid sequencing approaches. In this study, we evaluated three commercial non-bead-beating DNA extraction kits designed for high-molecular-weight DNA recovery for Oxford Nanopore long-read whole-genome sequencing of two C. difficile and two C. perfringens strains, including one reference strain and one clinical or environmental isolate per species. Based on sequencing performance and kit ease of use, one kit was selected for additional sequencing of plasmid-carrying strains of both species. All three kits allowed correct identification of sequence types, toxin-encoding genes, and antimicrobial resistance determinants, confirming their suitability for clinical and epidemiological applications. However, the BT MasterPure Kit provided the highest DNA concentrations, longest fragment sizes, and superior read lengths and N50 values, particularly for C. difficile, achieving >100× coverage and enabling reliable circularization of chromosomes and plasmids, including a C. difficile metronidazole resistance plasmid and C. perfringens plasmids carrying toxin and antibiotic resistance genes. The other kits produced slightly lower DNA yields, resulting in shorter reads and reduced genome coverage for C. difficile, highlighting the challenge of extracting high-quality DNA from Gram-positive, spore-forming bacteria. Overall, this study provides practical guidance for selecting DNA extraction protocols optimized for Oxford Nanopore sequencing of C. difficile and C. perfringens, supporting high-quality genome assemblies and plasmid characterization and facilitating the routine genomic surveillance of clinically relevant spore-forming pathogens. IMPORTANCE: High-quality genomic data are essential for accurate characterization of Clostridioides difficile and Clostridium perfringens, two clinically and epidemiologically important Gram-positive, spore-forming pathogens. However, long-read sequencing performance can be strongly influenced by the choice of DNA extraction method, particularly for organisms with robust cell walls, where commonly used methods can lead to fragmented DNA. In this work, DNA of four strains was extracted using three commercial high-molecular-weight DNA extraction kits and sequenced using Oxford Nanopore Technologies. The best-performing kit was also evaluated using three additional strains known to harbor plasmids in order to assess its plasmid recovery efficiency. The results demonstrated successful plasmid recovery, circularization, and characterization. DNA extraction protocols optimized for Oxford Nanopore sequencing enable the rapid and cost-effective characterization of C. difficile and C. perfringens for genomic surveillance or outbreak investigations.

Clostridioides difficile

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Comparative genomics of ESKAPE pathogen species: Integrating pan-genome architecture, antimicrobial resistance, and virulence factor repertoires.

BACKGROUND: ESKAPE pathogens are major causes of hospital-acquired infections and are characterized by extensive antimicrobial resistance (AMR) and diverse virulence mechanisms. Although species-specific pan-genome studies have revealed substantial genomic diversity, the relationships among genome plasticity, resistance burden, and virulence remain incompletely understood across the ESKAPE complex. METHODS: We analyzed 120 high-quality genomes representing six single-species ESKAPE groups (20 genomes per species). Genome quality was assessed using CheckM2. Species-specific pan-genomes were constructed with Roary, AMR genes were identified using AMRFinderPlus, and virulence factors were detected against the VFDB database using DIAMOND. AMR genes were mapped to core and accessory genome compartments through integration of Prokka annotations and Roary outputs. Statistical associations were evaluated using Fisher's exact tests and correlation analyses, with false discovery rate correction applied within each test family. Core-genome maximum-likelihood phylogenies were reconstructed to provide an evolutionary framework. RESULTS: Pan-genome sizes ranged from 4720 to 17,272 genes, with Enterobacter and Pseudomonas possessing the largest accessory genomes. Multidrug resistance (MDR; resistance to ≥3 antimicrobial classes) was detected in 93.3% of strains. After false discovery rate correction, AMR genes remained significantly enriched in the accessory genomes of Enterobacter, Enterococcus, Klebsiella, and Staphylococcus, whereas Acinetobacter and Pseudomonas did not show significant enrichment in either genome compartment. Within-species analyses identified significant positive associations between accessory genome size and AMR class burden in Staphylococcus, Enterococcus, and Enterobacter, whereas the moderate Pearson correlation observed in Pseudomonas was not significant after FDR correction. Virulence factor repertoires varied markedly among species, with Pseudomonas exhibiting the highest burden and Enterococcus the lowest. CONCLUSIONS: ESKAPE pathogens display distinct patterns of resistance and virulence. Accessory genome expansion was associated with higher AMR burden in several species, whereas other species showed no significant association between accessory genome size and AMR burden and no significant enrichment of AMR genes in either genome compartment, highlighting the species-specific nature of AMR evolution.

Virulence Factors

Chromosome-level genome assembly of hawthorn spider mite, Amphitetranychus viennensis (Acari: Tetranychidae).

The hawthorn spider mite, Amphitetranychus viennensis, is a major pest of orchards and ornamentals in the Palaearctic region, with adaptability and acaricide resistance. The lack of high-quality genomic resources limits understanding of its detoxification mechanisms and the development of RNAi-based pest control strategies. In this study, we utilized Illumina, Pacific Biosciences (PacBio), and Hi-C sequencing technologies to assemble a chromosome-level reference genome of A. viennensis. The assembled genome spans 141.96 Mb, with a contig N50 of 1.35 Mb. BUSCO analysis confirmed a high level of completeness, covering 91.6% of annotated genes. The assembly includes 50.97 Mb of repetitive sequences, representing 35.93% of the genome, and annotates 13,968 protein-coding genes. Using Hi-C sequencing, we anchored 47 contigs to three chromosomes, accounting for 97.27% of the estimated nuclear genome and achieving a contig N50 of 45.83 Mb. This high-quality genome assembly provides a valuable foundation for evolutionary and genomic research on spider mites, while also serving as a genetic resource to inform molecular control strategies and support sustainable pest management.

Animals

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals

Diversity of Salmonella enterica isolates from urban river and sewage water in Blantyre, Malawi.

BACKGROUND: Salmonella enterica encompasses over 2,600 serovars, including several commonly associated with severe infection in humans. Salmonella is a major cause of sepsis in Africa; however, diagnosis requires clinical microbiology facilities. Environmental surveillance has the potential to play a role in Salmonella surveillance. METHODS: We undertook water-based environmental surveillance in Blantyre, Malawi, from 2018-2020, taking samples from rivers (87.9%), a sewage plant (8.85%) and other water sources (3.24%), isolating and storing 1,042 non-typhoidal Salmonella (NTS) isolates in this period. Of these, 341 NTS isolates were whole genome sequenced, genome quality was checked, duplicate genomes from any given sample were removed and core genome phylogeny was reconstructed. AMRFinder, PathogenWatch and SISTR were used to further investigate serovar, sequence type and antimicrobial resistance determinants. RESULTS: After quality checks, and removal of duplicate genomes, 270 NTS genomes remained for further analysis. Multiple Salmonella serovars associated with human infection were detected, of which S. Typhimurium (55/270 isolates) was the most common, including 44 of Sequence Type (ST) 313, a serovar commonly associated with severe invasive disease (iNTS). Six lineage 2 ST313 genomes possessed AMR genes predicting multidrug resistance (MDR), while 29 lineage 3 isolates contained no AMR predictive genes. PCR based detection of staG has been proposed as a diagnostic marker of S. Typhi; however, all eight genomes that contained staG identified as Salmonella enterica serovar Orion, raising concerns about the specificity of this marker as a monoplex for environmental surveillance of S. Typhi. DISCUSSION: The study identified diverse Salmonella serovars in the environment, including those reported to cause invasive disease, emphasizing the complex but potentially valuable contribution of implementing environmental surveillance for Salmonella in high burden areas lacking diagnostic microbiology capacity.

Sewage

Complete genomes from a xenic Dolichospermum flosaquae FBCC-A233 culture reveal genome-inferred metabolic asymmetry with associated bacteria.

Cyanobacteria form phycosphere communities with associated bacteria, but genome-resolved resources are needed to formulate testable hypotheses about their metabolic interactions. Here, we reconstructed three complete circular genomes from a unialgal xenic culture, including Dolichospermum flosaquae FBCC-A233 and two associated alphaproteobacterial genomes assigned to Sphingorhabdus sp. and Brevundimonas sp. Genome-wide read mapping and genome-quality assessment supported the three recovered genomes as high-quality circular reconstructions. Comparative genome analysis placed the cyanobacterial genome within the Dolichospermum flosaquae species cluster under the GTDB framework, while the associated bacterial genomes represented Sphingorhabdus sp. and a putative undescribed Brevundimonas species-level lineage. Genome architecture analysis indicated reduced genome size and gene content in Brevundimonas relative to genus-level references although additional metrics did not support a strong conclusion of classical genome streamlining. Selected KEGG module and KO-level reconstructions indicated genome-inferred metabolic asymmetries across the consortium. FBCC-A233 encoded photosynthesis- and nitrogen-related modules and a BioU-mediated de novo biotin biosynthesis route, whereas the associated bacteria lacked complete de novo biotin biosynthesis but retained biotin-dependent carboxylase genes. FBCC-A233 also encoded extensive anaerobic corrinoid biosynthesis potential; however, canonical DMB-containing cobalamin completion, cobamide identity, and complete transporter systems were not resolved. Together, these complete genomes provide a genome-resolved resource for investigating genome-inferred metabolic differentiation and ecological interactions in cyanobacteria-associated bacterial consortia.IMPORTANCEPhycosphere interactions between cyanobacteria and associated bacteria can shape aquatic microbial communities, but many proposed interactions remain difficult to evaluate without genome-resolved resources. This study provides three complete circular genomes from a unialgal xenic Dolichospermum flosaquae culture, capturing the cyanobacterium and two co-maintained bacterial associates. Our analysis identifies genome-inferred metabolic asymmetries, particularly in biotin- and cobamide-related pathways. D. flosaquae FBCC-A233 encoded candidate de novo biotin and corrinoid biosynthesis capacity, whereas the associated bacteria lacked complete de novo pathways but retained cofactor-dependent enzymes. These findings nominate cofactor-related dependencies as experimentally testable hypotheses while emphasizing unresolved uptake, export, cobamide identity, and growth-dependence mechanisms. The complete genomes and KO-level reconstructions generated here provide a resource for future studies of cyanobacteria-associated consortia.

Genome, Bacterial

Chromosome-level haplotype-resolved genome assembly of the giant honeycomb oyster, Hyotissa hyotis.

The giant honeycomb oyster, Hyotissa hyotis, a common bivalve inhabitant of tropical and subtropical coastal waters, holds significant ecological and economic importance due to its shell characteristics, rapid growth, and high-quality adductor muscle. However, the lack of high-quality genome has impeded the genetic study and artificial breeding of this species. In this study, we provided the first chromosomal-level haplotype-resolved assembly for the H. hyotis (2n = 20) by combining PacBio HiFi long-read and Hi-C sequencing. We obtained a haplotype-resolved assembly of 3.39 Gb in size, of which 96.69% were anchored to 20 chromosomes. The haplotype A and B genome (HapA and HapB) was 1,639.90 and 1,643.23 Mb in size, respectively. Accordingly, a total of 28,720 and 29,003 protein-coding genes were annotated from HapA and HapB. Through the BUSCO evaluation, the assembly and annotation results exhibited the completeness value of 94.65% and 94.03% for HapA, while 94.13% and 92.98% for HapB. This high-quality genome assembly provides valuable resource for further genetic studies and genetic improvement of the group of oysters.

Animals

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43 Mb, with a Scaffold N50 length of 19.05 Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75 Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals