PubMed HealthSearch

SEARCH · PubMed Health

Results for “Non-redundant dataset”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

An atlas of non-redundant sequences and structures of transcription factor assemblies across domains of life.

Transcription factors (TFs) regulate gene expression by controlling the recruitment of transcriptional machinery to regulatory regions of the genome. Nearly 10% of the human genome encodes TFs, making them one of the largest protein families. Despite their central roles in gene regulation, TFs are historically considered challenging therapeutic targets due to their complex interactions with DNA, RNA and associated proteins. Although recent progress in studying TFs both at molecular and structural level excels our understanding on their function, yet a universal rule decoding their recognition process remains elusive. Here, we present a curated non-redundant dataset of TFs with 3570 sequences and 377 structures. We further characterize "unique interfaces" by quantifying interface identity across interacting chains in TF assemblies. Surprisingly, our data shows that the "unique interfaces" have optimal size ranging from 2000 Å2 to 4000 Å2 irrespective of their quaternary assembly. To understand the functional diversity, we integrate sequence motifs, structural domains, subcellular localization and functional enrichment of TFs. We have also catalogued association of TFs with various human diseases. Our dataset provides a comprehensive platform to perform large scale analysis of TF-assemblies and aid in computational methods for their prediction across domains of life.

Gene regulation

FoldX force field revisited, an improved version.

MOTIVATION: The FoldX force field was originally validated with a database of 1000 mutants at a time when there were few high-resolution structures. Here, we have manually curated a database of 5556 mutants affecting protein stability, resulting in 2484 highly confident mutations denominated FoldX stability dataset (FSD), represented in non-redundant X-ray structures with <2.5&#x2009;&#xc5; resolution, not involving duplicates, metals, or prosthetic groups. Using this database, we have created a new version of the FoldX force field by introducing pi stacking, pH dependency for all charged residues, improving aromatic-aromatic interactions, modifying the Ncap contribution and &#x3b1;-helix dipole, recalibrating the side-chain entropy of methionine, adjusting the H-bond parameters, and modifying the solvation contribution of tryptophan and others. RESULTS: These changes have led to significant improvements for the prediction of specific mutants involving the above residues/interactions and a statistically significant increase of FoldX predictions, as well as for the majority of the 20 aa. Removing all training sets data from FSD [Validation FoldX Stability Dataset (VFSD) dataset] resulted in improved predictions from R&#x2009;=&#x2009;0.693 (RMSE&#x2009;=&#x2009;1.277&#x2009;kcal/mol) to R&#x2009;=&#x2009;0.706 (RMSE&#x2009;=&#x2009;1.252&#x2009;kcal/mol) when compared with the previously released version. FoldX achieves 95% accuracy considering an error of &#xb1;0.85&#x2009;kcal/mol in prediction and an area under the curve&#x2009;=&#x2009;0.78 for the VFSD, predicting the sign of the energy change upon mutation. AVAILABILITY AND IMPLEMENTATION: FoldX versions 4.1 and 5.1 are freely available for academics at https://foldxsuite.crg.eu/.

Databases, Protein

An integrated global resource of wetland microbiomes linking environmental metadata, community profiles, and genome-resolved metabolic traits.

Wetlands are biogeochemical hotspots pivotal to global carbon and nutrient cycling, yet genome-resolved studies across diverse wetland types remain limited. To address this, we constructed a global wetland metagenomic dataset, integrating environmental metadata, community profiles, and genome-resolved metabolic traits. This dataset comprises 1,962 samples-including 129 newly sequenced field-collected samples-from lakes, rivers, paddies, marshes, and coastal wetlands, spanning water, soil, and sediment habitats. We generated comprehensive taxonomic profiles for all 1,962 samples, and used 251 samples to reconstruct 5,704 sample-specific metagenome-assembled genomes (MAGs). These MAGs were subsequently dereplicated to establish a normalized, non-redundant catalog of 4,164 representative genomes. We further mapped gene repertoires to 549 KEGG modules to decode the metabolic potential of all 5,704 MAGs. This dataset depicts an overview of microbial genomic diversity across global wetlands and provides a comprehensive resource for understanding the metabolic capabilities, ecology, and evolution of wetland microbiomes.

Wetlands

Farming reshapes the gut resistome, virulome, and mobilome of Cervidae.

The rapid expansion of cervid farming raises concerns about antimicrobial resistance (AMR) dissemination, yet its impact on the Cervidae gut microbiome remains poorly characterized. We integrated 89 newly sequenced fecal metagenomes with 599 publicly available datasets, comprising 285 metagenomes from farmed cervids and 370 from wild cervids, to construct a catalog of 15,494 non-redundant metagenome-assembled genomes (MAGs) representing 2,401 species. Our analysis demonstrates that farming profoundly reshapes the gut microbiome's functional composition. Specifically, farmed cervids exhibited significantly higher relative abundance, diversity, and heterogeneity of antimicrobial resistance genes (ARGs) compared to wild counterparts. We observed a robust synergistic relationship between ARGs, virulence factor genes, and mobile genetic element (MGE)-associated genes, identifying 70 ARG-MGE combinations as evidence of potential horizontal gene transfer. Plasmid profiling further suggested that a subset of ARGs may be associated with conjugative plasmids, with plasmid-associated ARGs being significantly more abundant in farmed than in wild cervids. Virome analyses indicated that bacteriophages, particularly Siphoviridae, may serve as mobile reservoirs for ARGs. Notably, Cervidae shared 268 ARG types with humans, including 23&#xa0;high-risk genes associated with resistance to clinically important antibiotics (e.g. tetX1, vanRD, and bla-CTX-M-178), with Escherichia coli as a key cross-host carrier. These findings highlight that human-impacted cervid gut microbiomes are significant environmental reservoirs of clinically relevant AMR, underscoring the necessity for enhanced antibiotic stewardship and resistance surveillance in managed wildlife within a One Health framework.

Animals

Interkingdom remodeling of the intestinal bacteriome and virome during Toxoplasma gondii infection in rats.

Toxoplasma gondii infection is associated with intestinal microbiome disruption, but its effects on genome-resolved bacterial populations, the gut virome, and bacteriome-virome relationships remain poorly understood. Using previously generated shotgun metagenomic datasets from 36 intestinal samples collected from 18 Sprague-Dawley rats across control, acute, and chronic infection groups, we reconstructed 294 quality-filtered, non-redundant bacterial metagenome-assembled genomes (MAGs) and identified 899 medium-to-high-quality viral operational taxonomic units (vOTUs) from assembled metagenomic contigs. Infection was associated with reduced bacterial richness in the small intestine during both acute and chronic stages and lower Shannon diversity during chronic infection. In contrast, large-intestinal &#x3b1;-diversity remained stable despite significant compositional reorganization. Taxonomic changes included increased Lactobacillus intestinalis, Limosilactobacillus reuteri, and Prevotella sp900547005, together with decreased Rothia sp002492045 and Akkermansia muciniphila. Functional profiling revealed region- and stage-specific changes in predicted bacterial metabolic potential, including reduced energy-related pathways and carbohydrate-active enzyme abundance. The virome also showed significant compositional changes in both intestinal regions. Quimbyviridae and Podoviridae_crAss-like viruses decreased in the small intestine during chronic infection, while Quimbyviridae, Flandersviridae, and Podoviridae_crAss-like viruses showed stage-specific decreases in the large intestine. Predicted bacterial hosts were assigned to 48.39% of vOTUs, with Lachnospiraceae and Ruminococcaceae being the most frequently linked families. Trans-kingdom networks further revealed region-specific positive and negative abundance correlations between bacterial and viral taxa. These findings extend previous microbiota-metabolome observations by integrating genome-resolved bacteriome analysis with contig-based virome profiling, providing a foundation for future mechanistic studies of toxoplasmosis-associated microbiome remodeling.

Gut virome

Genome-resolved analysis of colonization factor repertoires reveals ecological stratification in cervid gut microbiomes.

INTRODUCTION: Colonization factors (CFs) are important microbial traits associated with persistence and host adaptation in the gut, yet their large-scale organization in cervid gut microbiomes remains unclear. METHODS: A total of 3,311 non-redundant high-quality metagenome-assembled genomes (MAGs), derived from 688 cervid gut metagenomic samples across 15 publicly available projects and one in-house dataset, were analyzed. CF-associated genes were identified by comparison against the GHA CF database, and CF repertoires were characterized at genome, host-species, and gastrointestinal-segment levels. RESULTS: A total of 138,729 CF-associated genes spanning 71 CF families were identified. MAGs from Cervinae contained richer CF repertoires than those from Caprinae, and CF47 (Peptidase_C69), CF24_29 (QueH), and CF18 (Glycos_transf_2) were among the most prevalent families. CF repertoires were strongly structured by taxonomy, showed a moderate association with bacterial phylogenetic distance, and formed two recurrent genome-level configurations with distinct KEGG functional profiles. Integration of sample metadata further revealed differentiation of CF repertoires across host species and gastrointestinal segments, representing the major ecological dimensions examined in this study. Segment-associated CF variation was accompanied by redistribution of broader functional profiles, including enrichment of carbohydrate and lipid metabolism in the jejunum, membrane transport in the ileum, xenobiotics biodegradation in the cecum, and environmental adaptation in the rumen. DISCUSSION: These findings provide a genome-resolved view of CF repertoire organization in cervid gut microbiomes and demonstrate that colonization-associated functions are structured across microbial lineages and ecological contexts. This study highlights the importance of considering microbial taxonomy and host-associated environments when interpreting the distribution of CF repertoires in mammalian gut ecosystems.

Cervidae

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics