PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Metadata”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S↗

Proposal for the creation of a Web-based heterogeneous distributed archive for psychological data.

This report presents a proposal to create archives of data from psychological research and associated metadata Web pages and link them into a heterogeneous distributed archive on the World-Wide Web. Several specific recommendations are made concerning some of the issues faced by the data archivist and data archive user hoping to use the Web. In particular, a recommendation is made to create a publicly accessible Web page for each data set and place keywords, experimental methods, data descriptions, pointers to journal articles, and pointers to other archive Web pages pertinent to this data set on this metadata Web page. If the archivist includes a special keyword (PsychologyDataArchive) on the metadata Web page, Web-based search engines will automatically be able to subset all participating data archives for indexing and semantic analysis. The secondary data analyst can then include the word PsychologyDataArchive in his Web search and will be able to effectively find relevant participating Web data archives.

Archives↗

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics↗

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis↗

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software↗

The need for standardization and improved open (meta)data practices in metaproteomics.

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices. Video Abstract.

Proteomics↗

Guidelines for the effective use of entity-attribute-value modeling for biomedical databases.

PURPOSE: To introduce the goals of EAV database modeling, to describe the situations where entity-attribute-value (EAV) modeling is a useful alternative to conventional relational methods of database modeling, and to describe the fine points of implementation in production systems. METHODS: We analyze the following circumstances: (1) data are sparse and have a large number of applicable attributes, but only a small fraction will apply to a given entity; (2) numerous classes of data need to be represented, each class has a limited number of attributes, but the number of instances of each class is very small. We also consider situations calling for a mixed approach where both conventional and EAV design are used for appropriate data classes. RESULTS AND CONCLUSIONS: In robust production systems, EAV-modeled databases trade a modest data sub-schema for a complex metadata sub-schema. The need to design the metadata effectively makes EAV design potentially more challenging than conventional design.

Database Management Systems↗

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry↗

Automated cryoEM data acquisition and analysis of 284742 particles of GroEL.

One of the goals in developing our automated electron microscopy data acquisition system, Leginon, was to improve both the ease of use and the throughput of the process of acquiring low dose images of macromolecular specimens embedded in vitreous ice. In this article, we demonstrate the potential of the Leginon system for high-throughput data acquisition by describing an experiment in which we acquired images of more than 280,000 particles of GroEL in a single 25 h session at the microscope. We also demonstrate the potential for an automated pipeline for molecular microscopy by showing that these particles can be subjected to completely automated procedures to reconstruct a three-dimensional (3D) density map to a resolution better than 8 A. In generating the 3D maps, we used a variety of metadata associated with the data acquisition and processing steps to sort and select the particles. These metadata provide a number of insights into factors that affect the quality of the acquired images and the resulting reconstructions. In particular, we show that the resolution of the reconstructed 3D density maps improves with decreasing ice thickness. These data provide a basis for assessing the capabilities of high-throughput macromolecular microscopy.

Chaperonin 60↗

Effect of repeated mass drug administration on the transmission of yaws: a retrospective genomic epidemiology study.

BACKGROUND: Yaws, a neglected tropical disease caused by Treponema pallidum subspecies pertenue (T p pertenue), has evaded eradication, in part due to a high proportion of asymptomatic cases. Repeated mass drug administration (MDA), whereby an entire population is repeatedly treated irrespective of disease, could provide a solution. Here, we aimed to investigate the effect of MDA on the genomic epidemiology of T p pertenue. METHODS: We conducted a retrospective genomic epidemiology study on samples collected during a cluster-randomised trial of mass administration of azithromycin for yaws eradication in the Namatanai District of Papua New Guinea. Participants were in 38 wards (administrative units encompassing several villages) in three local-level government areas (LLGs). The experimental group received an initial round of MDA followed by two further rounds 6 months and 12 months after the first round. The control group received one round of MDA followed by two rounds of treatment targeting clinical cases and contacts only, on the same schedule as the MDA in the experimental group. A follow-up survey on both groups was done 18 months after the first MDA round. Swab samples were collected at each round from ulcerative and nodular skin lesions, and blood was collected by finger-prick for serological testing at 18 months. Metadata on ulcer size (cm) and duration (days) were recorded at each round, and treponemal and non-treponemal antibodies were recorded at 18 months. Samples from swabs positive for T p pertenue underwent library preparation and whole-genome sequencing. We examined the phylogenetic relationships between genomes, linking them with geospatial and patient metadata to understand the impact of MDA on T p pertenue diversity and transmission. FINDINGS: Swabs collected from 297 individuals with active yaws from April 30, 2018, to Nov 2, 2019, yielded 222 good-quality Tp pertenue genomes. We identified 20 sublineages of T p pertenue in the control group and 21 in the experimental group at the beginning of the study. At the end of the study, there were 13 sublineages in the control group and three in the experimental group, of which two persisted in both groups. Three sublineages not detected at baseline were observed in the control group after commencing MDA. The two sublineages that persisted in both groups had non-synonymous mutations in penicillin-binding proteins. One of these sublineages evolved macrolide resistance in three individuals and was associated with lowered treponemal antibody (p=0&#xb7;0036) and longer ulcer duration (p=0&#xb7;015). Despite the study taking place within a small island, sublineages were geographically clustered, with pairs of samples from the same ward (odds ratio 7&#xb7;1, 95% CI 5&#xb7;7-8&#xb7;8; p<0&#xb7;0001) or neighbouring wards (4&#xb7;3, 3&#xb7;3-5&#xb7;4; p<0&#xb7;0001) more likely to share the same sublineages compared with pairs from different LLGs. Additionally, older individuals were more likely to share sublineages than were younger individuals (1&#xb7;5, 1&#xb7;2-1&#xb7;9; p<0&#xb7;0001). INTERPRETATION: Repeated MDA was successful in reducing and maintaining the genetic diversity of T p pertenue at a low level but was associated with the development of macrolide resistance. Yaws re-emergence after MDA was attributed to multiple sublineages, of which the majority were detected in the population before MDA. Participants within the same ward were more likely to share sublineages than those that were more widely geographically separated, suggesting that re-emergence was driven by local transmission. These findings could inform future yaws elimination strategies. FUNDING: European Research Council, EU, Provincial Deputation of Barcelona, Barber&#xe0; Solid&#xe0;ria Foundation, Wellcome, and Fundaci&#xf3; "la Caixa".

Adolescent↗

Modelling data across labs, genomes, space and time.

Logical models and physical specifications provide the foundation for storage, management and analysis of complex sets of data, and describe the relationships between measured data elements and metadata - the contextual descriptors that define the primary data. Here, we use imaging applications to illustrate the purpose of the various implementations of data specifications and the requirement for open, standardized, data formats to facilitate the sharing of critical digital data and metadata.

Animals↗

A dynamic problem to knowledge linking Semantic Web service based on clinical codes.

Many information needs arise during everyday clinical practice. Problem to knowledge linking aims to answer these needs by providing contextually appropriate medical knowledge in the right place and at the right time. Empirical evidence shows that well-informed physicians and patients are able to make better clinical decisions that positively affect healthcare outcomes. This paper reports on the design and development of a re-usable and flexible Semantic Web problem to knowledge linking service. The service makes use of metadata and clinical codes contextually to link disparate Electronic Patient Record clients to resources in an online medical knowledge service (HealthCyberMap). Clinical codes act as crisp knowledge hooks, providing a reliable common backbone language for communication between Electronic Patient Records and HealthCyberMap. Ideas to improve the service are also discussed. By minimizing irrelevant leads (noise) and reducing the time needed to find relevant information (the right contextually relevant knowledge is linked to real patient data in the Electronic Patient Record), the system is potentially beneficial. The actual success of the system will depend on the quality and granularity of metadata it uses and the topical coverage and quality of resources to which it points.

Databases as Topic↗

Evolving from bioinformatics in-the-small to bioinformatics in-the-large.

We argue the significance of a fundamental shift in bioinformatics, from in-the-small to in-the-large. Adopting a large-scale perspective is a way to manage the problems endemic to the world of the small-constellations of incompatible tools for which the effort required to assemble an integrated system exceeds the perceived benefit of the integration. Where bioinformatics in-the-small is about data and tools, bioinformatics in-the-large is about metadata and dependencies. Dependencies represent the complexities of large-scale integration, including the requirements and assumptions governing the composition of tools. The popular make utility is a very effective system for defining and maintaining simple dependencies, and it offers a number of insights about the essence of bioinformatics in-the-large. Keeping an in-the-large perspective has been very useful to us in large bioinformatics projects. We give two fairly different examples, and extract lessons from them showing how it has helped. These examples both suggest the benefit of explicitly defining and managing knowledge flows and knowledge maps (which represent metadata regarding types, flows, and dependencies), and also suggest approaches for developing bioinformatics database systems. Generally, we argue that large-scale engineering principles can be successfully adapted from disciplines such as software engineering and data management, and that having an in-the-large perspective will be a key advantage in the next phase of bioinformatics development.

Computational Biology↗

myGrid: personalised bioinformatics on the information grid.

MOTIVATION: The (my)Grid project aims to exploit Grid technology, with an emphasis on the Information Grid, and provide middleware layers that make it appropriate for the needs of bioinformatics. (my)Grid is building high level services for data and application integration such as resource discovery, workflow enactment and distributed query processing. Additional services are provided to support the scientific method and best practice found at the bench but often neglected at the workstation, notably provenance management, change notification and personalisation. RESULTS: We give an overview of these services and their metadata. In particular, semantically rich metadata expressed using ontologies necessary to discover, select and compose services into dynamic workflows.

Computational Biology↗

The Hospital Data Project: comparing hospital activity within Europe.

BACKGROUND: The ability to measure and compare hospital activity between EU member states is important for policy, planning, financing and assessment of population health. Earlier initiatives in this area have been largely directed at standardising high-level indicator definitions without proper account of differences in health systems and health information systems. The Hospital Data Project (HDP) develops a methodology for improved comparability of hospital inpatient and day case activity data across Europe and produces a pilot common data set. All EU members, Iceland and the World Health Organisation are participants. METHODS: The approach comprises a detailed inventory of patient-level hospital data, identification of common areas, specification of data transformations and production of pilot data sets and metadata in a common format. An expert group developed a new diagnosis shortlist based on ICD-10. The project takes account of current work in the area of health care and morbidity indictors and applies the functional specification of health systems developed by the OECD. RESULTS: Seventeen countries have submitted data and metadata in the common format for a single year. Data on inpatients and day cases are classified by age, gender, diagnosis and type of admission. Numbers of hospital discharges, mean and median lengths of stay and population rates are reported. Test data on selected hospital procedures has also been collected. The full data set contains approximately 500,000 records, and software has been designed to facilitate validation and use. CONCLUSION: Results to date are promising. It is a first step in a complex area, and further work is required to extend and refine this approach.

Diagnosis-Related Groups↗

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article↗

Enhancing the MeSH thesaurus to retrieve French online health resources in a quality-controlled gateway.

The amount of health information available on the Internet is considerable. In this context, several health gateways have been developed. Among them, CISMeF (Catalogue and Index of Health Resources in French) was designed to catalogue and index health resources in French. The goal of this article is to describe the various enhancements to the MeSH thesaurus developed by the CISMeF team to adapt this terminology to the broader field of health Internet resources instead of scientific articles for the medline bibliographic database. CISMeF uses two standard tools for organizing information: the MeSH thesaurus and several metadata element sets, in particular the Dublin Core metadata format. The heterogeneity of Internet health resources led the CISMeF team to enhance the MeSH thesaurus with the introduction of two new concepts, respectively, resource types and metaterms. CISMeF resource types are a generalization of the publication types of medline. A resource type describes the nature of the resource and MeSH keyword/qualifier pairs describe the subject of the resource. A metaterm is generally a medical specialty or a biological science, which has semantic links with one or more MeSH keywords, qualifiers and resource types. The CISMeF terminology is exploited for several tasks: resource indexing performed manually, resource categorization performed automatically, visualization and navigation through the concept hierarchies and information retrieval using the Doc'CISMeF search engine. The CISMeF health gateway uses several MeSH thesaurus enhancements to optimize information retrieval, hierarchy navigation and automatic indexing.

Abstracting and Indexing↗

A simple method for serving Web hypermaps with dynamic database drill-down.

BACKGROUND: HealthCyberMap http://healthcybermap.semanticweb.org aims at mapping parts of health information cyberspace in novel ways to deliver a semantically superior user experience. This is achieved through "intelligent" categorisation and interactive hypermedia visualisation of health resources using metadata, clinical codes and GIS. HealthCyberMap is an ArcView 3.1 project. WebView, the Internet extension to ArcView, publishes HealthCyberMap ArcView Views as Web client-side imagemaps. The basic WebView set-up does not support any GIS database connection, and published Web maps become disconnected from the original project. A dedicated Internet map server would be the best way to serve HealthCyberMap database-driven interactive Web maps, but is an expensive and complex solution to acquire, run and maintain. This paper describes HealthCyberMap simple, low-cost method for "patching" WebView to serve hypermaps with dynamic database drill-down functionality on the Web. RESULTS: The proposed solution is currently used for publishing HealthCyberMap GIS-generated navigational information maps on the Web while maintaining their links with the underlying resource metadata base. CONCLUSION: The authors believe their map serving approach as adopted in HealthCyberMap has been very successful, especially in cases when only map attribute data change without a corresponding effect on map appearance. It should be also possible to use the same solution to publish other interactive GIS-driven maps on the Web, e.g., maps of real world health problems.

Journal Article↗