PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Metadata”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Multidimensional prophage profiling of carbapenem-resistant Enterobacteriaceae in Thailand: a nationwide, multicentre, genomic study.

BACKGROUND: Prophages influence bacterial fitness, resistance, and evolution, yet their epidemiology remains poorly understood in carbapenem-resistant Enterobacteriaceae (CRE). In this nationwide study in Thailand, we aimed to describe prophage repertoires in clinical CRE isolates and to explore their potential relevance for molecular epidemiology. METHODS: We performed a nationwide, retrospective, genomic analysis of all CRE clinical isolates collected through our previous national surveillance study involving 11 hospitals in 11 provinces in Thailand between March 25, 2012, and Jul 21, 2017. Whole-genome sequencing data from 747 CRE isolates were analysed. Intact prophages were identified using PHAge Search Tool Enhanced Release (PHASTER) and clustered by nucleotide sequence similarity. Prophage profiles were compared across multilocus sequence types, carbapenemase genotypes, specimens, geography, and patient demographics (age and sex). FINDINGS: Of the included 747 CRE isolates, 170 (23%) were Escherichia coli and 577 (77%) were Klebsiella pneumoniae. 220 (29%) of 747 strains had been isolated from female patients and 264 (35%) from male patients; metadata on patient sex were missing for 263 (35%) isolates. The median patient age was 63 years (IQR 50-72). 71 (10%) of isolates were from blood, 283 (38%) from sputum, 284 (38%) from urine, and 109 (15%) from other specimens. 374 distinct prophage clusters were identified, with significantly more prophages per genome in K pneumoniae (mean 3&#xb7;01 [SD 1&#xb7;55]) than in E coli (1&#xb7;64 [1&#xb7;46]; p<0&#xb7;0001). Prophage repertoires largely mirrored bacterial multilocus sequence types. However, even within the highly clonal K pneumoniae sequence type 16 lineage, discrete prophage variation was identified, with common profiles observed in geographically dispersed patients. Respiratory K pneumoniae frequently carried a mosaic prophage with environmental signatures and a type VI secretion system, whereas blood-derived E coli harboured a prophage with immune-modulating genes. Distinct prophage clusters were observed across clinical specimens, age groups, carbapenemase genotype, and geographical region. Strains coharbouring blaNDM-1 plus blaOXA-232 (114 [15%] of 747) had the highest prophage loads. INTERPRETATION: The prophage content was shaped by the bacterial lineage, ecological niche, and temporal dynamics, providing an additional layer of epidemiological resolution beyond conventional genome typing. Integrating prophage profiling into molecular surveillance frameworks could help to identify transmission events, improve infectious source attribution, and enhance infection control strategies. FUNDING: Japan Agency for Medical Research and Development.

Female↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78&#xb7;3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48&#xb7;1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60&#xb7;2%) had mixed ethnicity, 90 (17&#xb7;8%) were White, 56 (11&#xb7;1%) were Black, 13 (2&#xb7;6%) were Indigenous, and six (1&#xb7;2%) were Asian. 277 (54&#xb7;9%) had symptomatic disease and 228 (45&#xb7;1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76&#xb7;9%] of 277 vs 195 [85&#xb7;5%] of 228; p=0&#xb7;37), THD (median 0&#xb7;39 [IQR 0&#xb7;06-0&#xb7;62] vs 0&#xb7;50 [0&#xb7;09-0&#xb7;65]; p=0&#xb7;12), or LBI (0&#xb7;00863 [0&#xb7;00810-0&#xb7;00988] vs 0&#xb7;00871 [0&#xb7;00829-0&#xb7;01020]; p=0&#xb7;088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0&#xb7;56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗

Database and tools for analysis of topographic organization and map transformations in major projection systems of the brain.

Integration of dispersed and complicated information collected from the brain is needed to build new knowledge. But integration may be hampered by rigid presentation formats, diversity of data formats among laboratories, and lack of access to lower level data. We have addressed some of the fundamental issues related to this challenge at the level of anatomical data, by producing a coordinate based digital atlas and database application for a major projection system in the rat brain: the cerebro-ponto-cerebellar system. This application, Functional Anatomy of the Cerebro-Cerebellar System in rat (FACCS), is available via the Rodent Brain WorkBench (http://www.rbwb.org). The data included are x,y,z-coordinate lists describing exact distributions of tissue elements (axonal terminal fields of axons, or cell bodies) that are labeled with axonal tracing techniques. All data are translated to a common local coordinate system to facilitate across animal comparison. A search capability allows queries based on, e.g. location of tracer injection sites, tracer category, size of the injection sites, and contributing author. A graphic search tool allows the user to move a volume cursor inside a coordinate system to detect particular injection sites having connections to a specific tissue volume at chosen density levels. Tools for visualization and analysis of selected data are included, as well as an option to download individual data sets for further analysis. With this application, data and metadata from different experiments are mapped into the same information structure and made available for re-use and re-analysis in novel combinations. The application is prepared for future handling of data from other projection systems as well as other data categories.

Anatomy, Artistic↗

Data input module for Birth Defects Systems Manager.

The need for a computational bioinformatics infrastructure to manage the vast digital information from functional genomics and proteomics motivated us to develop Birth Defects Systems Manager (BDSM) as an open resource to facilitate analysis and discovery in developmental biology and developmental toxicity. This report describes the design, development and implementation of the data loading module of BDSM, referred to as LoadBDSM. It includes a shared data directory resource that can be granted various levels of security for different research groups or investigators to manage experimental datasets individually or in groups. LoadBDSM allows the upload of data and experiment details using controlled semantics for developmental exposure (toxicant, dosing scenario, intervention), biological sample (species, tissue, stage) and disease outcome (time, risk, phenotype). It adheres to existing controlled vocabulary plus rules of inference (ontologies) for experiment, data and metadata annotations. LoadBDSM extends the capabilities of BDSM to support the emergence of "embryo-formatics" defined here as the data, information and knowledge from genomic sciences applied to, or derived from, an embryological context. This includes, but is not limited to, delineating pathways and biological regulatory networks for specific chemicals or classes of developmental toxicants, developing novel biomarkers indicative of exposure and/or predictive of adverse effects, and integrating modern computing and information technology with data from molecular biology.

Abnormalities, Drug-Induced↗

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma↗

Protocol for histology-anchored macroscopic staging of gonadal maturity in exploited fishes.

Here, we present a protocol to assign gonadal maturity stages in commercially exploited fishes using a histology-anchored workflow. We describe steps for recording field metadata, photographing gonads, fixing central gonadal tissue, and processing paraffin sections. We then detail procedures for staining sections with hematoxylin and eosin, diagnosing gametogenic features, and assigning stages using a common reproductive-phase framework with species- and sex-specific reference descriptors. This protocol standardizes documentation and decision logic rather than proposing a new maturity scale.

Developmental biology↗

Continuing dental education on the World Wide Web.

Continuing dental education (CDE) courses delivered on the World Wide Web (Web CDE) offer numerous advantages over traditional CDE; however, two major issues--location of suitable courses and course quality--need resolution. Locating high-quality courses is difficult due to the lack of the standardized metadata that allows search engines to match courses to practitioners' needs. Web directories created by professional organizations are beginning to show promise, but require further development. Search engines and Web directories are discussed and improvements currently underway summarized. Course quality remains a highly significant concern. A national effort to create Web CDE course quality standards is underway that includes proposed standards. These proposed standards are summarized and used to comment on the current state of Web CDE courses. Examples are given when possible. Three emerging Web CDE technologies and a look to the future of Web CDE are discussed.

Computer-Assisted Instruction↗

Bioconductor: an open source framework for bioinformatics and computational biology.

This chapter describes the Bioconductor project and details of its open source facilities for analysis of microarray and other high-throughput biological experiments. Particular attention is paid to concepts of container and workflow design, connections of biological metadata to statistical analysis products, support for statistical quality assessment, and calibration of inference uncertainty measures when tens of thousands of simultaneous statistical tests are performed.

Animals↗

A system for simultaneous multiple subject, multiple stimulus modality, and multiple channel collection and analysis of sensory evoked potentials.

A system has been developed for collecting sensory evoked potentials simultaneously from multiple channels for multiple subjects at up to 80 kHz sample rate per channel. Sample rates up to 200 kHz are available for four or less chambers and a single channel per chamber. A variety of visual, somatosensory, and auditory stimuli may be presented singly or simultaneously. Collected waveforms are associated with searchable text (metadata) to allow convenient selection from a relational database. Multiple waveforms can then be easily grouped for analysis and processed. Results can be exported to other software for further graphics or statistical processing. Scripting and event logging are available to provide automation and improve data confidence. Sample data are presented from control animals for each of the sensory modalities for comparison with historical data collected from other systems.

Animals↗

The impact of Life Science Identifier on informatics data.

Since the Life Science Identifier (LSID) data identification and access standard made its official debut in late 2004, several organizations have begun to use LSIDs to simplify the methods used to uniquely name, reference and retrieve distributed data objects and concepts. In this review, the authors build on introductory work that describes the LSID standard by documenting how five early adopters have incorporated the standard into their technology infrastructure and by outlining several common misconceptions and difficulties related to LSID use, including the impact of the byte identity requirement for LSID-identified objects and the opacity recommendation for use of the LSID syntax. The review describes several shortcomings of the LSID standard, such as the lack of a specific metadata standard, along with solutions that could be addressed in future revisions of the specification.

Computational Biology↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

ThermoData Engine (TDE): software implementation of the dynamic data evaluation concept.

The first full-scale software implementation of the dynamic data evaluation concept {ThermoData Engine (TDE)} is described for thermophysical property data. This concept requires the development of large electronic databases capable of storing essentially all experimental data known to date with detailed descriptions of relevant metadata and uncertainties. The combination of these electronic databases with expert-system software, designed to automatically generate recommended data based on available experimental data, leads to the ability to produce critically evaluated data dynamically or 'to order'. Six major design tasks are described with emphasis on the software architecture for automated critical evaluation including dynamic selection and application of prediction methods and enforcement of thermodynamic consistency. The direction of future enhancements is discussed.

Journal Article↗

Bringing chemical data onto the Semantic Web.

Present chemical data storage methodologies place many restrictions on the use of the stored data. The absence of sufficient high-quality metadata prevents intelligent computer access to the data without human intervention. This creates barriers to the automation of data mining in activities such as quantitative structure-activity relationship modelling. The application of Semantic Web technologies to chemical data is shown to reduce these limitations. The use of unique identifiers and relationships (represented as uniform resource identifiers, URIs, and resource description framework, RDF) held in a triplestore provides for greater detail and flexibility in the sharing and storage of molecular structures and properties.

Journal Article↗

ChemSem: an extensible and scalable RSS-based seminar alerting system for scientific collaboration.

A seminar announcement system based on the extensive use of XML-based data structures, CML/MathML for carrying more domain-specific molecular content, and open source software components is described. The output is a resource description framework (RDF) site summary (RSS) feed, which potentially carries many advantages over conventional announcement mechanisms, including the ability to aggregate and then sort multiple and diverse RSS feeds on the basis of declared metadata and to feed into RDF-based mechanisms for establishing links between different subject areas.

Journal Article↗

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics↗

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗

Temporal stability and lack of variance in microbiome composition and functionality in fit recreational athletes.

Human gut microbiome composition and function is influenced by environmental and lifestyle factors, including exercise and fitness. We studied the composition and functionality of the faecal microbiome of recreational (non-elite) runners (n&#x2009;=&#x2009;62) with serial shotgun metagenomics, at 4 time points over a 7-week period. Gut microbiome composition and function was stable over time. Grouping of samples on the basis of their fitness level (fair, good, excellent, and superior) or habitual training (low (4-6&#xa0;h/week), medium (7-9&#xa0;h/week), high (10-12&#xa0;h/week), and extreme (13&#x2009;+&#x2009;hours/week)) revealed no significant microbiome-related differences. Overall, the species Faecalibacterium prausnitzii, Blautia wexlerae, and Prevotella copri were the most abundant members of the gut microbiome. Analysis of co-abundance groups (CAGs) revealed no significant relationship between CAGs and fitness levels or training subgroups. Functional pathways were similar across all samples and timepoints with no clustering based on associated metadata. The most abundant genes identified within samples corresponded to pathways for nucleoside and nucleotide biosynthesis, amino acid biosynthesis, and cell wall biosynthesis. Collectively, these results describe the microbiome of active recreational runners and note temporal stability amongst participants.

Humans↗

Analysis of molecular data of Arabidopsis thaliana (L.) Heynh. (Brassicaceae) with Geographical Information Systems (GIS).

A Geographical Information System (GIS) is used to analyse allelic information of 13 sequenced loci of natural populations of Arabidopsis thaliana and to identify geographical structures. GIS provides tools for visualization and analysis of geographical population structures using molecular data. The geographical distribution of the number of variable positions in the alignments, the distribution of recombinant sequence blocks, and the distribution of a newly defined measure, the differentiation index, are studied. The differentiation index is introduced to measure the sequence divergence among individual plants sampled from various geographical localities. The numbers of variable positions and the differentiation index are also used for a metadata analysis covering about 26 kb of the genome. This analysis reveals, for the first time, differences in DNA sequence structures of geographically different populations of A. thaliana. The broadly defined west Mediterranean region consists of accessions with the highest numbers of polymorphic positions followed by the west European region. The GIS technology Kriging is used to define Arabidopsis specific diversity zones in Europe. The highest genetic variability is observed along the Atlantic coast from the western Iberian Peninsula to southern Great Britain, while lowest variability is found in central Europe.

Arabidopsis↗