PubMed HealthSearch

SEARCH · PubMed Health

Results for “interoperability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans

Pragmatic gynecologic cancer clinical trials: statements and roadmap from the Gynecologic Cancer InterGroup Chicago Brainstorming Meeting.

Randomized controlled trials remain fundamental to evidence generation in oncology but are increasingly complex, costly, and often misaligned with real-world practice. Traditional explanatory trials, designed under ideal, controlled conditions, frequently enroll highly selected populations, limiting generalizability and underrepresenting key groups such as older adults, patients with comorbidities, and those from low- and middle-income countries. Pragmatic clinical trials offer an alternative by evaluating interventions under routine care conditions, with broader eligibility, simplified procedures, and patient-centered outcomes. To address these challenges, the Gynecologic Cancer InterGroup convened an international brainstorming meeting in May 2025 with multi-disciplinary experts, patients, and advocates to define priorities and develop a roadmap for pragmatic trials in gynecologic oncology. Key discussions emphasized embedding trial design within routine care, aligning eligibility criteria and procedures with standard practice, minimizing non-essential data collection, and prioritizing outcomes meaningful to patients, including quality of life. Innovative designs such as registry-based randomized trials, trials-within-cohorts, and cluster randomization were highlighted as feasible approaches to improve efficiency while preserving internal validity. Integration of patient-reported outcomes and real-world data was considered achievable when carefully streamlined. Major challenges identified included regulatory heterogeneity, consent complexity, data interoperability, and funding limitations, particularly in multi-national settings. Proposed solutions include simplified consent models, centralized ethics processes, hybrid funding strategies, and the responsible use of artificial intelligence to enhance patient identification, recruitment, and potential development of synthetic control arms. Patient engagement was recognized as essential to ensure relevance, feasibility, and equity. Incorporation of patient-reported outcomes was discussed as key to informing acceptance and tolerability. In summary, pragmatic trials within Gynecologic Cancer InterGroup represent a critical pathway to generate efficient, inclusive, and practice-changing evidence in gynecologic cancers across diverse health care settings.

Humans

Project ODIN: advancing environmental genomic surveillance for public health across sub-Saharan Africa.

Persistent SARS-CoV-2 transmission, ongoing mpox outbreaks, and the continued spread of endemic diseases such as typhoid fever and cholera underscore the urgent need for global, multiomics surveillance. In this Personal View, we present Project ODIN, a consortium of European and African partners launched in 2023 that aims to meet this challenge by deploying innovative systems for near real-time pathogen detection and actionable public health insights. The project is a collaboration between high-income and low-income countries in northern Europe and sub-Saharan Africa. Focusing on low-income and middle-income countries, ODIN integrates metagenomics with mobile laboratory systems for comprehensive pathogen monitoring across diverse environments. ODIN emphasises standardised sampling, bioinformatics pipelines, and data-sharing protocols to ensure reliable, interoperable results while addressing infrastructure and resource limitations. By bridging gaps in genomic surveillance, these initiatives seek to strengthen outbreak preparedness, improve pathogen detection, monitor antimicrobial resistance, and provide a holistic approach to One Health challenges. Together, these innovations could advance global surveillance capacity-particularly in under-resourced regions-paving the way for effective disease control and evidence-based policy making.

Humans

Genetics of sensory nutrition.

Sensory nutrition is an emerging research area that examines how chemosensory perception, particularly taste and smell, shapes dietary behaviours, nutritional status, and disease risk. Variation in how individuals perceive the same foods may help explain differences in diet quality and responsiveness to behavioural dietary interventions, yet chemosensory phenotypes are rarely measured at the population level. Genetic variation contributes to this perceptual diversity and provides a framework for investigating sensory determinants of diet using genomic approaches. This review summarises evidence linking chemosensory genetics to perception and dietary behaviours, and discusses applications for causal inference and for precision and personalised nutrition. Twin studies reveal moderate to high heritability for bitter taste traits, with more modest and phenotype-dependent estimates for sweetness, sourness, saltiness, fat-related traits, and olfactory measures. Genome-wide association studies have identified loci in taste and olfactory receptor genes associated with specific chemosensory traits as well as liking and intake of various foods, although the evidence remains concentrated on bitter taste and populations of European ancestry. These genetic variants have been used in Mendelian randomisation, a genetics-based approach that strengthens causal inference, to test whether sensory traits influence dietary behaviour. For precision nutrition, evidence for taste genotype-stratified interventions remains limited and mixed. Realising the promise of sensory nutrition will require scalable and standardised chemosensory phenotyping, Findable, Accessible, Interoperable, and Reusable (FAIR) data infrastructure, expanded research in diverse populations, and integration with broader biological and sociocultural determinants of dietary intake.

Genetics

Giotto Suite: a multiscale and technology-agnostic spatial multiomics analysis ecosystem.

Emerging spatial multiomics technologies provide an increasingly large amount of information content at multiple scales. However, it remains challenging to efficiently represent and harmonize diverse spatial datasets. Here we present Giotto Suite, a suite of modular packages that provides scalable and extensible end-to-end solutions for multiscale and multiomic data analysis, integration and visualization. At its core, Giotto Suite is centered around an innovative data framework, allowing the representation and integration of spatial omics data in a technology-agnostic manner. Giotto Suite integrates molecular, morphology, spatial and annotated feature information to create a responsive and flexible workflow, as demonstrated by applications to several state-of-the-art spatial technologies. Furthermore, Giotto Suite builds upon interoperable interfaces and data structures that bridge the established fields of genomics and spatial data science in R, thereby enabling independent developers to create custom-engineered pipelines. As such, Giotto Suite creates an immersive and multiscale ecosystem for spatial multiomic data analysis.

Genomics

MelanoDB: A dataset of clinical and molecular features of patients with advanced melanoma treated with MAPK inhibitors.

MAPK inhibitors (MAPKi) have revolutionized the treatment of patients with advanced melanoma. However, primary and acquired resistance mechanisms limit their efficacy. Predicting MAPKi response from the tumor baseline features remains challenging due to the limited size of patient cohorts. Therefore, we collected data from nine different patient cohorts (total n = 417 patients with advanced melanoma treated with MAPKi) to identify clinical and molecular features. Our curated dataset, named MelanoDB, includes whole or partial exome sequencing data for 191 patients, copy number alteration information for 66 patients, and gene expression data for 132 patients. We provide a web application to explore the integrated dataset and data distribution across the collected studies, and we share this dataset with the scientific community according to the Findable, Accessible, Interoperable, Reusable (FAIR) principles.

Humans

CAUSAL artificial intelligence and data-driven decision intelligence in personalized medicine: a review of healthcare informatics systems.

This review examines the integration of causal artificial intelligence (AI) and data-driven decision intelligence within healthcare informatics systems to advance personalized medicine and clinical decision-making. A narrative review methodology was employed, synthesizing interdisciplinary literature from major databases, including PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect. Studies focusing on causal inference, decision intelligence, and healthcare informatics applications in personalized medicine were included. Data were extracted on methodological approaches, healthcare settings, analytical techniques, and clinical applications, followed by thematic synthesis. Findings indicate that causal AI enhances clinical decision support by enabling estimation of treatment effects and simulation of intervention outcomes at the individual patient level. Integration of multimodal health data such as electronic health records, genomic data, and real-time monitoring improves prediction accuracy and supports tailored treatment strategies. Additionally, causal models improve interpretability, fostering clinician trust and facilitating transparent decision-making. Robust healthcare informatics infrastructures, including interoperable systems and data warehouses, were identified as critical enablers of causal analytics. Overall, causal AI represents a transformative advancement in healthcare analytics, supporting more informed, individualized, and evidence-based clinical decisions. Its integration within healthcare informatics systems has significant potential to improve patient outcomes and guide the future of intelligent, personalized healthcare delivery.

Precision Medicine

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

The Lipid Interactome: an interactive and open access platform for exploring cellular lipid-protein interactions.

SUMMARY: Lipid-protein interactions play essential roles in cellular signaling and membrane dynamics, yet their systematic characterization has long been hindered by the inherent biochemical properties of lipids. Recent advances in functionalized lipid probes-equipped with photoactivatable crosslinkers, affinity handles, and photocleavable protecting groups-have enabled proteomics-based identification of lipid interacting proteins with unprecedented specificity and resolution. Despite the growing number of published lipid interactomes, there remains no centralized effort to harmonize, compare, or integrate these datasets. The Lipid Interactome addresses this gap by providing a structured, interactive web portal that adheres to FAIR data principles-ensuring that lipid interactome studies are Findable, Accessible, Interoperable, and Reusable. Through standardized data formatting, interactive visualizations, and direct cross-study comparisons, this resource enables researchers to systematically explore the protein-binding partners of diverse bioactive lipids. By consolidating and curating lipid interactome proteomics data from multiple studies, the Lipid Interactome database serves as a critical tool for deciphering the biological functions of lipids in cellularsystems. AVAILABILITY AND IMPLEMENTATION: This site can be viewed at LipidInteractome.org. All data are available for download. No user information is collected or necessary for data navigation, interaction, or download.

Proteins

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

Atlas-level single-cell integration and clustering-free differential expression analysis with GEDI 2.0.

MOTIVATION: GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS: We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98× mean single-threaded speedups, and up to 11.5× acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.

Single-Cell Analysis

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

Conference report: the third Bacterial Genome Sequencing Pan-European Network conference.

The third Bacterial Genome Sequencing Pan-European Network conference, held in Engelberg, Switzerland (12-15 January 2026), brought together experts from six European countries to discuss the implementation of bacterial genome sequencing in clinical microbiology and public health. Key themes included regulatory frameworks (In Vitro Diagnostic Regulation, General Data Protection Regulation), standardization, quality control, data sharing, economic evaluation, and the integration of artificial intelligence and long-read sequencing into diagnostic workflows. Across presentations, panel discussions, and workshops, participants emphasized that successful implementation of genome sequencing requires more than technical capacity: it depends on robust validation, sustainable funding, interoperable data standards, ethical governance, and interdisciplinary collaboration. The meeting highlighted that sequencing should remain question-driven and clinically meaningful, balancing cost, turnaround time, and public health impact. Overall, the conference reinforced the need for coordinated European efforts to advance responsible, standardized, and sustainable genomic surveillance and diagnostics.

bacterial genome sequencing

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

KG-Microbe: Building modular and scalable knowledge graphs for microbiome and microbial sciences.

BACKGROUND: The integration of many disparate forms of data is essential for understanding the microbial world and its interaction with the environment and human health. Doing so is particularly challenging in the context of microbe-host and microbe-microbe interactions that contribute to health or environmental outcomes. There are thousands of relevant microbial species, and millions of interactions among those microbes and with their environment or host. Integrated information (e.g., about host and microbial physiology, genetics, and metabolism) facilitates deeper understanding of complex mechanisms and helps interpret correlative results. RESULTS: The KG-Microbe construction framework is a novel approach to harmonizing bacterial and archaeal data in the form of a findable, accessible, interoperable, reusable and AI-ready knowledge graph (KG). Starting from a core KG with organismal traits, environments, and growth preferences and the integration of established ontologies, the framework generates a hierarchy of related KGs targeting specific use cases, including the human microbiome in the context of disease, or environmental microbiomes. The framework supports customizable taxa subsets representing communities or clades of interest. Evaluations of the KG-Microbe KGs through a series of competency questions demonstrate the accuracy and effectiveness of the data harmonization, and the utility of the resulting KGs in studies of inflammatory bowel disease and Parkinson's disease. Finally, the predictive and environmental capabilities of the KGs are demonstrated by predicting growth preferences using graph features. CONCLUSIONS: The KG-Microbe framework unifies microbial contexts in a single resource to support integrative analyses across biomedical, host, and environmental domains. KG-Microbe is a flexible, modular enabling technology for humans and machine learning methods to uncover candidate mechanistic explanations of microbial associations.

Microbiota

Modular synthetic cross-kingdom promoters enable coordinated expression in Escherichia coli and Saccharomyces cerevisiae.

Synthetic biology and metabolic engineering increasingly demand predictable and interoperable gene expression across phylogenetically distant organisms, as the need for portable genetic systems and transferable metabolic pathways continues to grow. However, fundamental differences in promoter architecture and transcriptional logic across kingdoms remain a key bottleneck in developing universal expression platforms. Here, we designed a set of modular hybrid promoters that enable tunable and quantitatively consistent gene expression in both Escherichia coli and Saccharomyces cerevisiae. These promoters integrate bacterial -10/-35 motifs and Shine-Dalgarno sequences with minimal yeast TATA boxes and Kozak sequences to ensure transcriptional and translational compatibility. The promoter set supported weak, moderate, and strong expression with high relative consistency across species. Applied to the biosynthetic pathway for the valuable pigment prodeoxyviolacein, the hybrid promoters enabled coordinated production in both hosts. This work establishes a broadly compatible promoter architecture and provides a foundational toolkit for cross-kingdom, multi-host synthetic biology.

Promoter Regions, Genetic

A general strategy for generating expert-guided, simplified views of ontologies.

Annotation of biomedical entities with widely used, well-structured ontologies and ontology-aware tools ensures data and analyses are Findable, Accessible, Interoperable, and Reusable (FAIR). Standardized terms with synonyms support lexical search, while ontology structure enables biologically meaningful grouping of annotations, such as by location and type. However, ontologies serving diverse communities are often more complex than needed for specific applications, creating barriers to adoption by researchers and resource developers. For example, cell atlases often attempt simplifications by manually building term hierarchies linking to cell type and anatomy ontologies, but these may include relationship types unsuitable for grouping annotations. We present tools for validating human expert curated term hierarchies, developed in two human reference atlas projects, against ontology structures. The tools provide tabular statistics plus graphical views of matching and non-matching terms and relationships to support discussion and conflict resolution. The HuBMAP Human Reference Atlas (HRA) effort is used to validate the approach and tools, and the Human Developmental Cell Atlas is featured as a use case.

Journal Article