PubMed HealthSearch

SEARCH · PubMed Health

Results for “data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Privacy-Preserving Linkage of Distributed Biological, Clinical, and Imaging Data Supporting Artificial Intelligence in Pediatric Oncology.

BACKGROUND: Cancer remains the leading cause of disease-related mortality in children over the age of one in Europe, with over 35,000 new pediatric cases and more than 6,000 deaths annually. Due to the rarity of pediatric cancers, clinical trial protocols often substitute for formal treatment guidelines, resulting in many children being enrolled in multiple trials, with biological samples and genomic data stored in various biobanks. Data collection in pediatric oncology is challenging, with sparse data acquired over extended periods, underscoring the need for optimal utilization of all available information through linked, privacy-preserving datasets. METHODS: Here, we report the development of a distributed, privacy-preserving data infrastructure for the PRIMAGE project, a European initiative aimed at supporting artificial intelligence (AI)-driven image analysis for pediatric cancer prognostics. The infrastructure leverages the European Patient Identity (EUPID) Services for Privacy-Preserving Record Linkage, enabling pseudonymized data integration across clinical, biological, and imaging sources. The system incorporates EUPID's hashing and phonetic matching protocols to pseudonymize patient identifiers and link distributed datasets, facilitating secondary data use in compliance with the General Data Protection Regulation. RESULTS: Data from over 700 neuroblastoma patients from European trials and hospitals were linked and uploaded to the PRIMAGE platform, where AI models predict clinical outcomes. CONCLUSION: This infrastructure successfully facilitated AI model development, advancing pediatric oncology research, and offering a scalable framework for future European health data initiatives, such as the European Health Data Space.

Journal Article

Bridging genotype, phenotype, and clinical insight: the role of multi-omics in cardiovascular disease.

INTRODUCTION: It is increasingly evident that the multifactorial nature of cardiovascular disease requires the combination of different omics approaches for improving our mechanistic understanding, identifying novel drug targets, and developing accurate diagnostic, predictive, and prognostic biomarker panels. AREAS COVERED: We review the current state and the potential of multi-omics in cardiovascular disease, with a specific focus on plasma-, spatial-, and single-cell approaches. We discuss lipidomics as a genotype‑to‑phenotype bridge, the utility of remote longitudinal monitoring via microsampling/dried blood spots, and emerging clinical‑trial integrations of multi-omics approaches. We outline critical gaps in standardization and how to overcome these, pre‑analytical challenges and constraints that are often neglected, and data‑integration methods spanning from canonical correlation analysis to modern machine learning approaches. EXPERT OPINION: Multi‑omics can shape cardiovascular care by identifying drug targets in diseased tissue and by yielding small, usable biomarker panels.

Humans

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

The conflict between relational databases and the hierarchical structure of clinical trials data.

Relational database software has become popular for the management of certain types of commercial data. Its use is being given serious consideration in the management of data from clinical trials. Relational systems have a number of advantages over hierarchical or network systems for some types of data. However, as illustrated by an example, the data from clinical trials typically have an inherent hierarchical structure. The incorporation of hierarchically structured data into a relational database raises difficult problems of data integrity versus the complexity of the database structure. These problems, together with the long execution times of many relational operations, indicate that relational systems are not necessarily well suited for clinical trials data management.

Clinical Trials as Topic

The significance of teeth in pollution detection.

The general population is experiencing lifelong exposure to old and new hazardous substances. By using data collected from a subject's own teeth, accuracy in determining the effects of exposure is assured since extrapolation is excluded. The establishment of a common tooth bank can provide means to integrate data from multiple sources. Comprehensive pollution information shared by the environmental, scientific and medical communities can lead to a more efficient approach to a worldwide problem.

Animals

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms

Human Systems Immunology in the Omics Era: Challenges, Methods, and Emerging Directions.

The human immune system is a highly complex, dynamic, and heterogeneous network shaped by genetic, environmental, and temporal influences. Advances in high-throughput omics technologies have transformed our ability to study this complexity directly and comprehensively in human cohorts. These developments have positioned systems immunology as a powerful framework for investigating coordinated immune responses, identifying regulatory mechanisms, and linking molecular patterns to clinical phenotypes. However, the analytical challenges inherent to large-scale, multimodal datasets-including batch effects, small sample sizes, high dimensionality, and substantial interindividual heterogeneity-require rigorous study design, robust statistical modeling, and thoughtful data analysis strategies. In this review, we summarize key technological foundations enabling modern human systems immunology, outline common analytical pitfalls and effective mitigation approaches, discuss data integration concepts, and highlight emerging opportunities in the field. Together, these technological and analytical advances are redefining how immune function is measured and interpreted in real-world human biology and hold significant promise for enhancing mechanistic insight, biomarker discovery, and precision medicine across immunological diseases and interventions.

Humans

Automated data processing for chromatographic assay method validations.

A set of computer programs has been developed for the automated transfer, storage, and processing of data related to chromatographic pharmaceutical assay method validations. These programs calculate and report statistical parameters relating to the resolution, linearity of response, and precision of an assay method based upon peak width, retention time, and integration data provided by a commercial chromatographic data system. Utilization of an interface between the data system and a microcomputer minimizes manual handling of all data. Techniques and equations used in the development of the software are described, and an application to the validation of a high performance liquid chromatographic assay method for a bulk drug substance is demonstrated.

Chromatography, High Pressure Liquid

[Transmission of Dusseldorf Integrated Medical Documentation data by a personal computer. A new concept in the flexibility of the Dusseldorf Integrated Medical Documentation system].

The software package HOST PC was created as an important expansion of the oncological aftercare program INMEDD (Integrated Medical Documentation Düsseldorf). INMEDD itself provides only small evaluation capabilities. HOST PC is able to transmit data from INMEDD on a host computer to a personal computer. This transmission is fully automated. On the personal computer the data is stored in a database, which is completely compatible with the standards of dBase III. These created databases can be evaluated and analyzed by a lot of standard software packages. Therefore a wide range of individual statistical evaluation, analyzing with optional criteria and graphic presentation can be realized with small expenses of time and money. HOST PC increases the attractiveness of INMEDD and therefore improves the aftercare of cancer patients.

Aftercare

Normal sleep: patterns and mechanisms.

In the last three decades, research in the sleep laboratory has decisively contributed to a much deeper knowledge of sleep physiologic and pathologic states. Parallel to clinical research related to sleep disorders, multifaceted basic research has greatly contributed to a better understanding of the mechanisms underlying sleep and wakefulness. This basic research in the realm of the neurosciences integrates data derived from the application of various methodologic approaches. Currently, the prevailing concepts about sleep mechanisms generally favor the idea of a dynamic interaction among systems rather than that of a unidimensional explanation for sleep generation. Examples of these integrative concepts in current sleep research are the revised model of reciprocal interaction for the control of REM sleep and the two-process model comprising the seemingly incompatible homeostatic and circadian sleep mechanisms. In sleep disorders medicine, the prevailing approach is also integrative: findings from the sleep laboratory are considered in conjunction with those from clinical experience in understanding the nature of sleep disorders. Applying this integrative model, physicians in sleep disorders medicine are able to manage patients with sleep disorders comprehensively. Based on an understanding of sleep physiology, clinicians can make the diagnosis of most sleep disorders in the office setting.

Adolescent

Genetic architecture of endometriosis: risk factors, comorbidities and clinical implications.

BACKGROUND: In 1999, Dr Susan Treloar and colleagues conducted a landmark twin study in Australia and reported their estimate of 51% for the heritability of endometriosis. This important result led several groups to begin mapping genetic factors contributing to increased endometriosis risk. Despite early challenges, advances in genome-wide association studies (GWAS) have identified multiple genetic risk factors and some target genes implicated in follow-up studies on genetic regulation of transcription. Access to large publicly available genetic datasets and analysis with endometriosis GWAS results is also providing new opportunities to answer important questions about comorbid conditions associated with endometriosis and their implications for clinical practice. OBJECTIVE AND RATIONALE: The objective of the review is to summarize the last 25 years of genetic studies in endometriosis, outline contributions to our understanding of the disease, and suggest future directions to accelerate biological insights from genetic studies to improve clinical outcomes. SEARCH METHODS: A comprehensive review of scientific literature on the genetics of endometriosis was conducted through searches in PubMed and Google Scholar up to June 2026. Search terms included "endometriosis AND (genetics OR GWAS OR genetic risk factors)", For studies addressing the functional characterization of genetic risk loci, additional searches employed the terms "endometriosis AND (genotype-phenotype associations OR colocalization OR eQTL OR mQTL OR multi omics methods)". To identify studies examining shared genetic risk between endometriosis and comorbid conditions, the search strategy included "endometriosis AND (genetic correlation OR colocalization OR Mendelian randomisation)". Publications reporting discoveries related to genetic risk factors for endometriosis and studies interpreting their biological and clinical significance were critically evaluated, and 144 publications were discussed in the review. OUTCOMES: Discovery of genetic risk factors started slowly and has accelerated in recent years with developments in technology and international collaborations to combine data and increase statistical power. GWAS have mapped 80 genetic risk factors that implicate gene regulation of hormonal targets, development of the reproductive tract, regulation of cell proliferation, and regulation of epithelial cell differentiation. In common with most other complex diseases, effects of individual common genetic risk factors are small. However, several examples demonstrate that small effect sizes are not a good predictor for the impact of drugs developed against genetically validated targets. Genetic risk factors implicate five genes regulating gonadotrophin release and oestrogen action, the major target pathway of current drugs for treatment of endometriosis demonstrating proof-of-principal for biologically meaningful results. Genetic correlation and Mendelian Randomization studies highlight important causal relationships between endometriosis and comorbid conditions including a possible role for testosterone during development and shared genetic risk factors for gynaecological, gastrointestinal, pain, psychiatric, and inflammatory conditions. Understanding causal relationships between endometriosis and related conditions will aid clinical management and more personalized treatments. WIDER IMPLICATIONS: Genetic studies provide novel insights into endometriosis pathogenesis and associations with related comorbid conditions. Genetic factors modifying gene regulation and disease risk likely act in specific cell types, and access to datasets from genetically informed cell-based models, single-cell and spatial omics data are needed to accelerate progress. Future studies should address critical questions of heterogeneity and disease subtypes, expand the search for genetic risk factors to non-European populations, evaluate the role of rare and structural variants, and better integrate data from functional, genomics, genetics, and clinical studies to reduce diagnostic delay, develop novel treatment strategies, and translate discoveries into personalized management strategies for affected individuals. REGISTRATION NUMBER: N/A.

comorbid conditions

Work experiences of minority managers and professionals: individual and organizational costs of perceived bias.

The present study examined the relations of the way minority managers and professionals described their treatment within their organizations, and their organizations' acceptance and openness to minorities within measures of satisfaction, commitment, skill utilization, and integration. Data were collected from 81 minority managers and professionals in early career stages using questionnaires completed anonymously. Minority managers experiencing more positive treatment in their organization, and employed in organizations more accepting of minorities, were more satisfied, committed, and integrated.

Achievement

An optically scanned EMS reporting form and analysis system for statewide use: development and five years' experience.

Analysis of emergency medical services (EMS) systems data is crucial to planning, education, research, and quality assurance programs. Currently, comparative analysis of EMS data between regions or states is virtually impossible due to wide variations in data collection and analysis methods. To devise a practical and uniform EMS reporting system, we referenced the minimum data set (MDS) established by the federal government in 1974 and surveyed 22 states known to be using uniform reporting systems. In developing our final data set, elements were added based on inclusion in the MDS, national survey results, a review of current EMS literature, and consensus of local EMS providers. This set of 48 elements then was incorporated into a reporting form using narrative and optically scanned formats, allowing automated data collection for computer analysis. After a pilot study, the system was improved to allow high-speed ink reading and large volume data storage and analysis using a microcomputer. This system has subsequently been adopted by seven states. The combined data base exceeds 250,000 cases. Error screening algorithms ensure data integrity and are also used for quality assurance. Customized output reports can be generated within minutes and have assisted in EMS quality assurance, planning, and research. We believe that the successful performance of this system supports the use of the suggested data elements as well as optical scanning and microcomputer analysis of EMS data.

Data Collection

A multiuser system for whole body plethysmographic measurements and interpretation.

A multiuser system for whole body plethysmographic measurements and interpretation which has been developed under clinical conditions is described. The following measurements can be carried out in a rapid way and in one session with the patient: specific airway resistance during spontaneous breathing, determination of functional residual capacity, static lung volumes, and maximal forced expiratory data. Each section is normally measured twice and can be repeated up to ten times. The final results are displayed and printed together with a consistent system of normal reference values. All values and selected original curves are stored automatically in an integrated data base system. Obstructive patients are measured again after the inhalation of a bronchodilator. All results are evaluated by an automatic interpretation program. This analyzes and graduates airway obstruction, lung volumes, and pharmacological airway reversibility using standardized texts which are written below all numerical printouts and graphical plots. The interpretation algorithm is tree structured and uses the normal reference values as a knowledge base. The system supports up to four online laboratories with their own A/D converter and up to 20 video terminals, printers, plotters, and modems. Our laboratory performs 8,297 such complete measurements on 4,671 different patients per year with one body box.

Diagnosis, Computer-Assisted

Modification of the OMED nomenclature: a system approach based on the SISCOPE data model.

The OMED nomenclature represented a turning point in endoscopic computer systems by supplying software developers with an internationally recognized scientific document on which prototypes could be based. The main pitfalls of the OMED system are related to its hierarchical structure, probably not the most effective design to represent endoscopic findings. Based on our experience during the development of SISCOPE, an integrated data management system for endoscopy, an alternative scheme is proposed: Endoscopic descriptions are modeled as a set of objects represented by a data structure whose elements are location, morphology, associated lesions and hemorrhage. 72 objects appear to be sufficient for an accurate representation of all endoscopic scenes and a consistent data model could be created with this approach. Efforts should be made to decrease redundancy in the OMED nomenclature, but extension to other endoscopic data types, such as clinical and pathological diagnosis, is more urgently required. Furthermore, if data exchange between systems is desired, the definition of an Endoscopy Metafile is an absolute requirement.

Database Management Systems

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics

Trans-omics integration underscores distinct roles of polyunsaturated phospholipids in bidirectional offspring birth weight deviations.

BACKGROUND: Abnormal birth weights are associated with adverse pregnancy outcomes and future metabolic consequences. We aimed to examine cord blood lipidomes from low, normal and high birth weight (LBW, NBW, HBW) infants to identify core lipid signatures associated with non-optimum birth weight, and to derive biological insights through trans-omics data integration with placental proteome, maternal plasma lipidome and clinical phenome. METHODS: We conducted quantitative lipidomics of cord blood samples from two independent cohorts: a retrospective discovery cohort (n = 147) and a prospective validation cohort (n = 73). Integration with placental proteomics, maternal plasma lipidomics and clinical phenomics was conducted to elucidate potential biological implications. FINDINGS: We identified substantial reductions in cord blood polyunsaturated phospholipids (PUFA-PLs) (FDR <0.05) associated with placental vesicle trafficking and formation in LBW, and altered neutrophil degranulation in HBW. Combinatorial analyses of paired maternal plasma and cord blood samples indicated that cord blood PUFA-PL reductions were not attributable to deficient maternal supply, but rather to impeded assimilation (LBW) and increased utilisation (HBW). INTERPRETATION: Our findings provide biological insights that may inform targetable, lipid-oriented nutritional and/or pharmacological strategies to modulate foetal growth and development, with the goal of optimising clinical outcomes for both mother and child. FUNDING: This work was supported by the National Natural Science Foundation of China (82170854, 81870579, 81870545, 82571043, 2357308); National High Level Hospital Clinical Research Funding (2022-PUMCH-C-019); Noncommunicable Chronic Diseases-National Science and Technology Major Project (2024ZD0530200 and 2024ZD0530204); Beijing Municipal Science & Technology Commission (Z201100005520011); Peking University Clinical Scientist Training Program (No. BMU2023PYJH022); Beijing Municipal Natural Science Foundation (7202163, 7184252).

Humans

Automated data collection and presentation in the operating room.

An 'Operating Room Data Integration System', is described which is used to collect, present and archive all important physiological parameters during open heart surgery. The system requires very little attention, and provides an easy to understand and coherent interface to the user. The system is adaptable to a large extend and thus data can be presented to the user in a manner, with which he or she is already familiar. Simple drivers can be written to enable connection of the system to almost any other piece of medical equipment, if the latter provides an analog or digital, output signal. Automatic logging of the acquired signals is then possible.

Computer Systems