PubMed HealthSearch

SEARCH · PubMed Health

Results for “data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

DIVAS: an R package for identifying shared and individual variations of multiomics data.

MOTIVATION: Multiomics data integration aims to identify biological patterns shared across molecular modalities. Most existing methods detect either jointly shared variation, across all modalities, or individual variation, unique to a single modality, but overlook partially shared variation, shared by only a subset of modalities. This is a critical limitation, because many biological mechanisms manifest in some but not all molecular modalities. RESULTS: We present an open-source R package implementing data integration via analysis of subspaces (DIVAS), a framework for systematically identifying jointly shared, partially shared and individual variations across multiple data types. DIVAS combines angle-based subspace analysis with inference through rotational bootstrap, hierarchically searching all combinations of modalities to decompose multiomics data into interpretable components with scores and loadings. In simulations with a known sharing structure, DIVAS recovered every component across a wide range of noise levels, whereas existing methods did not. Applied to multi-modal COVID-19 data, it reveals partially shared immune and metabolic dysregulation patterns underpinning disease severity that conventional approaches would miss. AVAILABILITY AND IMPLEMENTATION: DIVAS is available at https://github.com/ByronSyun/DIVAS, with documentation and vignettes. The COVID-19 case study vignette is available at https://byronsyun.github.io/DIVAS_COVID19_CaseStudy/.

Multiomics

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n = 91), with fewer studies examining samples from individuals of advanced reproductive age (n = 45) and fetal (n = 16) samples. Transcriptome analyses were most common (n = 103, 85%), followed by proteome (n = 19, 16%) and epigenome (n = 14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female

A comprehensive integration of data on the association of ITPKC polymorphisms with susceptibility to Kawasaki disease: a meta-analysis.

BACKGROUND: This study aims to conduct a comprehensive meta-analysis of existing research to define clear associations between variations in the ITPKC gene and the risk of developing Kawasaki disease (KD). METHODS: A comprehensive search was conducted across multiple databases, including but not limited to PubMed, Scopus, EMBASE, and CNKI, up to June 1, 2024, to gather relevant information. This search utilized keywords and MeSH terms related to hyperbilirubinemia and genetic factors. The inclusion criteria encompassed original case-control, longitudinal, or cohort studies. Correlations were analyzed as odds ratios (ORs) with 95% confidence intervals (CIs) using Comprehensive Meta-Analysis software. RESULTS: Eighteen case-control studies with 5,434 KD cases and 9,419 controls were analyzed. Of these, ten studies assessed 3,129 KD cases and 6,172 controls for the rs28493229 variant, four examined 1,039 cases and 1,688 controls for the rs2290692 variant, two focused on 595 cases and 820 controls for the rs7251246 variant, and two investigated 671 cases and 739 controls for the rs10420685 variant. Results showed a significant association between the rs28493229 polymorphism and increased KD risk across all five genetic models. Subgroup analysis indicated this polymorphism correlates with KD susceptibility in Asians but not in the Chinese population. In contrast, no associations were found between the rs2290692, rs7251246, and rs10420685 polymorphisms and KD risk. CONCLUSIONS: Our pooled data indicate a significant association between the ITPKC rs28493229 polymorphism's minor allele and an increased risk of developing KD, suggesting this variant may enhance susceptibility. Conversely, SNPs rs2290692, rs7251246, and rs10420685 do not demonstrate a statistically significant relationship with KD.

Humans

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products

Multi-omics technologies: Novel tools and methods for assessing nerve injury and regeneration.

Recently, with the rapid advancement of multi-omics technologies, including genomics, transcriptomics, proteomics, and metabolomics, new tools and approaches have been introduced for studying nerve injury and regeneration. This review highlights the application and progress of multi-omics in uncovering the mechanisms of nerve injury, guiding the development of regenerative strategies, and promoting clinical translation. By integrating multi-omics datasets, researchers can comprehensively track dynamic molecular changes following nerve injury, including abnormal gene expression, disrupted protein signaling, altered metabolic programs, and shifts in the immune microenvironment. Single-cell multi-omics technologies resolve cellular heterogeneity, revealing the distinct functions of neurons, glial cells, and immune cell subpopulations during the injury response. Spatially resolved transcriptomics maintain the spatial context of lesion and regeneration sites, enabling precise localization for targeted interventions. Multi-omics technologies not only identify key molecular players involved in nerve regeneration but also create opportunities for personalized medicine. Nonetheless, integrating multi-omics data poses technical challenges, including high dimensionality, batch effects, and algorithmic constraints, while ethical concerns related to stem cell therapy and gene editing require stringent oversight. To transition from structural reconstruction to functional remodeling, future research should emphasize artificial intelligence-driven data integration, organ-on-a-chip modeling, and cross-disciplinary collaboration to overcome existing technical barriers and accelerate the clinical application of neuroregenerative therapies.

artificial intelligence

An automated clinic management system for a family planning network.

The medical information, financial, and logistic aspects of a comprehensive computer-based Appointment, Registration, Information System, and Evaluation (ARISE) are analyzed for the management of a family planning program serving 30,000 patients annually. An overview of the existing computer system network is presented with descriptions of the interactive master patient index, the batch appointment process, the management statistics package, and Department of Health, Education, and Welfare (HEW) reporting. Emphasis is placed on the financial management control system which includes 1) procedures for third-party submission of claims for payment, in particular Titles IVA, XX, and XIX (Social Security Act), together with discussion of related administrative requirements; 2) technics of auditing data integrity including systematic sampling of collected data; and 3) the process of billing and receipts collection. Methodology and implementation aspects of ARISE may have wide applicability to other family planning and similarly structured clinical programs.

Computers

A data structure model for a health information system.

The Ministry of Health and Environmental Control of Berlin is developing a Health Information System (HIS) on the basis of a multi-satellite network system, comprising a central, regional, local and functional unit level. The main part of this paper describes the conceptual structure of the Common Data Base (CDB) of HIS with special regard to the patient-oriented medical information originating from the various institutions of the health care system. This structure comprises the following five levels: 1. PATIENT 2. PROBLEM 3. CASE 4. EVENT 5. ACT Each of the levels represents a node in the structure model. A node is an entity with a set of "local properties" being specified for each level, referring to selected data on inferior levels. The structures of these five levels are described in detail. In the last part, so-called "data-manipulation procedures" are treated. These are descriptions covering any data manipulation and represent the basis of data integrity through system controlled transaction with the data base.

Berlin

Monitoring adverse drug reactions-the problem of integration of heterogeneous data.

Effective systems for a meaningful integration and interpretation of data of heterogeneous origin are essential for successful monitoring of adverse drug reactions internationally or in multicenter programs. Particular difficulties are encountered in the area of suspected drug adverse reactions and in the area of drugs. Both problem areas are discussed in the paper.

Drug-Related Side Effects and Adverse Reactions

Quality control methods for data entry in pathology using a computerized data management system based on an extended data dictionary.

In pathology, computerized data management systems have been used increasingly to facilitate a more efficient supply of information. Since data entry precedes data utilization, the reliability of the information stored strongly depends on the quality of data input. Despite its potential capability, most personal computer-based database software does not provide versatile and user-friendly data validation procedures. Therefore, we developed a data dictionary-driven data management system that enables the user to perform extensive validation routines without the need for hard programming. Using examples from an existing database for endometrial carcinomas, different types of data errors and their error traps are explained. It is pointed out that data type definitions, defaults, templates, or picture clauses are suitable means to avoid formal errors. Validations on data domains and ranges test whether data fall into a predefined scope. Relational checks control data validity within a context of different data items, whereas process routines provide automatic data computation, thereby circumventing user input. By exploiting the facilities of an extended data dictionary, a powerful tool is made available to secure various aspects of data integrity simultaneously with input. In this way, computerized data quality control can improve the efficiency and reliability of data management tasks in pathology.

Medical Informatics Computing

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics

SESAM: a relational database for structure and sequence of macromolecules.

A system is described that provides ways of integrating data on protein structure, sequence, and survey results, with molecular graphics and molecular mechanics software. Its major component is the relational database SESAM, presently implemented under the commercial package SYBASE. By design, the database allows full integration--within the same data organization--of raw data on protein structure, sequence, ligands, and heterogroups, obtained from the Brookhaven Protein Databank, with pure sequence information available from other databanks such as SWISS-PROT. It contains in addition higher level descriptions of structural and topological properties, as well as survey results, obtained by executing specialized computer programs. Aside from the very useful attribute of closely combining structural and nonstructural information, other important features distinguish it from analogous systems developed elsewhere. It includes a molecular dictionary with complete description of geometric properties and energy parameters used in modeling and conformational energy calculations. Using this dictionary, structural data are validated by checking for localized inconsistencies in atomic coordinates, atomic symbols, chirality definitions, and flagging errors and incomplete entries. Because of both the dictionary and the validation procedures, SESAM can be readily interfaced with conventional molecular graphics and mechanics software packages, or with other specialized application programs. With the aid of appropriate interfaces, data access is sufficiently fast for SESAM to be interrogated interactively. Prototypes of user interfaces, as well as an interface with the molecular graphics package BRUGEL, are described and the power of the system is illustrated in applications such as homology-based protein modeling, computer-aided protein design, protein structure predictions, analysis of local structure motifs, and of relationships between protein sequence and structure.

Amino Acid Sequence

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans

An integrative network approach for longitudinal stratification in Parkinson's disease.

Parkinson's disease (PD) is a neurodegenerative disorder characterized by motor symptoms resulting from the loss of dopamine-producing neurons in the brain. Currently, there is no cure for the disease which is in part due to the heterogeneity in patient symptoms, trajectories and manifestations. There is a known genetic component of PD and genomic datasets have helped to uncover some aspects of the disease. Understanding the longitudinal variability of PD is essential as it has been theorised that there are different triggers and underlying disease mechanisms at different points during disease progression. In this paper, we perform longitudinal and cross-sectional experiments to identify which data modalities or combinations of modalities are informative at different time points. We use clinical, genomic, and proteomic data from the Parkinson's Progression Markers Initiative. We validate the importance of flexible data integration by highlighting the varying combinations of data modalities for optimal stratification at different disease stages in idiopathic PD. We show there is a shared signal in the DNAm signatures of participants with a mutation in a causal gene of PD and participants with idiopathic PD. We also show that integration of SNPs and DNAm data modalities has potential for use as an early diagnostic tool for individuals with a genetic cause of PD.

Parkinson Disease

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans

Unlocking the Full Potential of Spatial Omics in Plants: Practical Challenges, Solutions, and a Path Forward.

Spatial omics technologies are providing new opportunities for plant biology by enabling molecular profiling within structurally intact tissues, revealing spatially organised cell states, developmental gradients, and regulatory interactions. While spatial transcriptomics has driven early advances, the field is rapidly expanding toward integrated spatial multi-omics by combining single-cell and spatial transcriptomic, epigenomic, proteomic, and metabolomic data. These approaches offer new opportunities to study development, physiology, and plant biotic and abiotic interactions in spatially preserved cellular contexts. However, despite rapid adoption, the field remains constrained by plant-specific challenges when applying technologies largely developed for animal systems. Compared with animal systems, plant tissues pose additional challenges due to rigid cell walls, and diverse chemistries, complicating sample preparation, cell and subcellular segmentation, signal detection, and data integration. As a result, many studies rely on bespoke protocols and analysis pipelines that are often difficult to reproduce or generalise. Here, we provide a practical, solution-oriented synthesis of current bottlenecks across experimental and computational pipelines, highlight emerging strategies to overcome these limitations, and propose a roadmap for community-driven protocol sharing, benchmarking, and integration across spatial and multi-omics modalities. Addressing these challenges will be essential to establish spatial omics as a routine and scalable tool for plant biology.

Journal Article

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans