PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Oncopacket: integration of cancer research data using GA4GH phenopackets.

SUMMARY: Lack of data integration remains a significant impediment to cancer research, and many analyses still require customized software to transform and prepare cancer data. We describe a software package to harmonize genetic and clinical cancer data into the GA4GH Phenopacket schema, an ISO standard for representing clinical case data. We integrated demographic, mutation, morphology, diagnosis, intervention, and survival data using case data from the National Cancer Institute for 12 cancer types. The Phenopacket standard provides a foundation for downstream use, including sophisticated statistical and AI/ML analyses. We demonstrate fitness for purpose by using the integrated data to recapitulate a known association between mutations in the gene encoding isocitrate dehydrogenase 1 and survival time in brain cancer patients. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at: https://github.com/monarch-initiative/oncopacket (archived at 10.5281/zenodo.15353125).

Humans

JASMINE: A powerful representation learning method for enhanced analysis of incomplete multi-omics data.

Integrative analysis of multi-omics data provides a more comprehensive and nuanced view of a subject's biological state. However, high-dimensionality and ubiquitous modality missingness present significant analytical challenges. Existing methods for incomplete multi-omics data are scarce, do not fully leverage both modality-specific and shared information, and produce task-biased representations. We propose JASMINE, a self-supervised representation learning method for incomplete multi-omics data that preserves both modality-specific and joint information and enhances sample similarity structure. JASMINE produces embeddings that achieve superior performance across multiple tasks for two different incomplete multi-omics datasets while requiring only a single round of training per dataset.

missing data

Investing in Canada's nursing workforce: a comprehensive review to inform policy innovations and directions.

BACKGROUND: Health systems worldwide face persistent health workers challenges including nursing shortages, workforce strain, and inequities. In Canada, these challenges have prompted renewed national and provincial reforms to strengthen recruitment, retention, leadership, and sustainability. This paper compares nursing workforce policy directions across Canada, and international jurisdictions to inform policy and planning. METHODS: A cross-country comparative analysis of policies building on a comprehensive national funded review that included an umbrella review of 69 systematic reviews, a comparative policy review of nursing workforce strategies in five jurisdictions, and validation through national horizon-scanning and policy dialogues (n >100). Evidence was analyzed across system, organizational, and individual levels. RESULTS: At the system level, international jurisdictions demonstrate comprehensive, legislated approaches integrating data, governance, and multi-year funding have advanced key nursing strategies. In Canada, the advances show the importance of strategies to have national and provincial/territorial alignment emphasizing leadership, flexibility, and inclusion as key levers. Organizational and individual-level reforms such as mentorship, leadership development, and wellness initiatives are expanding but remain variably evaluated. Experts identified national workforce data strategies and policy integration with embedded evaluation as key enablers to inform scalability and sustainability of implemented strategies. CONCLUSIONS: Canada's nursing workforce reforms are advancing toward coordinated, equity-driven, and evidence-informed strategies. Continued investment in evaluation, leadership, and national integrated data systems along with integrating nursing workforce planning within broader intersectoral planning will consolidate these gains and position Canada as an international leader in sustainable nursing workforce policy.

Canada

AI-Based 3D Heterogeneous Network Model for Functional Prediction of Epigenetics.

Human biology and diseases are the result of constantly evolving processes within an intricately complex molecular network of interactions, such as epigenetic regulation. Epigenetics refers to heritable changes in gene expression that occur without alterations to the underlying DNA sequence. These changes, driven by mechanisms such as DNA methylation, histone modifications, and noncoding RNAs, play critical roles in regulating chromatin structure and gene activity. Epigenetic regulation offers valuable insights into biological systems, and when integrated with sophisticated analyses, it enables us to gain insights into gene regulation and cellular behavior. Here, we describe an artificial intelligence (AI)-based model that is capable of generating 3-dimensional (3D) heterogeneous network by integrating multimodal data for the functional prediction of epigenetic mechanisms, emphasizing its applications in medicine, developmental biology, and personalized therapeutics. Heterogeneous networks in biology are powerful tools for understanding the complex interactions and interdependencies within biological systems. Key advancements in AI and multiomics data integration have propelled this field, offering new insights into disease mechanisms, biomarker discovery, and therapeutic interventions.

Epigenesis, Genetic

scMGCL: accurate and efficient integration representation of single-cell multi-omics data.

MOTIVATION: Single-cell multi-omics data integration is essential for understanding cellular states and disease mechanisms, yet integrating heterogeneous data modalities remains a challenge. We present scMGCL, a graph contrastive learning framework for robust integration of single-cell ATAC-seq and RNA-seq data. Our approach leverages self-supervised learning on cell-cell similarity graphs, in which each modality's graph structure serves as an augmentation for the other. This cross-modality contrastive paradigm enables the learning of biologically meaningful, shared representations while preserving modality-specific features. RESULTS: Benchmarking against state-of-the-art methods demonstrates that scMGCL outperforms others in cell-type clustering, label transfer accuracy, and preservation of marker-gene correlations. Additionally, scMGCL significantly improves computational efficiency, reducing runtime and memory usage. The method's effectiveness is further validated through extensive analyses of cell-type similarity and functional consistency, providing a powerful tool for multi-omics data exploration. AVAILABILITY AND IMPLEMENTATION: Code and datasets are released at https://github.com/zlCreator/scMGCL.

Single-Cell Analysis

The ASH HematOmics Program supports integrative analysis of genomic and clinical data in hematologic diseases.

The increasing availability of genomic and transcriptomic sequencing has uncovered diverse genomic alterations and distinct gene expression profiles driving hematologic diseases, yet a data integration and sharing platform dedicated to hematology remains lacking. We developed the American Society of Hematology (ASH) HematOmics Program (ASHOP; ashop.hematology.org), a resource for exploring somatic alterations and gene fusions, transcriptomic results, and clinical data from 5960 patients spanning B-cell precursor and T-cell acute lymphoblastic leukemia, acute myeloid leukemia, myelodysplastic syndromes, and chronic lymphocytic leukemia. Users can explore genomic alteration landscapes and comutation patterns via lollipop and matrix plots and analyze significantly altered genes in user-defined subcohorts. Transcriptomes can be explored through interactive uniform manifold approximation and projections, clustering, differential expression, and pathway enrichment. Genomic, transcriptomic features, and clinical outcomes can be correlated in a user-driven manner or combined to precisely define study cohorts. We illustrate the following 4 use cases of ASHOP: (1) stratification of DUX4-rearranged B-cell leukemias into Early/Multipotent and Committed subgroups with distinct outcomes, (2) characterization of HOXA/HOXB expression patterns in acute myeloid leukemias, (3) correlating mutational burden with mismatch repair deficiency and mutational signatures, and (4) investigation of TP53 alteration landscape. ASHOP is an open-access resource to inform genomic and transcriptomic data interpretation for hematologic malignancies and will expand to support additional diseases and data modalities from the ASH community.

Humans

A profile for managing sensory integrative test data.

A concise method for compiling a data profile from a general sensory integrative test battery has been presented. Subtests from each test used were categorized according to the sensory integrative and motor functions being tested. These categories have been defined and include: tactile-kinesthetic perception, visual perception-figure ground, visual perception-constancy, ocular control, gross motor control, fine motor control, integration of function-two sides of the body, orientation in space, body awareness, and auditory discrimination. A method for converting the various scores into descriptive terminology is provided in which the test results are reported as above age expectancy, appropriate for age, somewhat deficient for age, and markedly deficient for age. The clinical implications of the technique are discussed.

Auditory Perception

RNAcare: integrating clinical data with transcriptomic evidence using rheumatoid arthritis as a case study.

BACKGROUND: Gene expression analysis is a crucial tool for uncovering the biological mechanisms that underlie differences between patient subgroups, offering insights that can inform clinical decisions. However, despite its potential, gene expression analysis remains challenging for clinicians due to the specialised skills required to access, integrate, and analyse large datasets. Existing tools primarily focus on RNA-Seq data analysis, providing user-friendly interfaces but often falling short in several critical areas: they typically do not integrate clinical data, lack support for patient-specific analyses, and offer limited flexibility in exploring relationships between gene expression and clinical outcomes in disease cohorts. Users, including clinicians with a general knowledge of transcriptomics, however, who may have limited programming experience, are increasingly seeking tools that go beyond traditional analysis. To overcome these issues, computational tools must incorporate advanced techniques, such as machine learning, to better understand how gene expression correlates with patient symptoms of interest. RESULTS: Our RNAcare platform, addresses these limitations by offering an interactive and reproducible solution specifically designed for analysing transcriptomic data from patient samples in a clinical context. This enables researchers to directly integrate gene expression data with clinical features, perform exploratory data analysis, and identify patterns among patients with similar diseases. By enabling users to integrate transcriptomic and clinical data, and customise the target label, the platform facilitates the analysis of the relationships between gene expression and clinical symptoms like pain and fatigue. This allows users to generate hypotheses and illustrative visualisations/reports to support their research. As proof of concept, we use RNAcare to link inflammation-related genes to pain and fatigue in rheumatoid arthritis (RA) and detect signatures in the drug response group, confirming previous findings. CONCLUSION: We present a novel computational platform allowing the interpretation of clinical and transcriptomics data in real-time. The platform can be used for data generated by the user, such as the patient data presented here or using published datasets. The platform is available at https://rna-care.mvls.gla.ac.uk/ , and its source code is https://github.com/sii-scRNA-Seq/RNAcare/ .

Humans

Toward large-scale mass spectrometry-based omics for clinical applications.

INTRODUCTION: As healthcare advances toward personalized medicine, mass spectrometry-based research is advancing our understanding of cellular biology and disease states, and translating these findings into clinical applications. This review highlights recent advances in methodology and technology that demonstrate the capabilities of mass spectrometry-based proteomics, lipidomics, and metabolomics in clinical practice. AREAS COVERED: The ability to directly analyze functional molecules with mass spectrometry uncovers crucial clinical information. Each data modality (proteins, lipids, and metabolites) provides essential insight into healthy and disease states. As technology advances, integrating data from different modalities unlocks new possibilities for clinical research. To gain the most from this multi-omic data, unsupervised integration methods can provide detailed insights into complex biological processes. As the field applies this knowledge, healthcare could experience significant leaps in the near future. This review examines recent advancements in mass spectrometry-based proteomics, lipidomics, and metabolomics, focusing on how improvements in sample preparation, automation, and multi-omics data integration are making large-scale clinical studies more accessible. EXPERT OPINION: Recent technical and methodological advancements in mass spectrometry analysis have propelled healthcare toward a tipping point, shifting from traditional RNA- and DNA-based research to downstream analysis of protein, lipid, and metabolite effectors.

Humans

VIJB: a companion of the JBROWSE genome browser for the visually impaired people.

MOTIVATION: The availability of touch-sensitive and haptic devices has been a keystone development for the inclusion of visually impaired people (VIPs) in modern, highly digitized work environments. Braille displays have proven efficient and versatile enough to parse large and complex text files, making bioinformatics and text-heavy programming accessible to VIPs. However, the complex graphical objects -combining numerous datasets- typically generated during data integration remain challenging, even with the aid of descriptive AI. This is particularly true in functional genomics. Here, we present VIJB, a simple application that displays the multilayered output of the JBROWSE genome browser on a Braille reader, enabling VIPs to fully participate in data integration in functional genomics. AVAILABILITY AND IMPLEMENTATION: VIJB is programmed in Python and relies on the scientific library NumPy, the braillegraph and pyBigWig libraries, and the TABIX software. The architecture is summarized in Supplementary Material 1, available as supplementary data at Bioinformatics online. VIJB is available for download at the GitHub repository https://GitHub.com/NiBuMNHN/VIJB and is licenced under the GPL 3.0.

Persons with Visual Disabilities

DIVAS: an R package for identifying shared and individual variations of multiomics data.

MOTIVATION: Multiomics data integration aims to identify biological patterns shared across molecular modalities. Most existing methods detect either jointly shared variation, across all modalities, or individual variation, unique to a single modality, but overlook partially shared variation, shared by only a subset of modalities. This is a critical limitation, because many biological mechanisms manifest in some but not all molecular modalities. RESULTS: We present an open-source R package implementing data integration via analysis of subspaces (DIVAS), a framework for systematically identifying jointly shared, partially shared and individual variations across multiple data types. DIVAS combines angle-based subspace analysis with inference through rotational bootstrap, hierarchically searching all combinations of modalities to decompose multiomics data into interpretable components with scores and loadings. In simulations with a known sharing structure, DIVAS recovered every component across a wide range of noise levels, whereas existing methods did not. Applied to multi-modal COVID-19 data, it reveals partially shared immune and metabolic dysregulation patterns underpinning disease severity that conventional approaches would miss. AVAILABILITY AND IMPLEMENTATION: DIVAS is available at https://github.com/ByronSyun/DIVAS, with documentation and vignettes. The COVID-19 case study vignette is available at https://byronsyun.github.io/DIVAS_COVID19_CaseStudy/.

Multiomics

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n = 91), with fewer studies examining samples from individuals of advanced reproductive age (n = 45) and fetal (n = 16) samples. Transcriptome analyses were most common (n = 103, 85%), followed by proteome (n = 19, 16%) and epigenome (n = 14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female

A comprehensive integration of data on the association of ITPKC polymorphisms with susceptibility to Kawasaki disease: a meta-analysis.

BACKGROUND: This study aims to conduct a comprehensive meta-analysis of existing research to define clear associations between variations in the ITPKC gene and the risk of developing Kawasaki disease (KD). METHODS: A comprehensive search was conducted across multiple databases, including but not limited to PubMed, Scopus, EMBASE, and CNKI, up to June 1, 2024, to gather relevant information. This search utilized keywords and MeSH terms related to hyperbilirubinemia and genetic factors. The inclusion criteria encompassed original case-control, longitudinal, or cohort studies. Correlations were analyzed as odds ratios (ORs) with 95% confidence intervals (CIs) using Comprehensive Meta-Analysis software. RESULTS: Eighteen case-control studies with 5,434 KD cases and 9,419 controls were analyzed. Of these, ten studies assessed 3,129 KD cases and 6,172 controls for the rs28493229 variant, four examined 1,039 cases and 1,688 controls for the rs2290692 variant, two focused on 595 cases and 820 controls for the rs7251246 variant, and two investigated 671 cases and 739 controls for the rs10420685 variant. Results showed a significant association between the rs28493229 polymorphism and increased KD risk across all five genetic models. Subgroup analysis indicated this polymorphism correlates with KD susceptibility in Asians but not in the Chinese population. In contrast, no associations were found between the rs2290692, rs7251246, and rs10420685 polymorphisms and KD risk. CONCLUSIONS: Our pooled data indicate a significant association between the ITPKC rs28493229 polymorphism's minor allele and an increased risk of developing KD, suggesting this variant may enhance susceptibility. Conversely, SNPs rs2290692, rs7251246, and rs10420685 do not demonstrate a statistically significant relationship with KD.

Humans

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products

Multi-omics technologies: Novel tools and methods for assessing nerve injury and regeneration.

Recently, with the rapid advancement of multi-omics technologies, including genomics, transcriptomics, proteomics, and metabolomics, new tools and approaches have been introduced for studying nerve injury and regeneration. This review highlights the application and progress of multi-omics in uncovering the mechanisms of nerve injury, guiding the development of regenerative strategies, and promoting clinical translation. By integrating multi-omics datasets, researchers can comprehensively track dynamic molecular changes following nerve injury, including abnormal gene expression, disrupted protein signaling, altered metabolic programs, and shifts in the immune microenvironment. Single-cell multi-omics technologies resolve cellular heterogeneity, revealing the distinct functions of neurons, glial cells, and immune cell subpopulations during the injury response. Spatially resolved transcriptomics maintain the spatial context of lesion and regeneration sites, enabling precise localization for targeted interventions. Multi-omics technologies not only identify key molecular players involved in nerve regeneration but also create opportunities for personalized medicine. Nonetheless, integrating multi-omics data poses technical challenges, including high dimensionality, batch effects, and algorithmic constraints, while ethical concerns related to stem cell therapy and gene editing require stringent oversight. To transition from structural reconstruction to functional remodeling, future research should emphasize artificial intelligence-driven data integration, organ-on-a-chip modeling, and cross-disciplinary collaboration to overcome existing technical barriers and accelerate the clinical application of neuroregenerative therapies.

artificial intelligence