PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

AI-HOPE: an AI-driven conversational agent for enhanced clinical and genomic data integration in precision medicine research.

MOTIVATION: The growing complexity of clinical cancer research has fueled a surge in demand for automated bioinformatics tools capable of integrating clinical and genomic data to accelerate discovery efforts. RESULTS: We present the Artificial Intelligence Agent for High-Optimization and Precision Medicine (AI-HOPE), an AI-driven system that enables domain experts to conduct integrative data analyses through natural language interactions. Powered by Large Language Models, AI-HOPE interprets user instructions, converts them into executable code, and autonomously analyzes locally stored data. It supports flexible association studies, subset comparisons, clinical prevalence assessments and survival analyses. In addition, AI-HOPE enables global variable scans to identify features significantly associated with a user-defined outcome, making a powerful and intuitive tool for advancing precision medicine research. Importantly, its closed-system design prevents clinical data leakage. To demonstrate its utility, AI-HOPE was applied to The Cancer Genome Atlas data to address two clinical questions. First, it identified significant enrichment of TP53 mutations in late-stage colorectal cancer compared to early-stage cases. Second, it uncovered a strong association between KRAS mutations and poorer progression-free survival in FOLFOX-treated patients. These findings align with established literature and demonstrate AI-HOPE's ability to generate meaningful insights independently, without prior assumptions. By removing programming barriers and simplifying complex analyses, AI-HOPE bridges the gap between data complexity and research needs. With its scalable and adaptable framework, AI-HOPE has the potential to support diverse biomedical research fields, driving innovation and efficiency in translational studies. AVAILABILITY AND IMPLEMENTATION: The AI-HOPE software and demonstration data is available at https://github.com/Velazquez-Villarreal-Lab/AI-HOPE.

Precision Medicine

Scalable, open-access and multidisciplinary data integration pipeline for climate-sensitive diseases.

Climate-sensitive infectious diseases pose an important challenge for human, animal and environmental health and it has been estimated that over half of known human pathogenic diseases can be aggravated by climate change. While climatic and weather conditions are important drivers of transmission of vector-borne diseases, socio-economic, behavioural, and land-use factors as well as the interactions among them impact transmission dynamics. Analysis of drivers of climate-sensitive diseases require rapid integration of interdisciplinary data to be jointly analysed with epidemiological (including genomic and clinical) data. Current tools for the integration of multiple data sources are often limited to one data type or rely on proprietary data and software. To address this gap, we develop a scalable and open-access pipeline for the integration of multiple spatio-temporal datasets that requires only the declaration of the country and temporal range and resolution of the study. The tool is locally deployable and can easily be integrated into existing climate-disease-modelling applications. We demonstrate the utility of the tool for dengue modelling in Vietnam where epidemiological data are legally required to remain local. We include a pipeline for bias correction of climate data to enhance their quality for downstream modelling tasks. The Dengue Advanced Readiness Tools-Pipeline empowers users by simplifying complex download, correction, and aggregation steps, fostering data-driven discovery of relationships between infectious diseases and their drivers in space and time, and enhancing reproducibility in research. Additional modules and datasets can be added to the existing ones to make the pipeline extendable to use cases other than the ones presented here.

automated workflows

Integrated data mining and network pharmacology to explore the prescription patterns from a senior TCM oncologist's clinical practice in treating chemotherapy-induced hand-foot syndrome.

Hand-foot syndrome (HFS) is a common and refractory adverse effect of chemotherapy lacking specific therapeutic strategies currently. Traditional Chinese medicine (TCM) has shown empirical efficacy in clinical HFS management. This study integrated data mining and network pharmacology to systematically elucidate the medication principles and molecular mechanisms underlying Professor Gang Xie's prescriptions for HFS. All medical records from Professor Xie's specialist clinic (January 2020 to March 2025) were retrospectively collected and standardized in Excel. Prescriptions were analyzed through frequency statistics, association and clustering. Active ingredients of core herb pairs and their disease-related targets were identified using TCMSP, HERB, GeneCards, PharmGKB and GEO databases. Protein-protein interaction (PPI) networks, gene ontology (GO), and Kyoto encyclopedia of genes and genomes (KEGG) pathway analyses were performed. Molecular docking validated interactions between key bioactive compounds and targets. This study involved 217 prescriptions containing 150 herbs. Core herb combinations comprised Radix Astragali (Huangqi), Poria (Fuling), and Radix Pseudostellariae (Taizishen), predominantly classified as spleen-tonifying agents with warm properties, targeting lung, spleen, and stomach meridians. Network analysis identified 67 bioactive compounds and 899 disease targets. Quercetin, kaempferol, acacetin and luteolin were identified the key ingredients. The core targets (TP53, STAT3, PIK3CA, HSP90AA1, AKT1, CTNNB1, PI3KR1, MAPK1) were enriched in MAPK and PI3K-Akt signaling pathways. Molecular docking confirmed strong binding affinity between key compounds and targets. Professor Xie's therapeutic strategy for HFS emphasizes "spleen fortification, phlegm elimination, and stasis resolution." The core herb combination likely exerts anti-HFS effects via modulation of MAPK and PI3K-Akt pathways, providing a pharmacological basis for TCM-driven HFS management.

Network Pharmacology

Global inequities in hepatitis B and C genomic surveillance revealed through an interactive data integration dashboard.

OBJECTIVES: To assess global disparities in hepatitis B virus (HBV) and hepatitis C virus (HCV) genomic surveillance and to develop an integrated platform that links genomic data with epidemiological burden. STUDY DESIGN: Retrospective observational analysis. METHODS: We reviewed existing viral genomic repositories to identify structural and analytical limitations. Subsequently, we integrated 10 996 HBV and 3533 HCV whole-genome sequences (WGS) from public databases with Global Burden of Disease (GBD) estimates to quantify inequities in genomic surveillance across countries and genotypes. Using these data, we developed the open-access Hepatitis Dashboard, incorporating >14 000 sequences from 141 countries with GBD metrics to evaluate representativeness and sequencing coverage relative to disease burden. RESULTS: Marked inequities in hepatitis genomic surveillance were identified. Despite increasing HBV- and HCV-associated mortality, virus sequence availability remains geographically and genotypically skewed-dominated by China and the United States, with substantial underrepresentation of HBV genotype E and HCV genotypes 5 and 8. Many high-endemic countries in Africa and the Western Pacific remain severely undersampled. We detected circulating antiviral drug-resistance mutations and developed a burden-adjusted sequencing coverage metric, revealing that several high-burden countries, including China, Nigeria and India, are among the least represented in global genomic datasets. Projections to 2030 indicate that neither HBV nor HCV are currently on track to meet WHO elimination targets. CONCLUSIONS: The Hepatitis Dashboard provides an integrated, continuously updated resource that links genomic and epidemiological data to quantify and visualise global surveillance gaps. This analysis highlights a critical disconnect between sequencing efforts and public health needs, which may limit the effectiveness of surveillance-informed strategies to support progress toward WHO 2030 elimination goals. By enabling burden-adjusted prioritisation and longitudinal tracking of genomic coverage, the platform supports evidence-based sampling strategies, equitable resource allocation, and monitoring of global progress toward hepatitis elimination.

Humans

Research data integrity: a result of an integrated information system.

The toxicologic problems of today frequently require long-term, multidisciplinary experimentation involving large numbers of animals. In order to provide the extensive safety evaluation necessary to produce data that can be reasonably extrapolated to humans, automated research support systems have transcended the position of useful tools and have become an integral part of the total design of experimental protocols. For an automated information system to fully represent the reality of the experiment, it must be able to assure integrity, as well as provide for the storage, calculation, and retrieval of data values of the quality and quantity necessary for fulfilling protocol requirements. Guarantees against error and loss of data, in addition to flexibility and easy access, must be an inherent part of the system if the acceptance and condifence of the investigator are to be obtained. This paper discusses the criteria, philosophies, and benefits of integrated data systems that ensure integrity of toxicologic research support.

Computers

Interpretable data integration for single-cell and spatial multi-omics.

Integrating single-cell or spatial transcriptomic and epigenomic data enables scrutinizing the transcriptional regulatory mechanisms controlling cell fate. Current integration methods usually align multi-omics data into a shared latent space but fail to reveal the underlying connections between genes and regulatory elements. The correlation- or regression-based regulatory inference methods cannot dissect different transcriptional regulation codes for cells under different spatial and temporal states. To address both problems, we develop a feature-guided optimal transport (FGOT) method, which simultaneously uncovers cellular heterogeneity and their associated transcriptional regulatory links. FGOT also provides post hoc interpretability for existing integration methods. FGOT is applicable for paired/unpaired single-cell multi-omics data and paired spatial multi-omics data. Benchmarking and validating via histone modification data or three-dimensional (3D) genomics data show good robustness and accuracy in integration and inference of regulatory links. The method allows systematic screening of cell-state and spatial-location-specific regulatory elements in diseases at the single-cell level. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

Improving recombinant protein productivity in CHO cells via multi-omics data integration.

Chinese hamster ovary (CHO) cells represent the dominant host system for the production of recombinant therapeutic proteins. In recent decades, extensive research has focused on process/media optimization and cell line engineering to improve both the productivity and quality of biopharmaceutical proteins produced in CHO cells. Nevertheless, the inherent complexity of biological pathways and the heterogeneous cellular responses to different environmental conditions have posed substantial challenges to traditional methodologies. Recent advances in omics technologies have enabled comprehensive characterization of CHO cell physiology, providing multidimensional molecular and phenotypic insights that facilitate the enhancement of recombinant protein production. This review first summarizes the methodologies and advances in CHO omics research, including genomics, transcriptomics, proteomics, metabolomics, and epigenomics. It then examines contemporary approaches to integrate and analyze multi-omics data in CHO cells. The review further elucidates how these multi-omics datasets can be strategically applied across various developmental stages, including cell line selection, genetic engineering, expression vector design, and bioprocess optimization. Finally, we explore the transformative potential of integrating multi-omics with artificial intelligence and discuss promising future research directions in CHO cell studies. These emerging paradigms offer novel opportunities for data-driven cell engineering and bioprocess optimization in CHO-based biomanufacturing.

Bioprocessing

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software

SeqUIaSCOPE: multi-omics data integration platform for single-patient clinical oncology pathway exploration.

SUMMARY: SeqUIaSCOPE is an open-source platform designed for routine clinical oncology diagnostics through case-centric integration and visualization of genomic variants, fusion events, and expression profiles. The platform combines molecular-level validation via embedded genome browsing with systems-level interpretation through dynamic pathway visualization, enabling geneticists to assess how alterations converge across biological networks. Flexible reporting with customizable templates accommodates diverse institutional requirements, while secure cluster-based or local deployment ensures compliance with data protection policies, making advanced multi-omics diagnostics accessible to academic and clinical institutions. AVAILABILITY AND IMPLEMENTATION: SeqUIaSCOPE is freely available on GitHub at https://github.com/BioIT-CEITEC/sequiascope under the MIT license and archived at Zenodo (https://zenodo.org/records/21338445). Due to the sensitive nature of patient data, the repository provides simulated datasets that mimic the structure of real clinical data for testing and exploration. Documentation and a live demo accompany these datasets, allowing users to explore the application without any prior setup. The repository also includes a Helm chart for Kubernetes deployment and Docker containers for local deployment, ensuring compatibility across Linux, macOS, and Windows. No user registration is required, and all data remains on local or institutional infrastructure.

Humans

Prioritization of causal genes from genome-wide association studies by Bayesian data integration across loci.

MOTIVATION: Genome-wide association studies (GWAS) have identified genetic variants, usually single-nucleotide polymorphisms (SNPs), associated with human traits, including disease and disease risk. These variants (or causal variants in linkage disequilibrium with them) usually affect the regulation or function of a nearby gene. A GWAS locus can span many genes, however, and prioritizing which gene or genes in a locus are most likely to be causal remains a challenge. Better prioritization and prediction of causal genes could reveal disease mechanisms and suggest interventions. RESULTS: We describe a new Bayesian method, termed SigNet for significance networks, that combines information both within and across loci to identify the most likely causal gene at each locus. The SigNet method builds on existing methods that focus on individual loci with evidence from gene distance and expression quantitative trait loci (eQTL) by sharing information across loci using protein-protein and gene regulatory interaction network data. In an application to cardiac electrophysiology with 226 GWAS loci, only 46 (20%) have within-locus evidence from Mendelian genes, protein-coding changes, or colocalization with eQTL signals. At the remaining 180 loci lacking functional information, SigNet selects 56 genes other than the minimum distance gene, equal to 31% of the information-poor loci and 25% of the GWAS loci overall. Assessment by pathway enrichment demonstrates improved performance by SigNet. Review of individual loci shows literature evidence for genes selected by SigNet, including PMP22 as a novel causal gene candidate.

Genome-Wide Association Study

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning