PubMed HealthSearch

SEARCH · PubMed Health

Results for “data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

Integrating multi-omics approaches in acute myeloid leukemia (AML): Advancements and clinical implications.

Acute myeloid leukemia (AML) is a highly heterogeneous and aggressive hematologic malignancy characterized by clonal proliferation of myeloid precursors. Despite significant advancements in genomic profiling and targeted therapies, patient outcomes remain suboptimal due to disease complexity, resistance mechanisms, and high relapse rates. The integration of multi-omics approaches-spanning genomics, epigenomics, transcriptomics, proteomics, and metabolomics-has revolutionized AML research, offering a comprehensive understanding of leukemogenesis, tumor heterogeneity, and therapeutic vulnerabilities. Recent studies leveraging high-throughput sequencing, mass spectrometry, and advanced computational tools have uncovered novel biomarkers, clonal evolution dynamics, and microenvironmental interactions that drive AML progression and resistance. For instance, single-cell multi-omics has revealed chemotherapy-resistant leukemic stem cell populations, while proteogenomic analyses have identified actionable targets such as MCL1 and metabolic dependencies like OXPHOS. Clinically, integrated omics platforms are refining risk stratification, minimal residual disease (MRD) monitoring, and personalized therapy selection. However, challenges such as data integration complexity, cost barriers, and ethical considerations remain. This review highlights the transformative potential of multi-omics in AML, emphasizing recent advancements in technology, biomarker discovery, and therapeutic innovation. By bridging the gap between molecular insights and clinical practice, multi-omics integration promises to redefine AML management, paving the way for precision oncology and improved patient outcomes.

Humans

Multi-omics approaches in idiopathic pulmonary fibrosis: from molecular mechanisms to therapeutic targets and precision medicine.

Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease with limited therapeutic options and marked molecular heterogeneity. Despite available antifibrotic therapies, disease progression remains poorly predictable, highlighting the need for improved mechanistic understanding and therapeutic targeting. This review summarizes recent advances in multi-omics research to elucidate the molecular mechanisms underlying IPF and to identify potential biomarkers and pharmacological targets. Multi-omics studies, including genomics, epigenomics, transcriptomics, proteomics, metabolomics, microbiome profiling, and single-cell sequencing, have revealed key pathogenic mechanisms in IPF. Genetic susceptibility factors such as MUC5B promoter variants and telomere-related genes contribute to disease risk. Epigenetic regulation, including DNA methylation, histone modifications, and non-coding RNAs, plays a central role in fibrotic remodeling. Transcriptomic and proteomic analyses have identified dysregulated signaling pathways, including TGF-β, mTOR, cellular senescence, and extracellular matrix remodeling. Metabolomic alterations indicate disrupted lipid and amino acid metabolism. Importantly, integration of multi-omics datasets enables the identification of molecular endotypes, candidate biomarkers, and potential therapeutic targets. However, challenges including data integration, tissue heterogeneity, limited cohort size, and the need for functional validation remain important barriers to clinical translation. Continued development of multi-omics approaches may facilitate more accurate disease classification and support the development of personalized therapeutic strategies for IPF.

biomarkers

Defining drug use: a model for the integration of measures through the census tract.

The nature and severity of drug use has been measured both directly and indirectly by various studies employing different indicators, although the majority of studies still tend to use single measures of drug use. The need to employ multiple measures in examining drug abuse is constrained by the fact that available data may have been collected through diverse methodologies and measured on different levels or units. The purpose of this study was to develop and test in Philadelphia a model using qualitatively different types of data integrated by the common geographic unit of a census tract. The types of data used included: archival data, key informant data, and survey data. Using this approach the paper examines the relationships of drug use measures to each other, to the social environment, and to drug market factors. Major findings of the analysis indicate that there are several independent measures of drug use as reflected in five composite indicators which differentiate behavioral activities or consequences of drug use. Moreover, heroin use indicators exhibit relationships with social-environmental characteristics and drug market factors which are different from those existing with amphetamine or synthetic drug use.

Adult

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans

The social demography of drug use.

This review of the epidemiology of drug use and drug dependence/abuse describes the overall current pattern of use of mood-changing legal and illegal drugs based on the most recently available data, variations in drug use among subgroups in the population, and trends in drug use over time. This is the first systematic attempt to integrate data on patterns of use with data on drug dependence/abuse. In the course of the analysis an effort is made to account for two paradoxes: blacks report the lowest rate of drug use in general population studies, yet constitute the largest category of treated cases or drug-related casualties. Although prevalence of cocaine use in the general population decreased, beginning in 1986, morbidity and mortality related to cocaine increased sharply in 1990.

Age Factors

Advancing One Health genomics in Africa: opportunities and challenges for outbreak and antimicrobial resistance control.

SUMMARYAfrica's ongoing struggles with emerging epidemics and antimicrobial resistance (AMR) underscore the urgency of integrating pathogen genomics and surveillance systems into the continent's One Health strategy, particularly given the existing limitations in preparedness and technological resources. This review brings together current evidence on the growth of sequencing infrastructure, the development of regional genomic hubs, and the establishment of governance frameworks, while identifying critical challenges in data integration, bioinformatics capacity, and sustainable financing. Special focus is placed on the lack of African-based genomic data, with our analysis showing that only 1.82% of the global total is available. Case studies illustrate the immense potential and importance of pathogen genomics, giving policymakers a tangible sense of its impact. These examples demonstrate how genomic technologies integrated with artificial intelligence (AI) are transforming outbreak response, AMR surveillance, and stewardship programs by enabling early detection of zoonotic threats, mapping transmission pathways, and guiding vaccine development. However, to fully realize this scientific intel, it is essential to embed One Health pathogen surveillance within strong policy and system frameworks to ensure the translation of technical progress into lasting institutional capacity and sustainable impact. Long-term implementation depends on coordinated investment and advocacy across four interdependent pillars: data architecture, governance and sovereignty, human capital, and technical capacity.

Humans

Efforts towards a precision medicine approach in juvenile idiopathic arthritis.

Juvenile idiopathic arthritis (JIA) is the commonest group of childhood arthritides. Despite the availability of advanced therapeutics, many children and young people (CYP) with JIA experience disease flares, and in some, chronic joint damage. Tailoring treatment based on unique biological profiles would benefit CYP with JIA given their variable clinical presentation and disease course. To date, biomarkers to predict treatment response are lacking. With advances in single cell technologies, we are now able to profile the genes and proteins of target tissues at unprecedented resolution to define the biological basis of disease and guide novel treatment approaches. The complex analyses and combination of biological and clinical outcome data from large datasets across disease phenotypes have become possible with the development of computational and machine learning methods. Here, we summarize the strategies to integrate data through multimodal based approaches to maximize precision medicine and research priorities for CYP with JIA.

Humans

Threshold integration of bi-amplitude signals.

This study examined the pattern of intensity integration at threshold. The stimuli studied are unique in that they have a compound peak-to-peak amplitude envelope. This waveform was partitioned into two segments, each having a different peak-to-peak magnitude (bi-amplitude). All signals in the bi-amplitude series were the same duration (100 ms). Therefore, threshold differences between these signals are due solely to the integration of intensity in the amplitude dimension. A prediction of the pattern of thresholds, based on the diverted-input hypothesis, suggested that little or no integration would occur when the amplitude difference between segments is greater than a specific magnitude. Our results indicate that there are similarities in the integration process found with variable duration signals and with bi-amplitude signals. We conclude that previous estimates of the minimum intensity level based on temporal integration data underestimates the intensity levels that can contribute to threshold. Our results suggest that there are no apparent constraints on the intensity levels that can be integrated near threshold. The auditory system integrates distributed stimulus intensity in both the time and amplitude dimensions. Temporal integration in the auditory system can be viewed as a signal process, where an enhanced internal representation is given low-level stimuli.

Acoustic Stimulation

Surface reflections of cardiac excitation and the assessment of infarct volume in dogs. A comparison of methods.

Ventricular depolarization was analyzed in intact dogs by simultaneously recording body surface potential maps, McFee axial vectorcardiograms, and a 5 X 4 lead precordial grid of QRS complexes. The purpose of this study was to compare the effectiveness of subtraction approaches, using the simultaneously acquired data. The totally closed chest approach avoided the problem of volume conductor alteration by thoracotomy. Infarct volume was calculated morphologically from measurements of serial ventricular sections. The maximal correlation with anatomic infarct size using the precordial QRS grid approach was 0.51, using cumulative difference data between 1 and 38 msec when the postinfarction grid was substracted from the preinfarction grid. A correlation coefficient of 0.80 was achieved using the numerically integrated data between 1 and 31 msec from the vectorcardiogram, and the body surface potential map achieved a correlation coefficient above 0.88 when the electrical difference of msec 16 was used. These data suggest that estimates of infarct size from selected surface reflections of the activation process are feasible if some sort of preinfarction control data are available. Caution must be exercised to avoid inclusion of electrical effects late in the activation process which contain contamination by highly variable alterations in the excitation sequence due to delayed conduction or alteration in conduction pathway in or near the infarct zone.

Action Potentials

Privacy-Preserving Linkage of Distributed Biological, Clinical, and Imaging Data Supporting Artificial Intelligence in Pediatric Oncology.

BACKGROUND: Cancer remains the leading cause of disease-related mortality in children over the age of one in Europe, with over 35,000 new pediatric cases and more than 6,000 deaths annually. Due to the rarity of pediatric cancers, clinical trial protocols often substitute for formal treatment guidelines, resulting in many children being enrolled in multiple trials, with biological samples and genomic data stored in various biobanks. Data collection in pediatric oncology is challenging, with sparse data acquired over extended periods, underscoring the need for optimal utilization of all available information through linked, privacy-preserving datasets. METHODS: Here, we report the development of a distributed, privacy-preserving data infrastructure for the PRIMAGE project, a European initiative aimed at supporting artificial intelligence (AI)-driven image analysis for pediatric cancer prognostics. The infrastructure leverages the European Patient Identity (EUPID) Services for Privacy-Preserving Record Linkage, enabling pseudonymized data integration across clinical, biological, and imaging sources. The system incorporates EUPID's hashing and phonetic matching protocols to pseudonymize patient identifiers and link distributed datasets, facilitating secondary data use in compliance with the General Data Protection Regulation. RESULTS: Data from over 700 neuroblastoma patients from European trials and hospitals were linked and uploaded to the PRIMAGE platform, where AI models predict clinical outcomes. CONCLUSION: This infrastructure successfully facilitated AI model development, advancing pediatric oncology research, and offering a scalable framework for future European health data initiatives, such as the European Health Data Space.

Journal Article

Bridging genotype, phenotype, and clinical insight: the role of multi-omics in cardiovascular disease.

INTRODUCTION: It is increasingly evident that the multifactorial nature of cardiovascular disease requires the combination of different omics approaches for improving our mechanistic understanding, identifying novel drug targets, and developing accurate diagnostic, predictive, and prognostic biomarker panels. AREAS COVERED: We review the current state and the potential of multi-omics in cardiovascular disease, with a specific focus on plasma-, spatial-, and single-cell approaches. We discuss lipidomics as a genotype‑to‑phenotype bridge, the utility of remote longitudinal monitoring via microsampling/dried blood spots, and emerging clinical‑trial integrations of multi-omics approaches. We outline critical gaps in standardization and how to overcome these, pre‑analytical challenges and constraints that are often neglected, and data‑integration methods spanning from canonical correlation analysis to modern machine learning approaches. EXPERT OPINION: Multi‑omics can shape cardiovascular care by identifying drug targets in diseased tissue and by yielding small, usable biomarker panels.

Humans

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

The conflict between relational databases and the hierarchical structure of clinical trials data.

Relational database software has become popular for the management of certain types of commercial data. Its use is being given serious consideration in the management of data from clinical trials. Relational systems have a number of advantages over hierarchical or network systems for some types of data. However, as illustrated by an example, the data from clinical trials typically have an inherent hierarchical structure. The incorporation of hierarchically structured data into a relational database raises difficult problems of data integrity versus the complexity of the database structure. These problems, together with the long execution times of many relational operations, indicate that relational systems are not necessarily well suited for clinical trials data management.

Clinical Trials as Topic

The significance of teeth in pollution detection.

The general population is experiencing lifelong exposure to old and new hazardous substances. By using data collected from a subject's own teeth, accuracy in determining the effects of exposure is assured since extrapolation is excluded. The establishment of a common tooth bank can provide means to integrate data from multiple sources. Comprehensive pollution information shared by the environmental, scientific and medical communities can lead to a more efficient approach to a worldwide problem.

Animals

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms