PubMed HealthSearch

SEARCH · PubMed Health

Results for “data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

An automated clinic management system for a family planning network.

The medical information, financial, and logistic aspects of a comprehensive computer-based Appointment, Registration, Information System, and Evaluation (ARISE) are analyzed for the management of a family planning program serving 30,000 patients annually. An overview of the existing computer system network is presented with descriptions of the interactive master patient index, the batch appointment process, the management statistics package, and Department of Health, Education, and Welfare (HEW) reporting. Emphasis is placed on the financial management control system which includes 1) procedures for third-party submission of claims for payment, in particular Titles IVA, XX, and XIX (Social Security Act), together with discussion of related administrative requirements; 2) technics of auditing data integrity including systematic sampling of collected data; and 3) the process of billing and receipts collection. Methodology and implementation aspects of ARISE may have wide applicability to other family planning and similarly structured clinical programs.

Computers

A medical information relational database system (MIRDS).

A medical information relational database system (MIRDS) which is resident on a relational database machine and is accessed via microcomputers has been created for a pediatric pulmonary division of a research hospital. The power and flexibility of MIRDS has permitted the integration of clinical tasks, research interests, and laboratory functions. Procedures have been devised to assure data integrity, allow flexibility in data retrievals, produce standardized report formats, and permit data access for users with a wide range of query expertise. There are few impediments to the integration of additional clinical, research, and laboratory functions as the system evolves.

Child

A data structure model for a health information system.

The Ministry of Health and Environmental Control of Berlin is developing a Health Information System (HIS) on the basis of a multi-satellite network system, comprising a central, regional, local and functional unit level. The main part of this paper describes the conceptual structure of the Common Data Base (CDB) of HIS with special regard to the patient-oriented medical information originating from the various institutions of the health care system. This structure comprises the following five levels: 1. PATIENT 2. PROBLEM 3. CASE 4. EVENT 5. ACT Each of the levels represents a node in the structure model. A node is an entity with a set of "local properties" being specified for each level, referring to selected data on inferior levels. The structures of these five levels are described in detail. In the last part, so-called "data-manipulation procedures" are treated. These are descriptions covering any data manipulation and represent the basis of data integrity through system controlled transaction with the data base.

Berlin

[Complications of anesthesia in elderly patients].

Progress in surgery and anesthesia has contributed to lowering operative risk and expanding the indications for operations in higher age groups. The goal of treatment in the elderly is to achieve the best possible degree of reducing discomfort and increasing personal independence. Methods. A brochure with a clinical study on 1,021 patients chosen at random shows the frequency of complications arising during the peri- and post-operative course in patients around 60 years of age and older. Operative areas were general and emergency surgery, vascular surgery, neurosurgery, and urology. Operations were carried out in regional or general anesthesia. Patients were divided into groups below and above age 60. Evaluation of the data was carried out according to an integrated data processing concept. This program enables quantitative and qualitative data to be combined at will, taking into consideration that evaluating criteria can be varied considerably. Results. The results demonstrate that patients over 60 have significantly more complications than patients under 60. Analysis of the influence of the factors associated with surgical risk reveals that factors related to the operation such as type, length, and extent do not increase the risk as much as the numerous accompanying illnesses in both age groups. As far more elderly patients are affected by multimorbidity, the conclusion may be drawn that the increased risk observed is not due mainly to age, but rather to the patient's condition prior to surgery. The results indicate clearly that an exact analysis of the initial condition as well as avoiding failure or malfunction of certain organs must have priority in both age groups.

Aged

Auditory temporal integration and the power function model.

The auditory temporal integration function was studied with the objective of improving both its quantitative description and the specification of its principle independent variable, stimulus duration. In Sec. I, temporal integration data from 20 studies were subjected to uniform analyses using standardized definitions of duration and two models of temporal integration. Analyses revealed that these data were best described by a power function model used in conjunction with a definition of duration, termed assigned duration, that de-emphasized the rise/fall portions of the stimuli. There was a strong effect of stimulus frequency and, in general, the slope of the temporal integration function was less than 10 dB per decade of duration; i.e., a power function exponent less than 1.0. In Sec. II, an experimental study was performed to further evaluate the models and definitions. Detection thresholds were measured in 11 normal-hearing human subjects using a total of 24 single-burst and multiple-burst acoustic stimuli of 3.125 kHz. The issues addressed are: the quantitative description of the temporal integration function; the definition of stimulus duration; the similarity of the integration processes for single-burst and multiple-burst stimuli; and the contribution of rise/fall time to the integration process. A power function in conjunction with the assigned duration definition was again most effective in describing the data. Single- and multiple-burst stimuli both seemed to be integrated by the same central mechanism, with data for each type of stimulus being described by a power function exponent of approximately 0.6 at 3.125 kHz. It was concluded that the contribution of the rise/fall portions of the stimuli can be factored out from the rest of the temporal integration process. In Sec. III, the conclusions that emerged from the review of published work and the present experimental work suggested that auditory temporal integration is best described by a power function in conjunction with the assigned duration definition. The exponent for the power function is typically less than 1.0, and varies with frequency and hearing level. Second, a means of empirically assaying the contribution of the rise-fall portions of the stimuli is presented and evaluated. Finally, properties of a central auditory integrator are hypothesized.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

Monitoring adverse drug reactions-the problem of integration of heterogeneous data.

Effective systems for a meaningful integration and interpretation of data of heterogeneous origin are essential for successful monitoring of adverse drug reactions internationally or in multicenter programs. Particular difficulties are encountered in the area of suspected drug adverse reactions and in the area of drugs. Both problem areas are discussed in the paper.

Drug-Related Side Effects and Adverse Reactions

Quality control methods for data entry in pathology using a computerized data management system based on an extended data dictionary.

In pathology, computerized data management systems have been used increasingly to facilitate a more efficient supply of information. Since data entry precedes data utilization, the reliability of the information stored strongly depends on the quality of data input. Despite its potential capability, most personal computer-based database software does not provide versatile and user-friendly data validation procedures. Therefore, we developed a data dictionary-driven data management system that enables the user to perform extensive validation routines without the need for hard programming. Using examples from an existing database for endometrial carcinomas, different types of data errors and their error traps are explained. It is pointed out that data type definitions, defaults, templates, or picture clauses are suitable means to avoid formal errors. Validations on data domains and ranges test whether data fall into a predefined scope. Relational checks control data validity within a context of different data items, whereas process routines provide automatic data computation, thereby circumventing user input. By exploiting the facilities of an extended data dictionary, a powerful tool is made available to secure various aspects of data integrity simultaneously with input. In this way, computerized data quality control can improve the efficiency and reliability of data management tasks in pathology.

Medical Informatics Computing

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics

SESAM: a relational database for structure and sequence of macromolecules.

A system is described that provides ways of integrating data on protein structure, sequence, and survey results, with molecular graphics and molecular mechanics software. Its major component is the relational database SESAM, presently implemented under the commercial package SYBASE. By design, the database allows full integration--within the same data organization--of raw data on protein structure, sequence, ligands, and heterogroups, obtained from the Brookhaven Protein Databank, with pure sequence information available from other databanks such as SWISS-PROT. It contains in addition higher level descriptions of structural and topological properties, as well as survey results, obtained by executing specialized computer programs. Aside from the very useful attribute of closely combining structural and nonstructural information, other important features distinguish it from analogous systems developed elsewhere. It includes a molecular dictionary with complete description of geometric properties and energy parameters used in modeling and conformational energy calculations. Using this dictionary, structural data are validated by checking for localized inconsistencies in atomic coordinates, atomic symbols, chirality definitions, and flagging errors and incomplete entries. Because of both the dictionary and the validation procedures, SESAM can be readily interfaced with conventional molecular graphics and mechanics software packages, or with other specialized application programs. With the aid of appropriate interfaces, data access is sufficiently fast for SESAM to be interrogated interactively. Prototypes of user interfaces, as well as an interface with the molecular graphics package BRUGEL, are described and the power of the system is illustrated in applications such as homology-based protein modeling, computer-aided protein design, protein structure predictions, analysis of local structure motifs, and of relationships between protein sequence and structure.

Amino Acid Sequence

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans

An integrative network approach for longitudinal stratification in Parkinson's disease.

Parkinson's disease (PD) is a neurodegenerative disorder characterized by motor symptoms resulting from the loss of dopamine-producing neurons in the brain. Currently, there is no cure for the disease which is in part due to the heterogeneity in patient symptoms, trajectories and manifestations. There is a known genetic component of PD and genomic datasets have helped to uncover some aspects of the disease. Understanding the longitudinal variability of PD is essential as it has been theorised that there are different triggers and underlying disease mechanisms at different points during disease progression. In this paper, we perform longitudinal and cross-sectional experiments to identify which data modalities or combinations of modalities are informative at different time points. We use clinical, genomic, and proteomic data from the Parkinson's Progression Markers Initiative. We validate the importance of flexible data integration by highlighting the varying combinations of data modalities for optimal stratification at different disease stages in idiopathic PD. We show there is a shared signal in the DNAm signatures of participants with a mutation in a causal gene of PD and participants with idiopathic PD. We also show that integration of SNPs and DNAm data modalities has potential for use as an early diagnostic tool for individuals with a genetic cause of PD.

Parkinson Disease

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans

Precautions in topographic mapping and in evoked potential map reading.

First, we consider the main points that must be addressed when constructing topographic maps: types of projection, methods of interpolation, number and locations of recording electrodes, and color scales. Data integrity and precautions in map interpretation are then examined for the case of evoked potential data.

Brain

Unlocking the Full Potential of Spatial Omics in Plants: Practical Challenges, Solutions, and a Path Forward.

Spatial omics technologies are providing new opportunities for plant biology by enabling molecular profiling within structurally intact tissues, revealing spatially organised cell states, developmental gradients, and regulatory interactions. While spatial transcriptomics has driven early advances, the field is rapidly expanding toward integrated spatial multi-omics by combining single-cell and spatial transcriptomic, epigenomic, proteomic, and metabolomic data. These approaches offer new opportunities to study development, physiology, and plant biotic and abiotic interactions in spatially preserved cellular contexts. However, despite rapid adoption, the field remains constrained by plant-specific challenges when applying technologies largely developed for animal systems. Compared with animal systems, plant tissues pose additional challenges due to rigid cell walls, and diverse chemistries, complicating sample preparation, cell and subcellular segmentation, signal detection, and data integration. As a result, many studies rely on bespoke protocols and analysis pipelines that are often difficult to reproduce or generalise. Here, we provide a practical, solution-oriented synthesis of current bottlenecks across experimental and computational pipelines, highlight emerging strategies to overcome these limitations, and propose a roadmap for community-driven protocol sharing, benchmarking, and integration across spatial and multi-omics modalities. Addressing these challenges will be essential to establish spatial omics as a routine and scalable tool for plant biology.

Journal Article

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

Integrating multi-omics approaches in acute myeloid leukemia (AML): Advancements and clinical implications.

Acute myeloid leukemia (AML) is a highly heterogeneous and aggressive hematologic malignancy characterized by clonal proliferation of myeloid precursors. Despite significant advancements in genomic profiling and targeted therapies, patient outcomes remain suboptimal due to disease complexity, resistance mechanisms, and high relapse rates. The integration of multi-omics approaches-spanning genomics, epigenomics, transcriptomics, proteomics, and metabolomics-has revolutionized AML research, offering a comprehensive understanding of leukemogenesis, tumor heterogeneity, and therapeutic vulnerabilities. Recent studies leveraging high-throughput sequencing, mass spectrometry, and advanced computational tools have uncovered novel biomarkers, clonal evolution dynamics, and microenvironmental interactions that drive AML progression and resistance. For instance, single-cell multi-omics has revealed chemotherapy-resistant leukemic stem cell populations, while proteogenomic analyses have identified actionable targets such as MCL1 and metabolic dependencies like OXPHOS. Clinically, integrated omics platforms are refining risk stratification, minimal residual disease (MRD) monitoring, and personalized therapy selection. However, challenges such as data integration complexity, cost barriers, and ethical considerations remain. This review highlights the transformative potential of multi-omics in AML, emphasizing recent advancements in technology, biomarker discovery, and therapeutic innovation. By bridging the gap between molecular insights and clinical practice, multi-omics integration promises to redefine AML management, paving the way for precision oncology and improved patient outcomes.

Humans

Multi-omics approaches in idiopathic pulmonary fibrosis: from molecular mechanisms to therapeutic targets and precision medicine.

Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease with limited therapeutic options and marked molecular heterogeneity. Despite available antifibrotic therapies, disease progression remains poorly predictable, highlighting the need for improved mechanistic understanding and therapeutic targeting. This review summarizes recent advances in multi-omics research to elucidate the molecular mechanisms underlying IPF and to identify potential biomarkers and pharmacological targets. Multi-omics studies, including genomics, epigenomics, transcriptomics, proteomics, metabolomics, microbiome profiling, and single-cell sequencing, have revealed key pathogenic mechanisms in IPF. Genetic susceptibility factors such as MUC5B promoter variants and telomere-related genes contribute to disease risk. Epigenetic regulation, including DNA methylation, histone modifications, and non-coding RNAs, plays a central role in fibrotic remodeling. Transcriptomic and proteomic analyses have identified dysregulated signaling pathways, including TGF-β, mTOR, cellular senescence, and extracellular matrix remodeling. Metabolomic alterations indicate disrupted lipid and amino acid metabolism. Importantly, integration of multi-omics datasets enables the identification of molecular endotypes, candidate biomarkers, and potential therapeutic targets. However, challenges including data integration, tissue heterogeneity, limited cohort size, and the need for functional validation remain important barriers to clinical translation. Continued development of multi-omics approaches may facilitate more accurate disease classification and support the development of personalized therapeutic strategies for IPF.

biomarkers