PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The evolution of an integrated timeline for oncology patient healthcare.

The introduction of computers in the medical environment has contributed to the proliferation of medical data, often making it difficult to consolidate information on a single patient. In patients with complex medical problems, such as oncology patients, the lack of data integration can negatively impact on patient care. This paper presents an infrastructure for the creation of an integrated multimedia timeline that automatically combines patient information from distributed hospital information sources, and creates a visual summary of pertinent events in a patient's medical history. In this prototype, we focus on oncology patients under treatment for advanced cancers.

Database Management Systems↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

Precautions in topographic mapping and in evoked potential map reading.

First, we consider the main points that must be addressed when constructing topographic maps: types of projection, methods of interpolation, number and locations of recording electrodes, and color scales. Data integrity and precautions in map interpretation are then examined for the case of evoked potential data.

Brain↗

Cognitive evaluation of decision making processes and assessment of information technology in medicine.

This paper describes cognitive methods for analyzing medical decision making and evaluating medical information systems. The overall approach focuses on understanding the processes involved in the decision making and reasoning of health care workers, both with and without the use of information technologies. The issue of developing appropriate evaluation tools, for use in the design and analysis of medical information systems is considered to be of great importance. However, conventional methods are limited in their ability to identify and characterize the effects of information technology on the cognitive processes involved in decision making and reasoning. In this paper a range of methods are described involving video recording for collecting data on the use of information systems. The techniques described allow for the collection of an integrated data set consisting of transcripts of health care workers as they 'think aloud' in interacting with a medical system, along with complete video records of user-computer interaction. In addition, the methods can be extended to allow for the collection of process data from video recording of systems in actual clinical and emergency situations. The use of a variety of approaches, borrowing from research in cognitive science, is discussed. The development and application of these evaluation methods within the Canadian Centres of Excellence network HEALNet is subsequently described. Finally, implications for the development and evaluation of medical information systems are considered.

Cognition↗

Physical mapping: integrating computational and molecular genetic data.

A crucial step beyond the identification of genetic linkage of a disease to a chromosomal region is the production of a physical map that will allow the identification of candidate genes. Although the process of physical map building has been facilitated by the flow of data released by the Human Genome Project, gathering all the information together requires significant effort. In a previous study, we reported linkage between Bipolar Affective Disorder and the chromosomal location 4p15.3--p16.1. In this review we use this example to describe how to collect publicly available sequence, DNA fingerprint, and genetic marker data and integrate these with empirical data to build a large scale high resolution physical map of a region. Methods used to identify new genetic markers and candidate genes within a circumscribed region are also presented.

Databases, Factual↗

Health care informatics: the key to successful disease management.

Health services integration and disease state management (DSM) require improved health care informatics systems. Accurate, comprehensive patient information and an integrated data infrastructure are needed for all stages of DSM, from development and implementation of programs to evaluation and continuous program improvement. The lack of an integrated information infrastructure is one of the leading obstacles to achieving a comprehensive electronic patient data system. This article examines initiatives underway to make the computer-based patient record a reality.

Attitude of Health Personnel↗

Unlocking the Full Potential of Spatial Omics in Plants: Practical Challenges, Solutions, and a Path Forward.

Spatial omics technologies are providing new opportunities for plant biology by enabling molecular profiling within structurally intact tissues, revealing spatially organised cell states, developmental gradients, and regulatory interactions. While spatial transcriptomics has driven early advances, the field is rapidly expanding toward integrated spatial multi-omics by combining single-cell and spatial transcriptomic, epigenomic, proteomic, and metabolomic data. These approaches offer new opportunities to study development, physiology, and plant biotic and abiotic interactions in spatially preserved cellular contexts. However, despite rapid adoption, the field remains constrained by plant-specific challenges when applying technologies largely developed for animal systems. Compared with animal systems, plant tissues pose additional challenges due to rigid cell walls, and diverse chemistries, complicating sample preparation, cell and subcellular segmentation, signal detection, and data integration. As a result, many studies rely on bespoke protocols and analysis pipelines that are often difficult to reproduce or generalise. Here, we provide a practical, solution-oriented synthesis of current bottlenecks across experimental and computational pipelines, highlight emerging strategies to overcome these limitations, and propose a roadmap for community-driven protocol sharing, benchmarking, and integration across spatial and multi-omics modalities. Addressing these challenges will be essential to establish spatial omics as a routine and scalable tool for plant biology.

Journal Article↗

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans↗

Initiating informatics and GIS support for a field investigation of Bioterrorism: The New Jersey anthrax experience.

BACKGROUND: The investigation of potential exposure to anthrax spores in a Trenton, New Jersey, mail-processing facility required rapid assessment of informatics needs and adaptation of existing informatics tools to new physical and information-processing environments. Because the affected building and its computers were closed down, data to list potentially exposed persons and map building floor plans were unavailable from the primary source. RESULTS: Controlling the effects of anthrax contamination required identification and follow-up of potentially exposed persons. Risk of exposure had to be estimated from the geographic relationship between work history and environmental sample sites within the contaminated facility. To assist in establishing geographic relationships, floor plan maps of the postal facility were constructed in ArcView Geographic Information System (GIS) software and linked to a database of personnel and visitors using Epi Info and Epi Map 2000. A repository for maintaining the latest versions of various documents was set up using Web page hyperlinks. CONCLUSIONS: During public health emergencies, such as bioterrorist attacks and disease epidemics, computerized information systems for data management, analysis, and communication may be needed within hours of beginning the investigation. Available sources of data and output requirements of the system may be changed frequently during the course of the investigation. Integrating data from a variety of sources may require entering or importing data from a variety of digital and paper formats. Spatial representation of data is particularly valuable for assessing environmental exposure. Written documents, guidelines, and memos important to the epidemic were frequently revised. In this investigation, a database was operational on the second day and the GIS component during the second week of the investigation.

Journal Article↗

[Use of satellites for public health purposes in tropical areas].

The epidemiological hallmark of the new millennium has been the emergence or recrudescence of transmissible diseases with high epidemic potential. Disease tracking is becoming an increasingly global task requiring implementation of more and more sophisticated control strategies and facilities for sustainable development. A promising initiative involves the use of satellite technology to monitor and forecast the spread of disease. The Health Early Warning System (HEWS) was designed based on successful application of satellite data in food programs as well as in other areas (e.g. weather, farming and fishing). The HEWS integrates data from communications, remote-sensing and positioning satellites. The purpose of this review is to present the main studies containing satellite data on public health in tropical areas. Satellite data has allowed development of more reactive epidemiological tracking networks better suited to increasing population mobility, correlation of environmental factors (vegetation index, rainfall and ocean surface color) with human, animal and insect factors in epidemiological studies and assessment of the role of such factors in the development or reappearance of disease. Satellite technology holds great promise for more efficient management of public health problems in tropical areas.

Cholera↗

A computerized maintenance management system's requirements for standard operating procedures.

From this review of the 6 aspects of opportunity for inconsistency to corrupt or skew the reliability of data, it becomes apparent why members of management must provide the standards of operation and use within the CMMS for their employees. The possibility of poor data integrity due to any one of these aspects may not be severe; however, the severity is compounded and inevitable when different aspects are combined. Responding to information collected through the CMMS can be effective only if the data are reliable. With SOPs, management has provided their personnel with the necessary tools to ensure department-wide consistency. Management cannot afford to allow any one [table: see text] individual to apply personal interpretations of the importance and requirements in their approach to using the CMMS. If this is permitted, the loss of integrity due to one individual's judgment grows rapidly when data are analyzed at the departmental level. Standard operating procedures go beyond creating a "how to" for the CMMS; they provide the critical elements for collecting responsible and reliable data.

Biomedical Engineering↗

Quantitative quality control in microarray experiments and the application in data filtering, normalization and false positive rate prediction.

Data preprocessing including proper normalization and adequate quality control before complex data mining is crucial for studies using the cDNA microarray technology. We have developed a simple procedure that integrates data filtering and normalization with quantitative quality control of microarray experiments. Previously we have shown that data variability in a microarray experiment can be very well captured by a quality score q(com) that is defined for every spot, and the ratio distribution depends on q(com). Utilizing this knowledge, our data-filtering scheme allows the investigator to decide on the filtering stringency according to desired data variability, and our normalization procedure corrects the q(com)-dependent dye biases in terms of both the location and the spread of the ratio distribution. In addition, we propose a statistical model for false positive rate determination based on the design and the quality of a microarray experiment. The model predicts that a lower limit of 0.5 for the replicate concordance rate is needed in order to be certain of true positives. Our work demonstrates the importance and advantages of having a quantitative quality control scheme for microarrays.

Algorithms↗

GenBank.

GenBank (R) is a comprehensive sequence database that contains publicly available DNA sequences for more than 119 000 different organisms, obtained primarily through the submission of sequence data from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

GenBank: update.

GenBank is a comprehensive database that contains publicly available DNA sequences for more than 140 000 named organisms, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin program and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in the UK and the DNA Data Bank of Japan helps ensure worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design↗

High-throughput DNA sequencing on a capillary array electrophoresis system.

A capillary array electrophoresis apparatus capable of running and analyzing 48 DNA sequencing samples simultaneously has been constructed. The instrument uses a replaceable sieving buffer and incorporates a convenient method for introducing the buffer into the capillaries. Data from laser-induced fluorescence are collected as four separate images, one for each optical channel. The integrated data analysis software employs an open architecture that allows use of any DNA base-calling algorithm. DNA sequencing runs are completed in approx. 1 hr (approximately 500 bases), and instrument turnaround time between runs is less than 15 min. Overall, the instrument throughput is on the order of 720 templates/day, or 360,000 bases/day.

Animals↗

Web-based image review and data acquisition for multiinstitutional research.

OBJECTIVE: In this article, we describe a user-friendly Web-based interface that allows review of images combined with integrated data collection and entry for use at multiple sites involved in a large multicenter research project. CONCLUSION: The Web-based system that we present uses a commercially available Internet browser and Web platform and allows automated data entry that can be easily uploaded into standard data analysis programs. The system simplifies the complex logistics of using multiple sites and reviewers for radiology research and can preserve human subject confidentiality. We tested the system using a large-scale multicenter cohort study of pelvic fracture-related hemorrhage (the "Evaluating Pelvic Hemorrhage" study). Program testing revealed seamless remote image interpretation and data acquisition.

Biomedical Research↗

Improving data systems about juvenile victimization in the United States.

OBJECTIVE: To suggest improvements to 13 data sets and systems that collect information about juvenile victimization in United States. METHOD: The suggestions were gathered from a variety of sources, including data system users and administrators, as well as a special meeting convened on the topic by the National Consortium on Children, Families and the Law in Washington, DC (December 2000). RESULTS: Key areas of improvement were identified for each of 13 US data systems and possible solutions were identified. CONCLUSIONS: This paper suggests three broad categories of improvements that apply to a number of data systems. First, data systems could expand the coverage of the systems to include more jurisdictions or other segments of the population. Second, in order to be more comprehensive and specific to child victimization, the systems need to create more specific data items, questions, or response categories. Finally, the data systems need to be modified to provide continuity and interrelationships among systems, either by using uniform definitions, or integrating data systems to facilitate the tracking of children across systems.

Adolescent↗