PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data repository”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Exploiting the potential of routine data to better understand the disease burden posed by allergic disorders.

The Department of Health and Scottish Executive are currently undertaking independent reviews of allergy services in England (and Wales) and Scotland. Each review will assess the disease burden posed by allergic problems, involving secondary analyses of routine National Health Service (NHS) datasets. Major suggestions for re-structuring and/or re-focusing the NHS efforts to better deal with allergic disease are anticipated. The UK has some of the best datasets of routine health data in the world, but despite their strengths, they have important limitations. These include gaps in data collection, particularly in relation to monitoring of Accident & Emergency and out-patient consultations, and in-patient prescribing, thereby resulting in considerable under-estimates of hospital workload. The current gaps in service monitoring are likely to under-estimate the burden and workload associated with allergic problems, particularly in secondary care. One major limitation of existing data sources is the general inability to link individual patient level data between different datasets. By unlocking this potential there are very considerable potential gains to be made. Data linkage techniques currently being developed in the UK offer exciting new possibilities of looking across the primary-, secondary- and tertiary-care interfaces and also assessing short-and long-term social and educational outcomes in relation to allergic disorders. The current reviews of allergy services being undertaken need to be cognisant of these inherent limitations of existing data sources and would do well to recommend strategic initiatives that could enhance the availability, accessibility and quality of these datasets. Ideally, this should include investment in central data repositories staffed by teams with the necessary technical and statistical expertise, which would also take responsibility for progressing data linkage capabilities.

Drug Prescriptions↗

An integrated biomedical knowledge extraction and analysis platform: using federated search and document clustering technology.

High content screening (HCS) requires time-consuming and often complex iterative information retrieval and assessment approaches to optimally conduct drug discovery programs and biomedical research. Pre- and post-HCS experimentation both require the retrieval of information from public as well as proprietary literature in addition to structured information assets such as compound libraries and projects databases. Unfortunately, this information is typically scattered across a plethora of proprietary bioinformatics tools and databases and public domain sources. Consequently, single search requests must be presented to each information repository, forcing the results to be manually integrated for a meaningful result set. Furthermore, these bioinformatics tools and data repositories are becoming increasingly complex to use; typically they fail to allow for more natural query interfaces. Vivisimo has developed an enterprise software platform to bridge disparate silos of information. The platform automatically categorizes search results into descriptive folders without the use of taxonomies to drive the categorization. A new approach to information retrieval for HCS experimentation is proposed.

Biomedical Research↗

Emergency Department Allies: a Web-based multihospital pediatric asthma tracking system.

OBJECTIVE: To describe the development of a Web-based multihospital pediatric asthma tracking system and present results from the initial 18-month implementation of patient tracking experience. DESIGN: The Emergency Department (ED) Allies tracking system is a secure, password-protected data repository. Use-case methodology served as the foundation for technical development, testing, and implementation. Seventy-seven data elements addressing sociodemographics, wheezing history, quality of life, triggers, and ED managment were included for each subject visit. SETTING: The ED Allies partners comprised 1 academic pediatric ED and 5 community EDs. POPULATION: Subjects with a physician diagnosis of asthma who presented to the ED for acute respiratory complaints composed the asthma group; subjects lacking a physician diagnosis of asthma but presenting with wheezing composed the wheezing group. RESULTS: The tracking-system development and implementation process included identification of data elements, system database and use case development, and delineation of screen features, system users, reporting functions, and help screens. For the asthma group, 2005 subjects with physician-diagnosed asthma were enrolled between July 15, 2002 and January 14, 2004. These subjects accounted for 2978 visits; 10.4% had > or = 3 visits. Persistent asthma was noted in 68% of the subjects. During the same time period, 1297 wheezing subjects with a total of 1628 ED visits (wheezing group) were entered into the tracking system. After enrollment, 57% of the subjects with > or = 1 subsequent ED visits received a physician diagnosis of asthma. CONCLUSIONS: Our sophisticated tracking system facilitated data collection and identified key intervention opportunities for a diverse ED wheezing population. A significant asthma burden was identified with significant rates of hospitalization, acute care visits and persistent asthma in 68% of subjects. The surveillance component provided important insights into health care issues of both asthmatic subjects and wheezing subjects, many of whom subsequently were diagnosed with asthma.

Asthma↗

Integrating Biobanking Into Conservation Practice: The Development and Impact of the EAZA Biobank.

Zoological biobanks are becoming essential tools in conservation, offering a means to preserve genetic material and support in situ population management amid accelerating biodiversity loss. With rapid advances in genomics, cryopreservation, and assisted reproduction technologies, biobanks enable a proactive approach to providing insurance against genetic erosion and facilitating future research, supplementation, and genetic rescue. However, to be effective, zoological biobanks must be purposefully designed, strategically integrated into conservation frameworks such as the Convention on Biological Diversity (CBD) Kunming-Montreal Global Biodiversity Framework (KMGBF), and regularly evaluated for coverage and impact. Using the EAZA Biobank as an example, we outline the structure, development, and collaborative foundations that have enabled its rapid growth, built on community support and conservation impact. Leveraging EAZA's institutional network and data-sharing platforms such as ZIMS, the Biobank employs a decentralized, four-hub model of zoological institutions storing samples. A gap analysis, integrating threat status, breeding programs, genomic data repositories, and phylogenetic diversity, highlights current sampling strengths and deficiencies and guides future collection priorities. The integration of specimen-specific genomic data and the EAZA Biobank Cryonetwork of institutions with expertise in storing and generating gametes and cell lines will expand the Biobank's role in population management and conservation. Zoological biobanks must now evolve alongside advances in biotechnology and genomics. Sample collection strategies should serve conservation needs and anticipate future applications in genomics, cryobiology, and conservation medicine, linking biospecimens with the wealth of data generated from them. This approach should be scalable beyond EAZA, forming the foundation of a global standardized biobanking framework. Ultimately, zoological biobanks are not merely repositories of the past-they are essential infrastructures shaping the future potential of species conservation.

EAZA↗

Genomic biomarkers for cancer assessment: implementation challenges for laboratory practice.

Genomic biomarkers are an emerging class of laboratory tests, which present special implementation challenges for clinical laboratory services, compared to conventional laboratory tests. These challenges, which include analytical, bioinformatics, bioethical, interpretation and commercialization issues, represent real obstacles to widespread implementation of these tests. Technical challenges include the capacity to detect and identify many different kinds of markers for different diseases in a short time period, capacity to identify simultaneously gene rearrangements, amplification, inhibition, deletions and replications. Bioinformatics challenges include rapid analysis of genomic data, as well as the cross reference to other genomic data, and to other laboratory tests. Bioethical issues relate to consent to retain and use genetic data, which may be obtained inadvertently during analysis for genomic markers. Interpretation challenges include observations that the particular genomic markers may not be independent variables, as other undetected genomic alterations could invalidate or alter genomic marker interpretation. Further, as early experience with predictive genetic markers for cancer has shown, proprietary commercial interests may conflict with public health values of identifying genomic markers in subject populations. Based on our 10 years of experience with genomic biomarkers, important implementation strategies for genomic markers include development of:Standard high throughput analyzers capable of detecting any alteration of any genomic variant at any time. Bioinformatics analysis online, coupled to stored patient data. Laboratory service framework that preserves confidentiality but integrates genomic data with other laboratory tests. Laboratory service framework, which links consents, genomic analysis, reports to both specimen and data repositories. Overall, the laboratory service challenges for genomic markers are to manage very large analytical sets and very large data sets in finite time with responsible interpretation, all within finite funding. To meet these challenges, implementation strategies beyond the one disease, one diagnosis, one genomic marker concept must begin now.

Biomarkers, Tumor↗

MADGE: scalable distributed data management software for cDNA microarrays.

MOTIVATION: The human genome project and the development of new high-throughput technologies have created unparalleled opportunities to study the mechanism of diseases, monitor the disease progression and evaluate effective therapies. Gene expression profiling is a critical tool to accomplish these goals. The use of nucleic acid microarrays to assess the gene expression of thousands of genes simultaneously has seen phenomenal growth over the past five years. Although commercial sources of microarrays exist, investigators wanting more flexibility in the genes represented on the array will turn to in-house production. The creation and use of cDNA microarrays is a complicated process that generates an enormous amount of information. Effective data management of this information is essential to efficiently access, analyze, troubleshoot and evaluate the microarray experiments. RESULTS: We have developed a distributable software package designed to track and store the various pieces of data generated by a cDNA microarray facility. This includes the clone collection storage data, annotation data, workflow queues, microarray data, data repositories, sample submission information, and project/investigator information. This application was designed using a 3-tier client server model. The data access layer (1st tier) contains the relational database system tuned to support a large number of transactions. The data services layer (2nd tier) is a distributed COM server with full database transaction support. The application layer (3rd tier) is an internet based user interface that contains both client and server side code for dynamic interactions with the user. AVAILABILITY: This software is freely available to academic institutions and non-profit organizations at http://www.genomics.mcg.edu/niddkbtc.

Database Management Systems↗

Wireless application for complex wound management.

This project was to develop a web based wireless system to be used by community care nurses. Wound care constitutes approximately one half of home health care nursing visits. The system was implemented with a digital camera and a handheld computer with a wireless connection to a web based server. A database was used to control access to the clinical cases and to serve as a data repository. When new images were uploaded to the server a wound expert was notified via pager. The images and data could be viewed from any computer with an internet connection.

Community Health Nursing↗

The National Institutes of Health Clinical Center Digital Imaging Network, Picture Archival and Communication System, and Radiology Information System.

In this work, we describe the digital imaging network (DIN), picture archival and communication system (PACS), and radiology information system (RIS) currently being implemented at the Clinical Center, National Institutes of Health (NIH). These systems are presently in clinical operation. The DIN is a redundant meshed network designed to address gigabit density and expected high bandwidth requirements for image transfer and server aggregation. The PACS projected workload is 5.0 TB of new imaging data per year. Its architecture consists of a central, high-throughput Digital Imaging and Communications in Medicine (DICOM) data repository and distributed redundant array of inexpensive disks (RAID) servers employing fiber-channel technology for immediate delivery of imaging data. On demand distribution of images and reports to clinicians and researchers is accomplished via a clustered web server. The RIS follows a client-server model and provides tools to order exams, schedule resources, retrieve and review results, and generate management reports. The RIS-hospital information system (HIS) interfaces include admissions, discharges, and transfers (ATDs)/demographics, orders, appointment notifications, doctors update, and results.

Hospital Information Systems↗

Chemical effects in biological systems (CEBS) object model for toxicology data, SysTox-OM: design and application.

MOTIVATION: The CEBS data repository is being developed to promote a systems biology approach to understand the biological effects of environmental stressors. CEBS will house data from multiple gene expression platforms (transcriptomics), protein expression and protein-protein interaction (proteomics), and changes in low molecular weight metabolite levels (metabolomics) aligned by their detailed toxicological context. The system will accommodate extensive complex querying in a user-friendly manner. CEBS will store toxicological contexts including the study design details, treatment protocols, animal characteristics and conventional toxicological endpoints such as histopathology findings and clinical chemistry measures. All of these data types can be integrated in a seamless fashion to enable data query and analysis in a biologically meaningful manner. RESULTS: An object model, the SysBio-OM (Xirasagar et al., 2004) has been designed to facilitate the integration of microarray gene expression, proteomics and metabolomics data in the CEBS database system. We now report SysTox-OM as an open source systems toxicology model designed to integrate toxicological context into gene expression experiments. The SysTox-OM model is comprehensive and leverages other open source efforts, namely, the Standard for Exchange of Nonclinical Data (http://www.cdisc.org/models/send/v2/index.html) which is a data standard for capturing toxicological information for animal studies and Clinical Data Interchange Standards Consortium (http://www.cdisc.org/models/sdtm/index.html) that serves as a standard for the exchange of clinical data. Such standardization increases the accuracy of data mining, interpretation and exchange. The open source SysTox-OM model, which can be implemented on various software platforms, is presented here. AVAILABILITY: A universal modeling language (UML) depiction of the entire SysTox-OM is available at http://cebs.niehs.nih.gov and the Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. Currently, the public toxicological data in CEBS can be queried via a web application based on the SysTox-OM at http://cebs.niehs.nih.gov CONTACT: xirasagars@saic.com SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

Using the global proteome machine for protein identification.

This chapter describes the use of an open-source, freely available informatics system for the identification of proteins using tandem mass spectra of peptides derived from an enzymatic digest of a mixture of mature proteins. The chapter describes the use of features of the Global Proteome Machine (GPM) interface that assist in making comprehensive assignments between spectra and sequences, including the detection of point mutations, posttranslational modifications, and experimental artifacts. The use of this interface to validate results using the GPM Database is also described. This data repository allows analysts to compare their own results to those obtained by other scientists to determine the degree to which their data are consistent with previous measurements.

Amino Acid Sequence↗

National Bioterrorism Syndromic Surveillance Demonstration Program.

The National Bioterrorism Syndromic Surveillance Demonstration Program identifies new cases of illness from electronic ambulatory patient records. Its goals are to use data from health plans and practice groups to detect localized outbreaks and to facilitate rapid public health follow-up. Data are extracted nightly on patient encounters occurring during the previous 24 hours. Visits or calls with diagnostic codes corresponding to syndromes of interest are counted; repeat encounters are excluded. Daily counts of syndromes by zip code are sent to a central data repository, where they are statistically analyzed for unusual clustering by using a model-adjusted SaTScan approach. The results and raw data are displayed on a restricted website. Patient-level information stays at the originating health-care organization unless required by public health authorities. If a cluster surpasses a threshold of statistical aberration chosen by the corresponding public health department, an electronic alert can be sent to that department. The health department might then call a clinical responder, who has electronic access to records of cases contributing to clusters. The system is flexible, allowing for changes in participating organizations, syndrome definitions, and alert thresholds. It is transparent to clinicians and has been accepted by the health-care organizations that provide the data. The system's data are usable by local and national health agencies. Its software is compatible with commonly used systems and software and is mostly open-source. Ongoing activities include evaluating the system's ability to detect naturally occurring outbreaks and simulated terrorism events, automating and testing alerts and response capability, and evaluating alternative data sources.

Ambulatory Care↗

PlasmID: a centralized repository for plasmid clone information and distribution.

The Plasmid Information Database (PlasmID; http://plasmid.hms.harvard.edu) was developed as a community-based resource portal to facilitate search and request of plasmid clones shared with the Dana-Farber/Harvard Cancer Center (DF/HCC) DNA Resource Core. PlasmID serves as a central data repository and enables researchers to search the collection online using common gene names and identifiers, keywords, vector features, author names and PubMed IDs. As of October 2006, the repository contains >46 000 plasmids in 98 different vectors, including cloned cDNA and genomic fragments from 26 different species. Moreover, the clones include plasmid vectors useful for routine and cutting-edge techniques; functionally related sets of human cDNA clones; and genome-scale gene collections for Saccharomyces cerevisiae, Pseudomonas aeruginosa, Yersinia pestis, Francisella tularensis, Bacillus anthracis and Vibrio cholerae. Information about the plasmids has been fully annotated in adherence with a high-quality standard, and clone samples are stored as glycerol stocks in a state-of-the-art automated -80 degrees C freezer storage system. Clone replication and distribution is highly automated to minimize human error. Infor-mation about vectors and plasmid clones, including downloadable maps and sequence data, is freely available online. Researchers interested in requesting clone samples or sharing their own plasmids with the repository can visit the PlasmID website for more information.

Biological Specimen Banks↗

A consensus guide to preclinical indirect calorimetry experiments.

Understanding the complex factors influencing mammalian metabolism and body weight homeostasis is a long-standing challenge requiring knowledge of energy intake, absorption and expenditure. Using measurements of respiratory gas exchange, indirect calorimetry can provide non-invasive estimates of whole-body energy expenditure. However, inconsistent measurement units and flawed data normalization methods have slowed progress in this field. This guide aims to establish consensus standards to unify indirect calorimetry experiments and their analysis for more consistent, meaningful and reproducible results. By establishing community-driven standards, we hope to facilitate data comparison across research datasets. This advance will allow the creation of an in-depth, machine-readable data repository built on shared standards. This overdue initiative stands to markedly improve the accuracy and depth of efforts to interrogate mammalian metabolism. Data sharing according to established best practices will also accelerate the translation of basic findings into clinical applications for metabolic diseases afflicting global populations.

Calorimetry, Indirect↗

Decision support systems to manage outcomes.

It is essential that outcomes are evaluated and actions are taken to resolve problematic decision support issues. Many organizations are planning a transition to clinical information systems that will have data repositories and integrated databases within their systems. These transitions take years to complete, and systems must be put into place to collect data and report outcomes during the interim period. Careful and strategic planning is important so that the most effective use of existing technology and the most appropriate systems are implemented to manage outcomes. Multidisciplinary teams that use the most efficient technologic approaches will best meet the challenge to provide the most cost-effective care with the best quality outcomes.

Decision Support Techniques↗

Romanian perspective on health reporting.

Since 1990 the Romanian healthcare system has been crossing a stage of dramatic change. The healthcare reform, which is in progress, has structural, organizational and functional implications at the country level. The Ministry of Health is now preparing for the big IT changes. A Healthcare Management Information System (HMIS) will assure in the very next future automate data collection on health status of the population and on resource allocation and consumption. It is not an easy way towards the use of the Data Warehouse concept and On Line Analytical Processing technology for creating and accessing the data repository. Therefore we decided to integrate existing applications related to health system indicators until the HMIS will assure a high performance analysis and data interpretation.

Data Interpretation, Statistical↗

Data mining the NCI cancer cell line compound GI(50) values: identifying quinone subtypes effective against melanoma and leukemia cell classes.

Using data mining techniques, we have studied a subset (1400) of compounds from the large public National Cancer Institute (NCI) compounds data repository. We first carried out a functional class identity assignment for the 60 NCI cancer testing cell lines via hierarchical clustering of gene expression data. Comprised of nine clinical tissue types, the 60 cell lines were placed into six classes-melanoma, leukemia, renal, lung, and colorectal, and the sixth class was comprised of mixed tissue cell lines not found in any of the other five classes. We then carried out supervised machine learning, using the GI(50) values tested on a panel of 60 NCI cancer cell lines. For separate 3-class and 2-class problem clustering, we successfully carried out clear cell line class separation at high stringency, p < 0.01 (Bonferroni corrected t-statistic), using feature reduction clustering algorithms embedded in RadViz, an integrated high dimensional analytic and visualization tool. We started with the 1400 compound GI(50) values as input and selected only those compounds, or features, significant in carrying out the classification. With this approach, we identified two small sets of compounds that were most effective in carrying out complete class separation of the melanoma, non-melanoma classes and leukemia, non-leukemia classes. To validate these results, we showed that these two compound sets' GI(50) values were highly accurate classifiers using five standard analytical algorithms. One compound set was most effective against the melanoma class cell lines (14 compounds), and the other set was most effective against the leukemia class cell lines (30 compounds). The two compound classes were both significantly enriched in two different types of substituted p-quinones. The melanoma cell line class of 14 compounds was comprised of 11 compounds that were internal substituted p-quinones, and the leukemia cell line class of 30 compounds was comprised of 6 compounds that were external substituted p-quinones. Attempts to subclassify melanoma or leukemia cell lines based upon their clinical cancer subtype met with limited success. For example, using GI(50) values for the 30 compounds we identified as effective against all leukemia cell lines, we could subclassify acute lymphoblastic leukemia (ALL) origin cell lines from non-ALL leukemia origin cell lines without significant overlap from non-leukemia cell lines. Based upon clustering using GI(50) values for the 60 cancer cell lines laid out by the RadViz algorithm, these two compound subsets did not overlap with clusters containing any of the NCI's 92 compounds of known mechanism of action, a few of which are quinones. Given their structural patterns, the two p-quinone subtypes we identified would clearly be expected to possess different redox potentials/substrate specificities for enzymatic reduction in vivo. These two p-quinone subtypes represent valuable information that may be used in the elucidation of pharmacophores for the design of compounds to treat these two cancer tissue types in the clinic.

Algorithms↗

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility↗

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals↗