PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data repository”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Exploration and exploitation of clinical databases.

Clinical data repositories represent a potential gold mine of information and knowledge. Rapid access to such information can help bridge the gap between clinical care and research, support clinical and executive decision making, and improve the quality of care. A clinical database can be used in four ways: to display information about an individual patient (results reporting); to find data on a patient with similarities to one being seen (case finding); to describe a group of patients with at least one attribute in common (cohort description); and to analyze data patterns in terms of trends or relationships (predictive modeling). It seems unlikely that many important clinical questions will be subject to randomized clinical trials because of the ethics, logistics, and expense that would be involved. Evolving statistical and epidemiological methods allow us to approach these clinical data repositories with the purpose of building predictive models, but a clear understanding of the limitations of routinely collected clinical data and the inherent biases is necessary. The largest barrier to using routinely collected clinical data is not the limitations of the data themselves, but rather the lack of a data paradigm for the decision-maker. We present some of the problems and pitfalls in obtaining and using routinely collected data, based upon the use of ClinQuery at Boston's Beth Israel Hospital and the resources and traditions at the Mayo Clinic.

Databases, Factual↗

A review of interventions triggered by hepatitis A infected food-handlers in Canada.

BACKGROUND: In countries with low hepatitis A (HA) endemicity, infected food handlers are the source of most reported foodborne outbreaks. In Canada, accessible data repositories of infected food handler incidents are not available. We undertook a systematic review of such incidents to evaluate the extent of viral transmission through food contamination and the scope of post-exposure prophylaxis (PEP) interventions. METHODS: A systematic search of MEDLINE and EMBASE was conducted to identify published reports of incidents in Canada. An expanded search of a news repository (i.e., transcripts from newspapers and newscasts) was also conducted to identify the location and timing of an incident, which was used to retrieve the related report by contacting local public health departments. Data pertaining to case identification, public health risk, PEP interventions, and associated costs was independently abstracted by two reviewers and summarized according to incidents with and without large PEP interventions. RESULTS: A total of 16 incidents were identified from 1998-2004. There were approximately 3 incidents requiring public notification per year. Only 12.5% of incidents were described in published reports, indicating that published data significantly underestimated the number of incidents and PEP interventions. Data pertaining to the remaining incidents was unpublished, sparse and highly dispersed at the local public health level. Six of the 16 incidents required large PEP interventions to immunize on average 5000 potentially exposed individuals. Secondary transmission was low. Characteristics of incidents requiring large PEP interventions included potentially infectious food handlers working with uncooked food for a prolonged duration in high-volume grocery stores in high-density urban areas. CONCLUSION: Infected food handlers with hepatitis A virus (HAV) requiring public notification are not infrequent in Canada. Published data severely underestimated the burden of PEP intervention. Better and consistent reporting at the local and national level as well as a national data repository should be considered for the management of future incidents.

Canada↗

Using intermediate states to improve the ability of the Arden Syntax to implement care plans and reuse knowledge.

The Arden Syntax is one of a few knowledge representation languages currently in use for clinical decision support. While some of these languages are being used in active patient care settings, none have gained widespread acceptance as a clinical tool. Prior attempts to represent temporally complex care plans in the Arden Syntax have revealed difficulties in representing and tracking series of consecutive time-oriented events and recommendations, in sharing and reusing knowledge and in dealing with unobtainable data. In an attempt to improve Arden's ability to deal with these problems and demonstrate the importance of these factors, the clinical event monitor has been adapted to store coded data representing Intermediate States in the Columbia Presbyterian Medical Center (CPMC) central data repository. The Intermediate States define the current state of the patient as laid out in the care plan. Four care plans were constructed. The findings include an improved ability to track complex series of events and recommendations over long periods of time. The knowledge generated by the electronic care plans was able to be reused by the care plan that generated it, by other elements of the knowledge base and by non-decision support applications. Modular development, facilitated by the changes, simplified dealing with data not available to the central data repository by aiding the implementation of those parts of the care plan for which sufficient data is available.

Artificial Intelligence↗

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic↗

Predictors and impact of atrial fibrillation after isolated coronary artery bypass grafting.

OBJECTIVE: Although an extensive number of studies have attempted to identify predictors of new-onset atrial fibrillation (AFIB) after coronary artery bypass grafting (CABG), a strong predictive model does not exist. Prior studies have included patients recruited from multiple centers with variant AFIB prevalence rates and those who underwent CABG in combination with other surgical procedures. Also, most studies have focused on pre- and perioperative characteristics, with less attention given to the initial postoperative period. The purpose of this study was to comprehensively examine pre-, peri-, and postoperative characteristics that might predict new-onset AFIB in a large sample of patients undergoing isolated CABG in a single medical center, utilizing data readily available to clinicians in electronic data repositories. In addition, length of stay and selected postoperative complications and disposition were compared in patients with AFIB and no AFIB. DESIGN: Retrospective, comparative survey. SETTING: University-affiliated tertiary care hospital. PATIENTS: Patients with new-onset AFIB who underwent isolated standard CABG or minimally invasive direct vision coronary artery bypass were identified from an electronic clinical data repository. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: The prevalence of AFIB in the total sample (n = 814) was 31.9%. Predictors of AFIB included age (p =.0004), number of vessels bypassed (p =.013), vessel location (diagonal [p <.003] or posterior descending artery [p <.001]), and net fluid balance on the operative day (p =.015). Forward stepwise regression analysis produced a model that correctly predicted AFIB in only 24% of cases, with age (14%) and body surface area (9%) providing the most prediction. The incidence of embolic stroke was higher in AFIB (n = 8) vs. no AFIB (n = 4) patients, but stroke preceded AFIB onset in seven of eight cases. Subjects with AFIB had a longer stay (p =.0004), more intensive care unit readmissions (p =.0004), and required more assistance at hospital discharge (p =.017). CONCLUSIONS: Despite attempts to examine comprehensively predictors of new-onset AFIB, we were unable to identify a robust predictive model. Our findings, in combination with prior work, imply that it may not be feasible to predict the development of new-onset AFIB after CABG using data readily available to the bedside clinician. In this sample, stroke was uncommon and, when it occurred, preceded AFIB in all but one case. As anticipated, AFIB increased length of stay, and patients with this complication required more assistance at discharge.

Aged↗

Evolution of a legacy system to a Web patient record server: leveraging investment while opening the system.

A layered system is under development to enhance our legacy system as a backend in a WEB-enabled system. Each layer of the system has defined functionality, leverages the investment in the layer below, and follows the strategy of reducing support requirements for workstations. The mainframe system provides administrative integration of sub-systems, security, and the central data repository for most information. The second layer is a graphical user interface (GUI) to the system for Windows platforms. Support needs are limited by relying chiefly on X-terminals and application servers. The "Intranet" layer is a WEB Server building upon the second layer gateways to provide platform-independent access to selected information and images. The fourth layer, under evaluation, will extend access to the central data repository for Internet users of web browsers that support private-key/public-key encryption.

Computer Communication Networks↗

A simple, focused, computerized query to detect overutilization of laboratory tests.

CONTEXT: Although there is nearly universal agreement that laboratory tests are overutilized, the degree of overutilization in a given institution is difficult to quantify and monitor across time. OBJECTIVE: To detect and clearly document repetitive daily ordering of a commonly ordered laboratory test (serum sodium) by employing a simple, focused, computerized query of a test result database followed by chart review and validation. DESIGN: A retrospective computerized query of our clinical data repository was performed to find inpatients who displayed normal serum sodium test results on 4 or more consecutive days, without any abnormal values during the same admission. The search was limited to a 1-month period. A subset of these patients was selected for chart review. RESULTS: One hundred sixteen patients met our criteria, and the tests ordered for those patients comprised 5.1% of the monthly volume of serum sodium tests ordered in our institution. Chart review revealed a consistent lack of documentation of medical necessity for repeat testing as well as persistence of repeat serum sodium orders until the end of the patients' hospital course. CONCLUSIONS: We conclude that a focused query of data derived from a clinical data repository can detect and document overutilization of a common laboratory test in a convincing fashion within a given institution.

Clinical Laboratory Techniques↗

An object model and database for functional genomics.

MOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site.

Abstracting and Indexing↗

Mouse Phenome Database (MPD).

The Mouse Phenome Database (MPD; http://www.jax.org/phenome) is a repository of phenotypic and genotypic data on commonly used and genetically diverse inbred strains of mice. Strain characteristics data are contributed by members of the scientific community. Electronic access to centralized strain data enables biomedical researchers to choose appropriate strains for many systems-based research applications, including physiological studies, drug and toxicology testing and modeling disease processes. MPD provides a community data repository and a platform for data analysis and in silico hypothesis testing. The laboratory mouse is a premier genetic model for understanding human biology and pathology; MPD facilitates research that uses the mouse to identify and determine the function of genes participating in normal and disease pathways.

Animals↗

The importance of Java and CORBA in medicine.

One of the most powerful tools available for telemedicine is a multimedia medical record accessible over a wide area and simultaneously editable by multiple physicians. The ability to do this through an intuitive interface linking multiple distributed data repositories while maintaining full data integrity is a fundamental enabling technology in healthcare. We discuss the role of distributed object technology using Java and CORBA in providing this capability including an example of such a system (TeleMed) which can be accessed through the World Wide Web. Issues of security, scalability, data integrity, and usability are emphasized.

Computer Communication Networks↗

The cancer registry: a clinical repository of oncology data.

Health care institutions need complete and accurate data to plan, monitor, and evaluate their oncology programs. Although financial and discharge data are available, clinical repositories generally are not. For oncology, the cancer registry database serves as a clinical repository. The data in the registry are complete, accurate, and readily available. They can be used to plan new services, evaluate existing programs, and monitor patient care.

Community Health Planning↗

Model-driven user interfaces for bioinformatics data resources: regenerating the wheel as an alternative to reinventing it.

BACKGROUND: The proliferation of data repositories in bioinformatics has resulted in the development of numerous interfaces that allow scientists to browse, search and analyse the data that they contain. Interfaces typically support repository access by means of web pages, but other means are also used, such as desktop applications and command line tools. Interfaces often duplicate functionality amongst each other, and this implies that associated development activities are repeated in different laboratories. Interfaces developed by public laboratories are often created with limited developer resources. In such environments, reducing the time spent on creating user interfaces allows for a better deployment of resources for specialised tasks, such as data integration or analysis. Laboratories maintaining data resources are challenged to reconcile requirements for software that is reliable, functional and flexible with limitations on software development resources. RESULTS: This paper proposes a model-driven approach for the partial generation of user interfaces for searching and browsing bioinformatics data repositories. Inspired by the Model Driven Architecture (MDA) of the Object Management Group (OMG), we have developed a system that generates interfaces designed for use with bioinformatics resources. This approach helps laboratory domain experts decrease the amount of time they have to spend dealing with the repetitive aspects of user interface development. As a result, the amount of time they can spend on gathering requirements and helping develop specialised features increases. The resulting system is known as Pierre, and has been validated through its application to use cases in the life sciences, including the PEDRoDB proteomics database and the e-Fungi data warehouse. CONCLUSION: MDAs focus on generating software from models that describe aspects of service capabilities, and can be applied to support rapid development of repository interfaces in bioinformatics. The Pierre MDA is capable of supporting common database access requirements with a variety of auto-generated interfaces and across a variety of repositories. With Pierre, four kinds of interfaces are generated: web, stand-alone application, text-menu, and command line. The kinds of repositories with which Pierre interfaces have been used are relational, XML and object databases.

Computational Biology↗

Ongoing development of two-dimensional polyacrylamide gel electrophoresis data standards.

We present an approach toward standardizing two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) data in support of developing a globally relevant proteomics consensus in order to provide more efficient database querying and data comparisons through the establishment of the necessary definitions and interdisciplinary reference fields for both the 2-D PAGE community, particularly in the proteomics area, and the clinical and experimental biological research communities, in general. This article covers the need for unifying the 2-D PAGE data through a common data repository, and its usefulness in data standards and data interoperability.

Databases, Protein↗

Barriers and solutions in implementing occupational health and safety services at a large nuclear weapons facility.

The Hanford Nuclear Reservation is one of the U.S. Department of Energy's largest nuclear weapons sites. The enormous changes experienced by Hanford over the last several years, as its mission has shifted from weapons production to cleanup, has profoundly affected its occupational health and safety services. Innovative programs and new initiatives hold promise for a safer workplace for the thousands of workers at Hanford and other DOE sites. However, occupational health and safety professionals continue to face multiple organizational, economic, and cultural challenges. A major problem identified during this review was the lack of coordination of onsite services. Because each health and safety program operates independently (albeit with the guidance of the Richland field operations office), many services are duplicative and the health and safety system is fragmented. The fragmentation is compounded by the lack of centralized data repositories for demographic and exposure data. Innovative measures such as a questionnaire-driven Employee Job Task Analysis linked to medical examinations has allowed the site to move from the inefficient and potentially dangerous administrative medical monitoring assignment to defensible risk-based assignments and could serve as a framework for improving centralized data management and service delivery.

Contract Services↗

DictyMOLD-a Dictyostelium discoideum genome browser database.

UNLABELLED: With the Dictyostelium Genome Project nearing completion, we initiated the construction of a data repository for all Dictyostelium discoideum genomic data. Up to now this database, called DictyMOLD (Dicty Map Of Linked Data), incorporates the recently completed D.discoideum chromosomes 1 and 2 sequences together with related annotations. To visualise maps, sequences and annotations and to provide access for the scientific community a perl-based browser was developed. AVAILABILITY: The DictyMOLD database is freely accessible via http://genome.imb-jena.de/dictyostelium/ CONTACT: gernot@imb-jena.de.

Animals↗

An informatics infrastructure for patient safety and evidence-based practice in home healthcare.

The informatics infrastructure for patient safety and evidence-based practice (EBP) in home healthcare comprises data acquisition methods, healthcare standards including standardized terminologies, data repositories and clinical event monitors, data-mining techniques, digital sources of evidence, and communication technologies. Although the components of an informatics infrastructure are available and applications that bring these components together to promote patient safety and enable EBP have demonstrated positive or promising results in the acute care setting, a number of challenges hinder implementation in home healthcare. Resolution of these challenges requires commitment and collaboration among key stakeholders.

Benchmarking↗

PEDRo: a database for storing, searching and disseminating experimental proteomics data.

BACKGROUND: Proteomics is rapidly evolving into a high-throughput technology, in which substantial and systematic studies are conducted on samples from a wide range of physiological, developmental, or pathological conditions. Reference maps from 2D gels are widely circulated. However, there is, as yet, no formally accepted standard representation to support the sharing of proteomics data, and little systematic dissemination of comprehensive proteomic data sets. RESULTS: This paper describes the design, implementation and use of a Proteome Experimental Data Repository (PEDRo), which makes comprehensive proteomics data sets available for browsing, searching and downloading. It is also serves to extend the debate on the level of detail at which proteomics data should be captured, the sorts of facilities that should be provided by proteome data management systems, and the techniques by which such facilities can be made available. CONCLUSIONS: The PEDRo database provides access to a collection of comprehensive descriptions of experimental data sets in proteomics. Not only are these data sets interesting in and of themselves, they also provide a useful early validation of the PEDRo data model, which has served as a starting point for the ongoing standardisation activity through the Proteome Standards Initiative of the Human Proteome Organisation.

Animals↗

The tissue microarray data exchange specification: a community-based, open source tool for sharing tissue microarray data.

BACKGROUND: Tissue Microarrays (TMAs) allow researchers to examine hundreds of small tissue samples on a single glass slide. The information held in a single TMA slide may easily involve Gigabytes of data. To benefit from TMA technology, the scientific community needs an open source TMA data exchange specification that will convey all of the data in a TMA experiment in a format that is understandable to both humans and computers. A data exchange specification for TMAs allows researchers to submit their data to journals and to public data repositories and to share or merge data from different laboratories. In May 2001, the Association of Pathology Informatics (API) hosted the first in a series of four workshops, co-sponsored by the National Cancer Institute, to develop an open, community-supported TMA data exchange specification. METHODS: A draft tissue microarray data exchange specification was developed through workshop meetings. The first workshop confirmed community support for the effort and urged the creation of an open XML-based specification. This was to evolve in steps with approval for each step coming from the stakeholders in the user community during open workshops. By the fourth workshop, held October, 2002, a set of Common Data Elements (CDEs) was established as well as a basic strategy for organizing TMA data in self-describing XML documents. RESULTS: The TMA data exchange specification is a well-formed XML document with four required sections: 1) Header, containing the specification Dublin Core identifiers, 2) Block, describing the paraffin-embedded array of tissues, 3)Slide, describing the glass slides produced from the Block, and 4) Core, containing all data related to the individual tissue samples contained in the array. Eighty CDEs, conforming to the ISO-11179 specification for data elements constitute XML tags used in the TMA data exchange specification. A set of six simple semantic rules describe the complete data exchange specification. Anyone using the data exchange specification can validate their TMA files using a software implementation written in Perl and distributed as a supplemental file with this publication. CONCLUSION: The TMA data exchange specification is now available in a draft form with community-approved Common Data Elements and a community-approved general file format and data structure. The specification can be freely used by the scientific community. Efforts sponsored by the Association for Pathology Informatics to refine the draft TMA data exchange specification are expected to continue for at least two more years. The interested public is invited to participate in these open efforts. Information on future workshops will be posted at http://www.pathologyinformatics.org (API we site).

Community Health Services↗