PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Automatic query mapping among genomic databases: a pilot exploration.

As databases in the human genome project proliferate, it is important for users of one genomic database to identify similar or inconsistent data in other autonomously developed genomic databases. To do so, the user needs to issue the same query across multiple databases. We describe an approach that allows a query issued against one database to be automatically mapped to an equivalent query against another structurally different database. Our approach features two components: 1) a database designed to capture knowledge (metadata) that describes the correspondences among individual database components and 2) a module that utilizes the metadata to perform query mappings. As a demonstration, we apply our query mapping approach to two chromosome map databases (DB/12 and GDB).

Algorithms↗

The rat liver epithelial (RLE) cell protein database.

Computer databases of rat liver epithelial (RLE) cellular polypeptides have been established using high resolution two-dimensional gel electrophoresis and computer-assisted analysis. Databases have been constructed utilizing both [35S]methionine- and [32P]orthophosphate-labeled as well as silver-stained polypeptides from normal RLE cells. The RLE database, which contains both qualitative and quantitative annotations, includes experiments with normal, chemically and oncogene transformed as well as spontaneously transformed cell lines. A total of 2537 [35S]methionine-labeled polypeptides from whole cell lysates (1920 acidic and 617 basic, separated in the first dimension using isoelectric focusing and nonequilibrium pH gradient electrophoresis, respectively) were analyzed and databases constructed using the Elsie 5 gel analysis system. To increase the "viewing window" and hence the usefulness of the RLE database, subcellular fractionation of whole cell preparations was performed and high resolution two-dimensional maps of the individual subcellular components were constructed. Databases representing 1229 cytosolic, 1539 acidic and 674 basic nuclear, 1746 membrane-associated, 415 mitochondrial, 773 in vitro translated and 350 phosphoproteins were established from these maps. The RLE databases contain the Elsie 5 identification number, protein name (if known), molecular weight and pI information, quantitative and spot shape data, and specific information regarding transformation-sensitive, growth-related (exponentially proliferating versus confluent) cell populations as well as those polypeptides modulated by specific growth factors. The RLE databases represent initial efforts toward the establishment of comprehensive databases of rat liver proteins and serve as a vital resource for on-going as well as future studies regarding the regulation of growth and differentiation as well as transformation of RLE cells.

Animals↗

An efficient disk based data structure for rapid searching of quantitative two-dimensional gel databases.

Fast access of two-dimensional (2-D) gel quantitative databases is important for rapid searching for protein differences between sets of 2-D gels from an experiment. The GELLAB-II system organizes corresponding spots from the gels in the database into reference or "Rspot" sets. These Rspot numeric names index fixed regions in the paged composite gel database file. This is adequate for an existing database, but has several problems. (i) Building the initial database requires guessing how much disk space to pre-allocate for each corresponding spot (i.e. spots from different gels). If it ever runs out of pre-allocated space during this process, it must expand the size of each corresponding set of spots copying the old database data into the new in-place on the disk. (ii) When adding new gels or editing the database, if a new spot is created, the system may also go into this expansion mode. The time spent and wasted disk space can be appreciable--depending on the size of the database (order of 100 gel database). (iii) Because each set of corresponding spots is the same size, we waste space in most spot sets since they do not require the additional space a few spot sets require which contain additional fragmented spots. We present a new low-level disk object-based structure and algorithm, paged indexed buckets (PIB), which optimizes disk space usage while having similar retrieval speed to the original method.

Algorithms↗

Construction of HSC-2DPAGE: a two-dimensional gel electrophoresis database of heart proteins.

The dissemination of information relating to the characterisation of proteins from two-dimensional electrophoresis (2-DE) gel databases is essential for their effective utilisation in the study of protein expression in cell biology. Since the inception of the World Wide Web and the pioneering development of SWISS-2DPAGE as a tool for retrieving information on proteins separated by 2-DE, the Internet has become the method of choice for disseminating and accessing information on 2-DE protein databases. At Harefield we have established HSC-2DPAGE which is an advanced interface for accessing protein database relating to heart disease. The Web site currently includes databases of proteins from human, dog and rat ventricular tissue and a human endothelial cell line. The databases are searchable individually or as a whole by remote keyword searches. Each database is represented by both synthetic (computer generated) and real (scanned gel) clickable images upon which characterised protein spots are highlighted by hyperlinked symbols. The database conforms to all the rules proposed for federated 2-DE protein databases and individual protein entries are linked to other protein databases such as SWISS-PROT by active cross-references. This paper describes the construction of HSC-2DPAGE, its maintenance, and access via the Internet.

Animals↗

The 2DWG meta-database of two-dimensional electrophoretic gel images on the Internet.

The 2DWG meta-database is a searchable database of two-dimensional (2-D) electrophoretic gel images found on the Internet. A meta-database contains information about locating data in other databases - but not that data itself. This database was constructed because of a need for an enriched set of World Wide Web (WWW) locations (URLs) of 2-D gel images on the Internet. These gel images are used in conjunction with the National Cancer Institute (NCI) Flicker Server to manipulate and visually compare 2-D gel images across the Internet. User's gels may also be compared with those in the database. The 2DWG is organized as a spreadsheet table with each gel image being represented by a row sorted by tissue type. Data for each gel includes tissue type, species, cell-line, image URL, database URL, gel protocol, organization URL, image properties, map URL if it exists, etc. The 2DWG may be searched to find relevant subsets of gels. Searching is done using the dbEngine - a WWW database search engine which accesses selected rows of gels from the full 2DWG table. The 2DWG meta-database is accessible on the WWW at http://www-lecb.ncifcrf.gov/2dwgDB/ and the NCI Flicker server at http://www-lecb.ncifcrf.gov/flicker/

Computer Communication Networks↗

The RESID Database of Protein Modifications as a resource and annotation tool.

The RESID Database of Protein Modifications is a comprehensive collection of annotations and structures for protein modifications and cross-links including pre-, co-, and post-translational modifications. The database provides: systematic and alternate names, atomic formulas and masses, enzymatic activities that generate the modifications, keywords, literature citations, Gene Ontology (GO) cross-references, protein sequence database feature table annotations, structure diagrams, and molecular models. This database is freely accessible on the Internet through resources provided by the European Bioinformatics Institute (http://www.ebi.ac.uk/RESID), and by the National Cancer Institute--Frederick Advanced Biomedical Computing Center (http://www.ncifcrf.gov/RESID). Each RESID Database entry presents a chemically unique modification and shows how that modification is currently annotated in the protein sequence databases, Swiss-Prot and the Protein Information Resource (PIR). The RESID Database provides a table of corresponding equivalent feature annotations that is used in the UniProt project, an international effort to combine the resources of the Swiss-Prot, TrEMBL and PIR. As an annotation tool, the RESID Database is used in standardizing and enhancing modification descriptions in the feature tables of Swiss-Prot entries. As an Internet resource, the RESID Database assists researchers in high-throughput proteomics to search monoisotopic masses and mass differences and identify known and predicted protein modifications.

Databases, Factual↗

Comparison of 35 electronic databases for environmental risk assessment.

The objective of this work is to classify and evaluate the major factual electronic databases that could be efficiently questioned for environmental risk assessment of various chemicals. A series of 35 databases available on commercial CD-ROMs and/or freely accessible on the Internet were listed and compared for the presence or absence of information for 27 environmental criteria. A factorial correspondence analysis indicated that most of the 35 databases are specialized in physicochemical, toxicological or ecotoxicological data but a few of them are nonspecialized databases able to answer simultaneously in various areas. We then evaluated these 35 databases by querying them for 14 selected test chemicals. It appeared that the percentage of chemicals being listed in the databases was very unequal and none of the 35 databases obtained 14 positive responses. The quality of the responses was determined by calculating the number of criteria being documented for the previously listed chemicals in each database. It suggested that HSDB, DOSE, TOMES, IPCS are the most efficient databases to be used for an urgent search of environmental data.

Databases, Factual↗

Descriptive study comparing routine hospital administrative data with the Vascular Society of Great Britain and Ireland's National Vascular Database.

OBJECTIVE: To compare patient volume and outcomes in vascular surgery between an administrative data set (Hospital Episode Statistics) and a clinical database (National Vascular Database). DESIGN: Descriptive study. METHODS: Volume of cases determined by age, sex, year and procedure and in-hospital mortality by procedure for both datasets for patients undergoing either repair of abdominal aortic aneurysm, carotid endarterectomy or infrainguinal bypass over a three year period between 1st April 2001 and 31st March 2004. RESULTS: There were 32,242 admissions with a mention of the three selected vascular procedures within the administrative data set compared to 8462 within the clinical database. For NHS trusts common to both datasets, there were twice as many procedures (16,923) recorded within the administrative dataset compared to the clinical database. Patient characteristics were similar across both databases. Further analysis limiting the administrative data to records attributed to consultants known to contribute to the clinical database showed much closer agreement with only 11% more repairs of abdominal aortic aneurysm recorded within the administrative dataset compared to the National Vascular Database. CONCLUSIONS: There are significant differences in total numbers between HES and the NVD. If the National Vascular Database is to become a credible source of information on activity and outcomes for vascular surgery, there is a clear need to increase the number of contributing surgeons and to increase the completeness of data submitted. Further analysis at individual record level is needed to identify other reasons for discrepancies which could help to enhance data quality, both within Hospital Episode Statistics and within the National Vascular Database.

Adult↗

A new retrieval system for a database of 3D facial images.

A new retrieval system for a 3D facial image database was designed and its reliability was experimentally examined. This system has two steps, firstly to automatically adjust the orientation of all 3D facial images in a database to that of the 2D facial image of a target person, and then to identify the facial image of the target person from the adjusted 3D facial images in the database using a graph-matching method. From the experimental study [M. Yoshino, K. Imaizumi, T. Tanijiri, J.G. Clement, Automatic adjustment of facial orientation in 3D face image database, Jpn. J. Sci. Tech. Iden. 8 (2003) 41-47], it is concluded that the software developed for the first step will be applicable to the automatic adjustment of facial orientation in the 3D facial image database. In 28 out of 110 sets (25.5%), the 3D image of the target person was chosen as the best match (from a database of 132 3D facial images) according to the similarity of the facial image characteristics based on the graph matching. The 3D facial image of the target person was ranked in the top of 10 of the database in 75 out of 110 sets (68.2%). These results suggest that this system is inadequate for the identification level, but may be feasible for screening method in a small database. It will be necessary to further pursue the possibility of realization of a facial image retrieval system for a large database such as suspects' facial images in future.

Adult↗

Comparison of the efficacy of an existing versus a locally developed metabolic fingerprint database to identify non-point sources of faecal contamination in a coastal lake.

A comparison of the efficacy of an existing large metabolic fingerprint database of enterococci and Escherichia coli with a locally developed database was undertaken to identify the sources of faecal contamination in a coastal lake, in southeast Qld., Australia. The local database comprised of 776 enterococci and 780 E. coli isolates from six host groups. In all, 189 enterococci and 245 E. coli biochemical phenotypes (BPTs) were found, of which 118 and 137 BPTs were unique (UQ) to host groups. The existing database comprised of 295 enterococci UQ-BPTs and 273 E. coli UQ-BPTs from 10 host groups. The representativeness and the stability of the existing database were assessed by comparing with isolates that were external to the database. In all, 197 enterococci BPTs and 179 E. coli BPTs were found in water samples. The existing database was able to identify 62.4% of enterococci BPTs and 64.8% of E. coli BPTs as human and animal sources. The results indicated that a representative database developed from a catchment can be used to predict the sources of faecal contamination in another catchment with similar landuse features within the same geographical area. However, the representativeness and the stability of the database should be evaluated prior to its application in such investigation.

Animals↗

Spinal palpation: The challenges of information retrieval using available databases.

PURPOSE: This study addressed 2 questions: first, what is the yield of PubMed MEDLINE for complementary and alternative medicine (CAM) studies compared to other databases; second, what is an effective search strategy to answer a sample research question on spinal palpation? METHODS: We formulated the following research question: "What is the reliability of spinal palpation procedures?" We identified specific Medical Subject Headings (MeSH) and key terms as used in osteopathic medicine, allopathic medicine, chiropractic, and physical therapy. Using PubMed, we formulated an initial search template and applied it to 12 additional selected databases. Subsequently, we applied the inclusion criteria and evaluated the yield in terms of precision and sensitivity in identifying relevant studies. RESULTS: The online search result of the 13 databases identified 1189 citations potentially addressing the research question. After excluding overlapping and nonpertinent citations and those not meeting the inclusion criteria, 49 citations remained. PubMed yielded 19, while MANTIS (Manual Alternative and Natural Therapy Index System), a manual therapy database, yielded 35 citations. Twenty-six of the 49 online citations were repeatedly indexed in 3 or more databases. Content experts and selective manual searches identified 11 additional studies. In all, we identified 60 studies that addressed the research question. The cost of the databases used for conducting this search ranged from free-of-charge to $43,000 per year for a single network subscription. CONCLUSIONS: Commonly used databases often do not provide accurate indexing or coverage of CAM publications. Subject-specific specialized databases are recommended. Access, cost, and ease of using specialized databases are limiting factors.

Abstracting and Indexing↗

A specialist toxicity database (TRACE) is more effective than its larger, commercially available counterparts.

The retrieval precision and recall of a specialist bibliographic toxicity database (TRACE) and a range of widely available bibliographic databases used to identify toxicity papers were compared. The analysis indicated that the larger size and resources of the major bibliographic databases did not, for a series of test queries, assure superior retrieval of relevant papers. The specialist database, in which document selection and indexing is undertaken by the same expert toxicologists who use the database in their day-to-day work, achieved markedly better retrieval, using simpler search strategies, than the other databases. Specialist databases may offer a valuable alternative to the existing major bibliographic databases. The concept of relevance, as used to determine the effectiveness of bibliographic databases, is discussed.

Data Collection↗

Representation of chemical information in OASIS centralized 3D database for existing chemicals.

The present inventory of existing chemicals in regulatory agencies in North America and Europe, encompassing the chemicals of the European Chemicals Bureau (EINECS, with 61 573 discrete chemicals); the Danish EPA (159 448 chemicals); the U.S. EPA (TSCA, 56 882 chemicals; HPVC, 10 546 chemicals) and pesticides' active and inactive ingredients of the U.S. EPA (1379 chemicals); the Organization for Economic Cooperation and Development (HPVC, 4750 chemicals); Environment Canada (DSL, 10851 chemicals); and the Japanese Ministry of Economy, Trade, and Industry (16811), was combined in a centralized 3D database for existing chemicals. The total number of unique chemicals from all of these databases exceeded 185 500. Defined and undefined chemical mixtures and polymers are handled, along with discrete (hydrolyzing and nonhydrolyzing) chemicals. The database manager provides the storage and retrieval of chemical structures with 2D and 3D data, accounting for molecular flexibility by using representative sets of conformers for each chemical. The electronic and geometric structures of all conformers are quantum-chemically optimized and evaluated. Hence, the database contains over 3.7 million 3D records with hundreds of millions of descriptor data items at the levels of structures, conformers, or atoms. The platform contains a highly developed search subsystem--a search is possible on Chemical Abstracts Service numbers; names; 2D and 3D fragment searches; structural, conformational, or atomic properties; affiliation in other chemical databases; structure similarity; logical combinations; saved queries; and search result exports. Models (collections of logically related descriptors) are supported, including information on a model's author, date, bioassay, organs/tissues, conditions, administration, and so forth. Fragments can be interactively constructed using a visual structure editor. A configurable database browser is designed for the inspection and editing of all types of data items. Database statistics are maintained on the number and quality of structures, conformers, and descriptors. Reports can be generated presenting any chosen subset of structures and descriptors into different formats suitable for inclusion into documents. In addition to fixed report formats, there is a powerful report template designer module with a visual report template editor to produce a customized page layout. The database is compatible at the import/export level with SDF, MOL, SMILES, and other known formats. The precalculated centralized 3D database could be useful for quantitative structure-activity relationship developers avoiding the time-consuming and cumbersome 3D calculation phase of model development.

Algorithms↗

Validation of the Italian food composition database of the European institute of oncology.

OBJECTIVE: To compare nutrient intakes obtained by chemical analysis of food composite or duplicate portion of diets with those obtained by weighed record method using the database of the European Institute of Oncology (EIO). SETTING: Nutrition Section, Department of Internal Medicine, University of Perugia, Italy. SUBJECTS: Fifteen subjects aged 40-59 y in 1960 (41 observations in three seasons), twenty-six subjects in 1965, and only nine remaining subjects in 1970 and 1991 were examined in Crevalcore. In Montegiorgio sixteen subjects aged 40-59 y in 1960 (39 observations in three seasons), thirty-two in 1965, twenty in 1970 and nine in 1991 were assessed. Forty-four subjects in Gubbio area (Biscina, Belvedere and Scritto; 21 males, 23 females; age 56.2+/-14.4 y) were evaluated in 1993 and 1994. METHODS: For dietary appraisal the individual weighed record method was used for 7, 3 or 2 days. Equivalent food composites were made up from local foodstuffs and the duplicate portions were chemically analysed for total nitrogen, fat, saturated and polyunsaturated fatty acids, carbohydrates, retinol, beta-carotene, thiamin and riboflavin. RESULTS: In Crevalcore, a significant difference for protein intake was found between analysis and calculation with EIO database in 1965 and 1991 (P<0.05). Fat intake was significant different for EIO database compared to analysis in 1965 survey (P<0.05), but not for other years. In Montegiorgio, there was a significant difference for protein intake between analysis and calculation with EIO database in 1970 and 1991 (both P<0.001). EIO database showed a significant difference in regard to analysis for fat intake in 1960 IV, 1965, 1970 and 1991 (P<0.05). In both areas there was a significant difference between analysis and EIO database for starch and fibre, but not for polyunsaturated fatty acids and soluble carbohydrates (all P<0.05). In Gubbio area, a significant difference was found between analysis and calculation with EIO database for fat, retinol, beta-carotene and riboflavin intakes (all P<0.05). CONCLUSIONS: According to previous and present studies food composition tables and databases, such as the EIO database, cannot be considered a reliable method to determine nutrient intakes, particularly for some vitamins.

Adult↗

The use of federal and state databases to conduct health services research related to physical and occupational therapy.

OBJECTIVE: To describe the characteristics of a number of secondary databases that have the potential to answer questions related to the use of, access to, and effectiveness of physical therapy (PT) and occupational therapy (OT). DATA SOURCES: Federal and state databases maintained by the National Center for Health Statistics, the Agency for Healthcare Research and Quality, and the Centers for Medicare and Medicaid Services. STUDY SELECTION: Databases, described above, that were identified as having potential for answering questions related to the use and effectiveness of PT and OT were examined. DATA EXTRACTION: The databases were explored to determine if PT and OT had sufficient representation, and if so, to identify potential questions that could be answered by examining the databases in more detail. Some of the advantages, disadvantages, and methodologic issues of using the databases were identified. DATA SYNTHESIS: Several databases are available that can be used by researchers to increase our understanding of the use of, access to, and/or effectiveness of PT and OT. Many of the databases are most suited for examining issues related to the use of and access to these services. A few of the databases can be used to examine the effectiveness of PT and OT. CONCLUSION: Secondary data analyses are a particularly useful, cost-effective, and efficient means for preliminary exploration of topics that are not well understood, such as the use of, access to, and effectiveness of PT and OT.

Data Collection↗

Randomized sequence databases for tandem mass spectrometry peptide and protein identification.

Tandem mass spectrometry (MS/MS) combined with database searching is currently the most widely used method for high-throughput peptide and protein identification. Many different algorithms, scoring criteria, and statistical models have been used to identify peptides and proteins in complex biological samples, and many studies, including our own, describe the accuracy of these identifications, using at best generic terms such as "high confidence." False positive identification rates for these criteria can vary substantially with changing organisms under study, growth conditions, sequence databases, experimental protocols, and instrumentation; therefore, study-specific methods are needed to estimate the accuracy (false positive rates) of these peptide and protein identifications. We present and evaluate methods for estimating false positive identification rates based on searches of randomized databases (reversed and reshuffled). We examine the use of separate searches of a forward then a randomized database and combined searches of a randomized database appended to a forward sequence database. Estimated error rates from randomized database searches are first compared against actual error rates from MS/MS runs of known protein standards. These methods are then applied to biological samples of the model microorganism Shewanella oneidensis strain MR-1. Based on the results obtained in this study, we recommend the use of use of combined searches of a reshuffled database appended to a forward sequence database as a means providing quantitative estimates of false positive identification rates of peptides and proteins. This will allow researchers to set criteria and thresholds to achieve a desired error rate and provide the scientific community with direct and quantifiable measures of peptide and protein identification accuracy as opposed to vague assessments such as "high confidence."

Databases, Protein↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

Databases and software for the analysis of mutations in the human p53 gene, the human hprt gene and the lacZ gene in transgenic rodents.

We have created databases and software applications for the analysis of DNA mutations in the human p53 gene, the human hprt gene and the rodent transgenic lacZ locus. The databases themselves are stand-alone dBase files and the software for analysis of the databases runs on IBM- compatible computers. The software created for these databases permits filtering, ordering, report generation and display of information in the database. In addition, a significant number of routines have been developed for the analysis of single base substitutions. One method of obtaining the databases and software is via the World Wide Web (WWW). Open home page http://sunsite.unc.edu/dnam/mainpage.ht ml with a WWW browser. Alternatively, the databases and programs are available via public ftp from anonymous@sunsite.unc.edu. There is no password required to enter the system. The databases and software are found in subdirectory pub/academic/biology/dna-mutations. Two other programs are available at the WWW site, a program for comparison of mutational spectra and a program for entry of mutational data into a relational database.

Animals↗