PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Database Management Systems”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Of truth and pathways: chasing bits of information through myriads of articles.

Knowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data.

Algorithms↗

The EBI SRS server-new features.

MOTIVATION: Here we report on recent developments at the EBI SRS server (http://srs.ebi.ac.uk). SRS has become an integration system for both data retrieval and sequence analysis applications. The EBI SRS server is a primary gateway to major databases in the field of molecular biology produced and supported at EBI as well as European public access point to the MEDLINE database provided by US National Library of Medicine (NLM). It is a reference server for latest developments in data and application integration. The new additions include: concept of virtual databases, integration of XML databases like the Integrated Resource of Protein Domains and Functional Sites (InterPro), Gene Ontology (GO), MEDLINE, Metabolic pathways, etc., user friendly data representation in 'Nice views', SRSQuickSearch bookmarklets. AVAILABILITY: SRS6 is a licensed product of LION Bioscience AG freely available for academics. The EBI SRS server (http://srs.ebi.ac.uk) is a free central resource for molecular biology data as well as a reference server for the latest developments in data integration.

Computer Communication Networks↗

Rice Annotation Database (RAD): a contig-oriented database for map-based rice genomics.

A contig-oriented database for annotation of the rice genome has been constructed to facilitate map-based rice genomics. The Rice Annotation Database has the following functional features: (i) extensive effort of manual annotations of P1-derived artificial chromosome/bacterial artificial chromosome clones can be merged at chromosome and contig-level; (ii) concise visualization of the annotation information such as the predicted genes, results of various prediction programs (RiceHMM, Genscan, Genscan+, Fgenesh, GeneMark, etc.), homology to expressed sequence tag, full-length cDNA and protein; (iii) user-friendly clone / gene query system; (iv) download functions for nucleotide, amino acid and coding sequences; (v) analysis of various features of the genome (GC-content, average value, etc.); and (vi) genome-wide homology search (BLAST) of contig- and chromosome-level genome sequence to allow comparative analysis with the genome sequence of other organisms. As of October 2004, the database contains a total of 215 Mb sequence with relevant annotation results including 30 000 manually curated genes. The database can provide the latest information on manual annotation as well as a comprehensive structural analysis of various features of the rice genome. The database can be accessed at http://rad.dna.affrc.go.jp/.

Chromosomes, Plant↗

[Networking and integrated disease management. Advantages and disadvantages from the medical point of view].

At the moment the terms "networking", "cost reduction" and "integrated disease management" are frequently discussed in all branches of the German health care system. Unfortunately there are different interpretations of these terms. "Integrated disease management" in the meaning of communication between clinical and outpatient health care has al ready existed for years now. Traditional ways of communication lead to information loss. Losing information is a reason for low cost effectiveness and a prolonged healing process directly harming the patient. A computer network may prevent information loss and may in crease the performance of data transfer. Different sides have al ready started networking, and it is now necessary to bundle the interests. This necessity has been recognized by the German legislative. To lead this project to success it is important to know and to fulfil some medical criteria. Defining and describing these conditions is the topic of this paper. Our special intent is to show that digital technique is necessary to improve cooperation among physicians.

Computer Communication Networks↗

MPSS: an integrated database system for surveying a set of proteins.

SUMMARY: We design and implement an integrated database system called 'multi-protein survey system' (MPSS), which provides a platform to retrieve information about many proteins at a time. This system integrates several important and widely used databases including SwissProt, TrEMBL, PDB and InterPro, plus useful references such as GO and KEGG to other databases. Users may submit a group of protein IDs, entry names, SwissProt/TrEMBL accession numbers or GenBank GIs through MPSS' web interface, and obtain protein annotation information from public databases and pre-computed molecular properties speedily. MPSS can also supply comprehensive information about query proteins, including 3D structures, domains, pathway, gene ontology and visual presentation of mapping to the GO tree and KEGG pathway, to provide an up-to-date view of available knowledge with regard to the structures and molecular functions of proteins under study. AVAILABILITY: MPSS is freely accessible at http://www.scbit.org/mpss/

Database Management Systems↗

Microarray data warehouse allowing for inclusion of experiment annotations in statistical analysis.

MOTIVATION: Microarray technology provides access to expression levels of thousands of genes at once, producing large amounts of data. These datasets are valuable only if they are annotated by sufficiently detailed experiment descriptions. However, in many databases a substantial number of these annotations is in free-text format and not readily accessible to computer-aided analysis. RESULTS: The Multi-Conditional Hybridization Intensity Processing System (M-CHIPS), a data warehousing concept, focuses on providing both structure and algorithms suitable for statistical analysis of a microarray database's entire contents including the experiment annotations. It addresses the rapid growth of the amount of hybridization data, more detailed experimental descriptions, and new kinds of experiments in the future. We have developed a storage concept, a particular instance of which is an organism-specific database. Although these databases may contain different ontologies of experiment annotations, they share the same structure and therefore can be accessed by the very same statistical algorithms. Experiment ontologies have not yet reached their final shape, and standards are reduced to minimal conventions that do not yet warrant extensive description. An ontology-independent structure enables updates of annotation hierarchies during normal database operation without altering the structure. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://www.dkfz.de/tbi/services/mchips

Algorithms↗

SCA db: spinocerebellar ataxia candidate gene database.

UNLABELLED: The positional candidate gene approach accelerates the discovery of genes involved in disease. However, the properties of such disease genes are very diverse and the sample size of known disease genes is too small and does not warrant success by the use of a machine-learning approach. A user-defined scoring system may thus help to determine the priority of candidate genes. Spinocerebellar ataxia (SCA) is a good model to test this approach because most SCA subtypes are caused by an expansion of short tandem repeats (STRs). The SCA db is a candidate gene database for SCA, which collected 3185 genes for 17 types of SCA. Those SCA subtypes that have known disease genes can be used as positive controls to optimize the parameters. The users may browse the candidate genes of a given SCA subtype by using the default parameters. The known disease genes were found to be the top three candidates using the default parameters. Alternatively, the users may score the candidate genes by changing the weight or the scores on the basis of their own working hypothesis. AVAILABILITY: This database is available at http://ymbc.ym.edu.tw/sca/

Database Management Systems↗

BioMoby extensions to the Taverna workflow management and enactment software.

BACKGROUND: As biology becomes an increasingly computational science, it is critical that we develop software tools that support not only bioinformaticians, but also bench biologists in their exploration of the vast and complex data-sets that continue to build from international genomic, proteomic, and systems-biology projects. The BioMoby interoperability system was created with the goal of facilitating the movement of data from one Web-based resource to another to fulfill the requirements of non-expert bioinformaticians. In parallel with the development of BioMoby, the European myGrid project was designing Taverna, a bioinformatics workflow design and enactment tool. Here we describe the marriage of these two projects in the form of a Taverna plug-in that provides access to many of BioMoby's features through the Taverna interface. RESULTS: The exposed BioMoby functionality aids in the design of "sensible" BioMoby workflows, aids in pipelining BioMoby and non-BioMoby-based resources, and ensures that end-users need only a minimal understanding of both BioMoby, and the Taverna interface itself. Users are guided through the construction of syntactically and semantically correct workflows through plug-in calls to the Moby Central registry. Moby Central provides a menu of only those BioMoby services capable of operating on the data-type(s) that exist at any given position in the workflow. Moreover, the plug-in automatically and correctly connects a selected service into the workflow such that users are not required to understand the nature of the inputs or outputs for any service, leaving them to focus on the biological meaning of the workflow they are constructing, rather than the technical details of how the services will interoperate. CONCLUSION: With the availability of the BioMoby plug-in to Taverna, we believe that BioMoby-based Web Services are now significantly more useful and accessible to bench scientists than are more traditional Web Services.

Biology↗

2HAPI: a microarray data analysis system.

SUMMARY: 2HAPI (version 2 of High density Array Pattern Interpreter) is a web-based, publicly-available analytical tool designed to aid researchers in microarray data analysis. 2HAPI includes tools for searching, manipulating, visualizing, and clustering the large sets of data generated by microarray experiments. Other features include association of genes with NCBI information and linkage to external data resources. Unique to 2HAPI is the ability to retrieve upstream sequences of co-regulated genes for promoter analysis using MEME (Multiple Expectation-maximization for Motif Elicitation) AVAILABILITY: 2HAPI is freely available at http://array.sdsc.edu. Users can try 2HAPI anonymously with pre-loaded data or they can register as a 2HAPI user and upload their data.

Algorithms↗

GLAD: a system for developing and deploying large-scale bioinformatics grid.

MOTIVATION: Grid computing is used to solve large-scale bioinformatics problems with gigabytes database by distributing the computation across multiple platforms. Until now in developing bioinformatics grid applications, it is extremely tedious to design and implement the component algorithms and parallelization techniques for different classes of problems, and to access remotely located sequence database files of varying formats across the grid. In this study, we propose a grid programming toolkit, GLAD (Grid Life sciences Applications Developer), which facilitates the development and deployment of bioinformatics applications on a grid. RESULTS: GLAD has been developed using ALiCE (Adaptive scaLable Internet-based Computing Engine), a Java-based grid middleware, which exploits the task-based parallelism. Two bioinformatics benchmark applications, such as distributed sequence comparison and distributed progressive multiple sequence alignment, have been developed using GLAD.

Computational Biology↗

Visualizations for taxonomic and phylogenetic trees.

MOTIVATION: Despite substantial efforts to develop and populate the back-ends of biological databases, front-ends to these systems often rely on taxonomic expertise. This research applies techniques from human-computer interaction research to the biodiversity domain. RESULTS: We developed an interactive node-link tool, TaxonTree, illustrating the value of a carefully designed interaction model, animation, and integrated searching and browsing towards retrieval of biological names and other information. Users tested the tool using a new, large integrated dataset of animal names with phylogenetic-based and classification-based tree structures. These techniques also translated well for a tool, DoubleTree, to allow comparison of trees using coupled interaction. Our approaches will be useful not only for biological data but as general portal interfaces.

Algorithms↗

Deposition of macromolecular structures.

Macromolecular structures are being determined at an increasing rate, and are of interest to a wide diversity of researchers. Depositing a macromolecular structure with the Protein Data Bank makes it readily available to the community. Accuracy, consistency and machine-readability of the data are essential, as are clear indications of quality, and sufficient information to allow non-experimentalists to interpret the data. Good-quality depositions are necessary to allow this to be achieved. The PDB's AutoDep system allows deposition and some preliminary automatic checking to take place at multiple sites, prior to full processing and release of the structure by the PDB. However, depositing a structure currently requires the manual entry of a large amount of information at the time of deposition. The data-harvesting approach will allow much more information to be deposited, without placing an additional burden on the depositor. Deposition-ready files will be generated automatically during the course of a structure-determination experiment. The additional information will allow improved validation procedures to be applied to the structures, and the data to be made more useful to the wider scientific community.

Database Management Systems↗

A completely automated CAD system for mass detection in a large mammographic database.

Mass localization plays a crucial role in computer-aided detection (CAD) systems for the classification of suspicious regions in mammograms. In this article we present a completely automated classification system for the detection of masses in digitized mammographic images. The tool system we discuss consists in three processing levels: (a) Image segmentation for the localization of regions of interest (ROIs). This step relies on an iterative dynamical threshold algorithm able to select iso-intensity closed contours around gray level maxima of the mammogram. (b) ROI characterization by means of textural features computed from the gray tone spatial dependence matrix (GTSDM), containing second-order spatial statistics information on the pixel gray level intensity. As the images under study were recorded in different centers and with different machine settings, eight GTSDM features were selected so as to be invariant under monotonic transformation. In this way, the images do not need to be normalized, as the adopted features depend on the texture only, rather than on the gray tone levels, too. (c) ROI classification by means of a neural network, with supervision provided by the radiologist's diagnosis. The CAD system was evaluated on a large database of 3369 mammographic images [2307 negative, 1062 pathological (or positive), containing at least one confirmed mass, as diagnosed by an expert radiologist]. To assess the performance of the system, receiver operating characteristic (ROC) and free-response ROC analysis were employed. The area under the ROC curve was found to be Az = 0.783 +/- 0.008 for the ROI-based classification. When evaluating the accuracy of the CAD against the radiologist-drawn boundaries, 4.23 false positives per image are found at 80% of mass sensitivity.

Algorithms↗

The development of digital library system for drug research information.

The sophistication of computer technology and information transmission on internet has made various cyber information repository available to information consumers. In the era of information super-highway, the digital library which can be accessed from remote sites at any time is considered the prototype of information repository. Using object-oriented DBMS, the very first model of digital library for pharmaceutical researchers and related professionals in Korea has been developed. The published research papers and researchers' personal information was included in the database. For database with research papers, 13 domestic journals were abstracted and scanned for full-text image files which can be viewed by Internet web browsers. The database with researchers' personal information was also developed and interlinked to the database with research papers. These database will be continuously updated and will be combined with world-wide information as the unique digital library in the field of pharmacy.

Database Management Systems↗

Integration of digital angiography and gamma-camera diagnostic modalities to a generalized hospital information system.

Within the framework of the NIKA project for the development of a "Generalized System for Processing and Management of Medical Images", we have integrated a Digital Subtraction Angiography (DSA) station and a gamma-camera diagnostic modality to the newly developed generalized hospital-wide information system. The integration consists of acquiring, digitizing, converting and filing the information from the above diagnostic modalities to the hospital's Picture Archiving and Communication System (PACS). The PACS, which is installed in the Onasio Cardiosurgery Centre of Athens, Greece, is responsible for archival, cataloging, retrieval and viewing of the large volume of patient examination data accumulated during his/her stay in the Centre. The ultimate goal is for all patient data, from all different examinations, to be viewed on dedicated client workstations. Information from different modalities can then be simultaneously presented to the attending physician, for a complete picture of the patient history and condition.

Angiography, Digital Subtraction↗

Career History Archival Medical and Personnel System.

The Career History Archival Medical and Personnel System is a database that provides information on cancer, chronic diseases, occupational and preventive medicine, epidemiological research, and the use of health care in the Navy and Marine Corps. It was created at the Naval Health Research Center for enlisted Navy personnel, and it is being expanded to encompass all military personnel. Its objective is to provide a comprehensive, chronologically ordered database of career and medical events in all active duty military service members and to track career and disease events in order from the date of entry to service to the date service ended. Events include the dates of beginning and ending of each specific military occupation, all assignments to a military units or ships, all hospitalized diseases, and other events. The database contains detailed epidemiological data on more than six million members of the military services. It is the largest known epidemiological database in the United States.

Archives↗

Enhanced perceptual distance functions and indexing for image replica recognition.

The proliferation of digital images and the widespread distribution of digital data that has been made possible by the Internet has increased problems associated with copyright infringement on digital images. Watermarking schemes have been proposed to safeguard copyrighted images, but watermarks are vulnerable to image processing and geometric distortions and may not be very effective. Thus, the content-based detection of pirated images has become an important application. In this paper, we discuss two important aspects of such a replica detection system: distance functions for similarity measurement and scalability. We extend our previous work on perceptual distance functions, which proposed the Dynamic Partial Function (DPF), and present enhanced techniques that overcome the limitations of DPF. These techniques include the Thresholding, Sampling, and Weighting schemes. Experimental evaluations show superior performance compared to DPF and other distance functions. We then address the issue of using these perceptual distance functions to efficiently detect replicas in large image data sets. The problem of indexing is made challenging by the high-dimensionality and the nonmetric nature of the distance functions. We propose using Locality Sensitive Hashing (LSH) to index images while using the above perceptual distance functions and demonstrate good performance through empirical studies on a very large database of diverse images.

Algorithms↗

Concept-oriented indexing of video databases: toward semantic sensitive retrieval and browsing.

Digital video now plays an important role in medical education, health care, telemedicine and other medical applications. Several content-based video retrieval (CBVR) systems have been proposed in the past, but they still suffer from the following challenging problems: semantic gap, semantic video concept modeling, semantic video classification, and concept-oriented video database indexing and access. In this paper, we propose a novel framework to make some advances toward the final goal to solve these problems. Specifically, the framework includes: 1) a semantic-sensitive video content representation framework by using principal video shots to enhance the quality of features; 2) semantic video concept interpretation by using flexible mixture model to bridge the semantic gap; 3) a novel semantic video-classifier training framework by integrating feature selection, parameter estimation, and model selection seamlessly in a single algorithm; and 4) a concept-oriented video database organization technique through a certain domain-dependent concept hierarchy to enable semantic-sensitive video retrieval and browsing.

Abstracting and Indexing↗