PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Database Management Systems”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

METIS: multiple extraction techniques for informative sentences.

SUMMARY: METIS is a web-based integrated annotation tool. From single query sequences, the PRECIS component allows users to generate structured protein family reports from sets of related Swiss-Prot entries. These reports may then be augmented with pertinent sentences extracted from online biomedical literature via support vector machine and rule-based sentence classification systems. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/metis/

Algorithms↗

JColorGrid: software for the visualization of biological measurements.

BACKGROUND: Two-dimensional data colourings are an effective medium by which to represent three-dimensional data in two dimensions. Such "color-grid" representations have found increasing use in the biological sciences (e.g. microarray 'heat maps' and bioactivity data) as they are particularly suited to complex data sets and offer an alternative to the graphical representations included in traditional statistical software packages. The effectiveness of color-grids lies in their graphical design, which introduces a standard for customizable data representation. Currently, software applications capable of generating limited color-grid representations can be found only in advanced statistical packages or custom programs (e.g. micro-array analysis tools), often associated with steep learning curves and requiring expert knowledge. RESULTS: Here we describe JColorGrid, a Java library and platform independent application that renders color-grid graphics from data. The software can be used as a Java library, as a command-line application, and as a color-grid parameter interface and graphical viewer application. Data, titles, and data labels are input as tab-delimited text files or Microsoft Excel spreadsheets and the color-grid settings are specified through the graphical interface or a text configuration file. JColorGrid allows both user graphical data exploration as well as a means of automatically rendering color-grids from data as part of research pipelines. CONCLUSION: The program has been tested on Windows, Mac, and Linux operating systems, and the binary executables and source files are available for download at http://jcolorgrid.ucsf.edu.

Biology↗

An experimental metagenome data management and analysis system.

The application of shotgun sequencing to environmental samples has revealed a new universe of microbial community genomes (metagenomes) involving previously uncultured organisms. Metagenome analysis, which is expected to provide a comprehensive picture of the gene functions and metabolic capacity for microbial communities, needs to be conducted in the context of a comprehensive data management and analysis system. We present in this paper IMG/M, an experimental metagenome data management and analysis system that is based on the Integrated Microbial Genomes (IMG) system. IMG/M provides tools and viewers for analyzing both metagenomes and isolate genomes individually or in a comparative context. IMG/M is available at http://img.jgi.doe.gov/m.

Bacterial Physiological Phenomena↗

The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networks.

UNLABELLED: Biological processes involve complex networks of interactions between molecules. Various large-scale experiments and curation efforts have led to preliminary versions of complete cellular networks for a number of organisms. To grapple with these networks, we developed TopNet-like Yale Network Analyzer (tYNA), a Web system for managing, comparing and mining multiple networks, both directed and undirected. tYNA efficiently implements methods that have proven useful in network analysis, including identifying defective cliques, finding small network motifs (such as feed-forward loops), calculating global statistics (such as the clustering coefficient and eccentricity), and identifying hubs and bottlenecks. It also allows one to manage a large number of private and public networks using a flexible tagging system, to filter them based on a variety of criteria, and to visualize them through an interactive graphical interface. A number of commonly used biological datasets have been pre-loaded into tYNA, standardized and grouped into different categories. AVAILABILITY: The tYNA system can be accessed at http://networks.gersteinlab.org/tyna. The source code, JavaDoc API and WSDL can also be downloaded from the website. tYNA can also be accessed from the Cytoscape software using a plugin.

Computer Graphics↗

Database techniques for biological materials & methods.

The Biological sciences produce an enormous research literature every year. Research papers are highly structured documents whose content is not captured using the traditional techniques of information retrieval: keywords and flat text. This is especially true of the Materials & Methods section of experimental papers. A great deal of highly structured information is packed into this section. It involves logical and temporal sequences of operations that combine and operate on materials using various instruments and depending on many parameters. We are designing and implementing databases that will allow this complex knowledge to be represented, stored in object-oriented databases and retrieved. We are developing an application of this technology called the Laboratory Notebook. This application is a software system that will contain personal laboratory information as well as have access to databases of Materials & Methods sections drawn from the literature.

Artificial Intelligence↗

[Current developments in DICOM and IHE].

The broadening use of imaging management systems in radiology and other disciplines is due in large part to the success of the DICOM standard (Digital Imaging and Communications in Medicine), which has been accepted worldwide for more than 10 years and meanwhile represents one of the most successful standards in medicine. The central intent of establishing the initiative "Integrating the Healthcare Enterprise" (IHE) was to ensure interoperability of different IT systems in medical processes. IHE is essentially based on widespread standards such as DICOM or HL7. This overview article briefly addresses the principles and organization of DICOM and IHE and describes the current developments in both domains.

Database Management Systems↗

A survey of current work in biomedical text mining.

The volume of published biomedical research, and therefore the underlying biomedical knowledge base, is expanding at an increasing rate. Among the tools that can aid researchers in coping with this information overload are text mining and knowledge extraction. Significant progress has been made in applying text mining to named entity recognition, text classification, terminology extraction, relationship extraction and hypothesis generation. Several research groups are constructing integrated flexible text-mining systems intended for multiple uses. The major challenge of biomedical text mining over the next 5-10 years is to make these systems useful to biomedical researchers. This will require enhanced access to full text, better understanding of the feature space of biomedical literature, better methods for measuring the usefulness of systems to users, and continued cooperation with the biomedical research community to ensure that their needs are addressed.

Abstracting and Indexing↗

Ligand-Info, searching for similar small compounds using index profiles.

MOTIVATION: The Ligand-Info system is based on the assumption that small molecules with similar structure have similar functional (binding) properties. The developed system enables a fast and sensitive index based search for similar compounds in large databases. Index profiles, constructed by averaging indexes of related molecules are used to increase the specificity of the search. The utilization of index profiles helps to focus on frequent, common features of a family of compounds. RESULTS: A Java-based tool for clustering and scanning of small molecules has been created. The tool can interactively cluster sets of molecules and create index profiles on the user side and automatically download similar molecules from a databases of 250 000 compounds. The results of the application of index profiles demonstrate that the profile based search strategy can increase the quality of the selection process. AVAILABILITY: The system is available at http://Ligand.Info. The application requires the Java Runtime Environment 1.4, which can be automatically installed during the first use on desktop systems, which support it. A standalone version of the program is available from the authors upon request.

Binding Sites↗

The AcCell series 2000 as a support system for training and evaluation in educational and clinical settings.

Providing effective training, retraining and evaluation programs, including proficiency testing programs, for cytoprofessionals is a challenge shared by many academic and clinical educators internationally. In cytopathology the quality of training has immediately transferable and critically important impacts on satisfactory performance in the clinical setting. Well-designed interactive computer-assisted instruction and testing programs have been shown to enhance initial learning and to reinforce factual and conceptual knowledge. Computer systems designed not only to promote diagnostic accuracy but to integrate and streamline work flow in clinical service settings are candidates for educational adaptation. The AcCell 2000 system, designed as a diagnostic screening support system, offers technology that is adaptable to educational needs during basic and in-service training as well as testing of screening proficiency in both locator and identification skills. We describe the considerations, approaches and applications of the AcCell 2000 system in education programs for both training and evaluation of gynecologic diagnostic screening proficiency.

Automation↗

Genescript: DNA sequence annotation pipeline.

UNLABELLED: Genescript uses a number of publicly available analysis programs to annotate a DNA sequence. It provides an integrated display of results from each program, and includes an evidence-based scoring system that gives informative summaries of predicted gene models. AVAILABILITY: Genescript is available for download from http:://tcag.bioinfo.sickkids.on.ca/genescript/

Computer Graphics↗

ClusterControl: a web interface for distributing and monitoring bioinformatics applications on a Linux cluster.

UNLABELLED: ClusterControl is a web interface to simplify distributing and monitoring bioinformatics applications on Linux cluster systems. We have developed a modular concept that enables integration of command line oriented program into the application framework of ClusterControl. The systems facilitate integration of different applications accessed through one interface and executed on a distributed cluster system. The package is based on freely available technologies like Apache as web server, PHP as server-side scripting language and OpenPBS as queuing system and is available free of charge for academic and non-profit institutions. AVAILABILITY: http://genome.tugraz.at/Software/ClusterControl

Computational Biology↗

Recognizing names in biomedical texts: a machine learning approach.

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective and efficient literature mining and knowledge discovery that can help biologists to gather and make use of the knowledge encoded in text documents. In order to make organized and structured information available, automatically recognizing biomedical entity names becomes critical and is important for information retrieval, information extraction and automated knowledge acquisition. RESULTS: In this paper, we present a named entity recognition system in the biomedical domain, called PowerBioNE. In order to deal with the special phenomena of naming conventions in the biomedical domain, we propose various evidential features: (1) word formation pattern; (2) morphological pattern, such as prefix and suffix; (3) part-of-speech; (4) head noun trigger; (5) special verb trigger and (6) name alias feature. All the features are integrated effectively and efficiently through a hidden Markov model (HMM) and a HMM-based named entity recognizer. In addition, a k-Nearest Neighbor (k-NN) algorithm is proposed to resolve the data sparseness problem in our system. Finally, we present a pattern-based post-processing to automatically extract rules from the training data to deal with the cascaded entity name phenomenon. From our best knowledge, PowerBioNE is the first system which deals with the cascaded entity name phenomenon. Evaluation shows that our system achieves the F-measure of 66.6 and 62.2 on the 23 classes of GENIA V3.0 and V1.1, respectively. In particular, our system achieves the F-measure of 75.8 on the "protein" class of GENIA V3.0. For comparison, our system outperforms the best published result by 7.8 on GENIA V1.1, without help of any dictionaries. It also shows that our HMM and the k-NN algorithm outperform other models, such as back-off HMM, linear interpolated HMM, support vector machines, C4.5, C4.5 rules and RIPPER, by effectively capturing the local context dependency and resolving the data sparseness problem. Moreover, evaluation on GENIA V3.0 shows that the post-processing for the cascaded entity name phenomenon improves the F-measure by 3.9. Finally, error analysis shows that about half of the errors are caused by the strict annotation scheme and the annotation inconsistency in the GENIA corpus. This suggests that our system achieves an acceptable F-measure of 83.6 on the 23 classes of GENIA V3.0 and in particular 86.2 on the "protein" class, without help of any dictionaries. We think that a F-measure of 90 on the 23 classes of GENIA V3.0 and in particular 92 on the "protein" class, can be achieved through refining of the annotation scheme in the GENIA corpus, such as flexible annotation scheme and annotation consistency, and inclusion of a reasonable biomedical dictionary. AVAILABILITY: A demo system is available at http://textmining.i2r.a-star.edu.sg/NLS/demo.htm. Technology license is available upon the bilateral agreement.

Abstracting and Indexing↗

AdOnco: a database for clinical and scientific documentation of head and neck oncology.

OBJECTIVES: The aim of this project was to design, develop, and implement a head and neck cancer computer database for clinical and scientific use. METHODS: A relational database based on Filemaker Pro 6.0 was developed and integrated into our local network. A precise and easy-to-handle interface should allow for a quick overview of the patient's oncological history and for optimized data acquisition. An automatic report function was integrated to enhance the quality of daily patient care. For evaluation purposes, statistical analysis functions were implemented. RESULTS: Over a 14-month period, 410 patient records were available through the local network. The automated report function and the well-organized screen desktop resulted in time-efficient and accurate patient care. Additionally, the quality of information presented to referring physicians increased notably. The statistical analysis data provided by the database were reliable and easy to export. CONCLUSIONS: We developed an oncology database for both clinical and scientific purposes and integrated it successfully into our patient documentation system. The combination of clinical and scientific features proved to be very effective in daily routine and research.

Algorithms↗

Marburger concept for computer-aided acquisition, processing and documentation of patient data in the intensive care unit.

The authors report on their experience with the computer-aided acquisition, processing and documentation of patient data in an intensive care unit. The goal of an effective data and information collection system in the intensive care unit is to make therapy, and the patients respond to it, recognizable and understandable through the clear and complete representation of the patients conditions. The focal point of the data documentation is the medical record. In the implementation of a computer-aided data documentation and processing system the application of one central PC is unpractical. Each bed will be equipped with its own PC, and all the individual PCs will be connected to one central PC, which functions as server, and to connected to another PC in the doctors room. The data will be managed in a powerful database-management-system and be stored on an optical-disk. The development of an effective man/machine interface is especially important. In addition we can promote the user acceptance in other ways, so as to ensure a concrete benefit for each user group.

Electronic Data Processing↗

Automated management of gene discovery projects.

UNLABELLED: The paper reports on programs, organization and Hypertext Mark-up Language (HTML) presentation provided by the Hierarchical Automated Gene Identification System (HAGIS) designed for managing sequencing projects. AVAILABILITY: On request from the authors. CONTACT: st23646@ggr.co.uk

Computational Biology↗

A model system for studying the integration of molecular biology databases.

MOTIVATION: Integration of molecular biology databases remains limited in practice despite its practical importance and considerable research effort. The complexity of the problem is such that an experimental approach is mandatory, yet this very complexity makes it hard to design definitive experiments. This dilemma is common in science, and one tried-and-true strategy is to work with model systems. We propose a model system for this problem, namely a database of genes integrating diverse data across organisms, and describe an experiment using this model. RESULTS: We attempted to construct a database of human and mouse genes integrating data from GenBank and the human and mouse genome-databases. We discovered numerous errors in these well-respected databases: approximately 15% of genes are apparently missing from the genome-databases; links between the sequence and genome-databases are missing for another 5-10% of the cases; about a third of likely homology links are missing between the genome-databases; 10-20% of entries classified as 'genes' are apparently misclassified. By using a model system, we were able to study the problems caused by anomalous data without having to face all the hard problems of database integration. CONTACT: nat@jax.org

Animals↗

KDE Bioscience: platform for bioinformatics analysis workflows.

Bioinformatics is a dynamic research area in which a large number of algorithms and programs have been developed rapidly and independently without much consideration so far of the need for standardization. The lack of such common standards combined with unfriendly interfaces make it difficult for biologists to learn how to use these tools and to translate the data formats from one to another. Consequently, the construction of an integrative bioinformatics platform to facilitate biologists' research is an urgent and challenging task. KDE Bioscience is a java-based software platform that collects a variety of bioinformatics tools and provides a workflow mechanism to integrate them. Nucleotide and protein sequences from local flat files, web sites, and relational databases can be entered, annotated, and aligned. Several home-made or 3rd-party viewers are built-in to provide visualization of annotations or alignments. KDE Bioscience can also be deployed in client-server mode where simultaneous execution of the same workflow is supported for multiple users. Moreover, workflows can be published as web pages that can be executed from a web browser. The power of KDE Bioscience comes from the integrated algorithms and data sources. With its generic workflow mechanism other novel calculations and simulations can be integrated to augment the current sequence analysis functions. Because of this flexible and extensible architecture, KDE Bioscience makes an ideal integrated informatics environment for future bioinformatics or systems biology research.

Biological Science Disciplines↗

A scalable mediator approach to process large biomedical 3-D images.

The Edinburgh Mouse Atlas is a spatial-temporal framework to store and analyze biological data including three-dimensional (3-D) images that relate to mouse embryo development. The purpose of the system is the analysis and querying of complex spatial patterns, in particular the patterns of gene activity during embryo development. The framework holds large 3-D gray level images and is implemented in part as an object-oriented database. In this paper, we propose a dynamic layered architecture, based on the mediator approach, for the design of a transparent and scalable distributed system which can process objects that can exceed 1 GB in size. The system's data are distributed and/or declustered across a number of image servers and are processed by specialized mediators.

Algorithms↗