PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

HICLAS: a taxonomic database system for displaying and comparing biological classification and phylogenetic trees.

MOTIVATION: Numerous database management systems have been developed for processing various taxonomic data bases on biological classification or phylogenetic information. In this paper, we present an integrated system to deal with interacting classifications and phylogenies concerning particular taxonomic groups. RESULTS: An information-theoretic view (taxon view) has been applied to capture taxonomic concepts as taxonomic data entities. A data model which is suitable for supporting semantically interacting dynamic views of hierarchic classifications and a query method for interacting classifications have been developed. The concept of taxonomic view and the data model can also be expanded to carry phylogenetic information in phylogenetic trees. We have designed a prototype taxonomic database system called HICLAS (HIerarchical CLAssification System) based on the concept of taxon view, and the data models and query methods have been designed and implemented. This system can be effectively used in the taxonomic revisionary process, especially when databases are being constructed by specialists in particular groups, and the system can be used to compare classifications and phylogenetic trees. AVAILABILITY: Freely available at the WWW URL: http://aims.cps.msu.edu/hiclas/ CONTACT: pramanik@cps.msu.edu; lotus@wipm.whcnc.ac.cn

Classification↗

Computational space reduction and parallelization of a new clustering approach for large groups of sequences.

MOTIVATION: The explosive growth of the biological sequences databases stimulated by genome projects has modified the framework of several applications in the biological sequence analysis area. In most cases, this new scenario is characterized by studies on large sets of sequences, suggesting the need for effective and automatic methods for their clustering. A more effective clustering of the database could be followed by the application of common family analysis schemes to the groups so formed. RESULTS: In this work, we present a new strategy to reduce the computational cost associated with the clustering of large sets of sequences which are expected to contain several families. The strategy is based on the grouping of the sequences into families by using a dynamic threshold on a pairwise sequence similarity criterion. Routine clustering of large data sets can now be done very efficiently. The method developed here achieves a computational space reduction of about an order of magnitude over more traditional ones of all-versus-all comparisons. The outcome of this approach produces family groupings that reproduce closely already accepted biological results. Our work includes a parallel implementation for distributed memory multiprocessors with a dynamic scheduling strategy for performance optimization. AVAILABILITY: By anonymous ftp at ftp.ac.uma.es (/pub/ots/pCluster directory), or from our Web site http://www.cnb. uam.es/www/software/software_index.html CONTACT: ots@ac.uma.es

Algorithms↗

WU-Blast2 server at the European Bioinformatics Institute.

Since 1995, the WU-BLAST programs (http://blast.wustl.edu) have provided a fast, flexible and reliable method for similarity searching of biological sequence databases. The software is in use at many locales and web sites. The European Bioinformatics Institute's WU-Blast2 (http://www.ebi.ac.uk/blast2/) server has been providing free access to these search services since 1997 and today supports many features that both enhance the usability and expand on the scope of the software.

Computational Biology↗

DNA databases.

This paper presents DNA algorithms for five relational algebra database operations, selection, projection, union, set difference, and Cartesian product on so-called DNA databases. A DNA database is a database where data records are encoded as DNA strands. The five operations mentioned before are fundamental in the field of databases and perform most of the data retrieval operations on current databases.

Algorithms↗

Functional genomics databases on the web.

Experiments involving high-throughput methods for measuring transcripts, proteins and metabolites constitute the area of functional genomics. These experiments are highly context dependent and require much more detail about the experimental design, sample and protocols used than in genomics. Functional genomics databases are needed that follow established and emerging standards. Functional genomic databases are not yet very common; however, there are a few focused on microbial genomes and a couple integrative systems are available for setting up functional genomics databases.

Computational Biology↗

Database system to identify biological risk in managed care organizations: implications for clinical care.

The investigators constructed an index measure of cardiovascular risk and scored 1.991 adults as having high, average, or low cardiovascular risk. High cardiovascular risk was positively associated with hospital admissions (odds ratio [OR] = 3.9, p < 0.0001), total hospital days (OR = 4.0, p < 0.001), primary care clinic visits (OR = 7.3, p < 0.0001), and subspecialty clinic visits (OR = 2.3, p = 0.0003), compared to low cardiovascular risk, after controlling in multivariate analyses for gender and age. The index can provide estimates of utilization, costs, and potential preventability of adverse cardiovascular events, can be used to identify groups of patients in need of various systematic interventions, and can provide population-based ways to evaluate the results of interventions.

Adult↗

Predicting protein subcellular location by fusing multiple classifiers.

One of the fundamental goals in cell biology and proteomics is to identify the functions of proteins in the context of compartments that organize them in the cellular environment. Knowledge of subcellular locations of proteins can provide key hints for revealing their functions and understanding how they interact with each other in cellular networking. Unfortunately, it is both time-consuming and expensive to determine the localization of an uncharacterized protein in a living cell purely based on experiments. With the avalanche of newly found protein sequences emerging in the post genomic era, we are facing a critical challenge, that is, how to develop an automated method to fast and reliably identify their subcellular locations so as to be able to timely use them for basic research and drug discovery. In view of this, an ensemble classifier was developed by the approach of fusing many basic individual classifiers through a voting system. Each of these basic classifiers was trained in a different dimension of the amphiphilic pseudo amino acid composition (Chou [2005] Bioinformatics 21: 10-19). As a demonstration, predictions were performed with the fusion classifier for proteins among the following 14 localizations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cytoplasm, (5) cytoskeleton, (6) endoplasmic reticulum, (7) extracellular, (8) Golgi apparatus, (9) lysosome, (10) mitochondria, (11) nucleus, (12) peroxisome, (13) plasma membrane, and (14) vacuole. The overall success rates thus obtained via the resubstitution test, jackknife test, and independent dataset test were all significantly higher than those by the existing classifiers. It is anticipated that the novel ensemble classifier may also become a very useful vehicle in classifying other attributes of proteins according to their sequences, such as membrane protein type, enzyme family/sub-family, G-protein coupled receptor (GPCR) type, and structural class, among many others. The fusion ensemble classifier will be available at www.pami.sjtu.edu.cn/people/hbshen.

Amino Acids↗

Comparative interactomics.

The behavior, morphology and response to stimuli in biological systems are dictated by the interactions between their components. These interactions, as we observe them now, are therefore shaped by genetic variations and selective pressure. Similar to what has been achieved by comparing genome structures and protein sequences, we hope to obtain valuable information about systems' evolution by comparing the organization of interaction networks and by analyzing their variation and conservation. Equally, significantly we can learn whether and how to extend the network information obtained experimentally in well-characterized model systems to different organisms. We conclude from our analysis that, despite the recent completion of several high throughput experiments aimed at the description of complete interactomes, the available interaction information is not yet of sufficient coverage and quality to draw any biologically meaningful conclusion from the comparison of different interactomes. Thus, the transfer of network information obtained from simple organism to evolutionary distant species should be carried out and considered with caution. By using smaller higher-confidence datasets, a larger fraction of interactions is shown to be conserved; this suggests that with the development of more accurate experimental and informatic approaches, we will soon be in the position to study the network evolution.

Animals↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

Ensembl 2002: accommodating comparative genomics.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of human, mouse and other genome sequences, available as either an interactive web site or as flat files. Ensembl also integrates manually annotated gene structures from external sources where available. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. These range from sequence analysis to data storage and visualisation and installations exist around the world in both companies and at academic sites. With both human and mouse genome sequences available and more vertebrate sequences to follow, many of the recent developments in Ensembl have focusing on developing automatic comparative genome analysis and visualisation.

Animals↗

Ensembl 2004.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organize biology around the sequences of large genomes. It is a comprehensive and integrated source of annotation of large genome sequences, available via interactive website, web services or flat files. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. The facilities of the system range from sequence analysis to data storage and visualization and installations exist around the world both in companies and at academic sites. With a total of nine genome sequences available from Ensembl and more genomes to follow, recent developments have focused mainly on closer integration between genomes and external data.

Animals↗