PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

NLProt: extracting protein names and sequences from papers.

Automatically extracting protein names from the literature and linking these names to the associated entries in sequence databases is becoming increasingly important for annotating biological databases. NLProt is a novel system that combines dictionary- and rule-based filtering with several support vector machines (SVMs) to tag protein names in PubMed abstracts. When considering partially tagged names as errors, NLProt still reached a precision of 75% at a recall of 76%. By many criteria our system outperformed other tagging methods significantly; in particular, it proved very reliable even for novel names. Names encountered particularly frequently in Drosophila, such as white, wing and bizarre, constitute an obvious limitation of NLProt. Our method is available both as an Internet server and as a program for download (http://cubic.bioc.columbia.edu/services/NLProt/). Input can be PubMed/MEDLINE identifiers, authors, titles and journals, as well as collections of abstracts, or entire papers.

Algorithms↗

BioThesaurus: a web-based thesaurus of protein and gene names.

UNLABELLED: BioThesaurus is a web-based system designed to map a comprehensive collection of protein and gene names to protein entries in the UniProt Knowledgebase. Currently covering more than two million proteins, BioThesaurus consists of over 2.8 million names extracted from multiple molecular biological databases according to the database cross-references in iProClass. The BioThesaurus web site allows the retrieval of synonymous names of given protein entries and the identification of protein entries sharing the same names. AVAILABILITY: BioThesaurus is accessible for online searching at http://pir.georgetown.edu/iprolink/biothesaurus

Animals↗

CROPPER: a metagene creator resource for cross-platform and cross-species compendium studies.

BACKGROUND: Current genomic research methods provide researchers with enormous amounts of data. Combining data from different high-throughput research technologies commonly available in biological databases can lead to novel findings and increase research efficiency. However, combining data from different heterogeneous sources is often a very arduous task. These sources can be different microarray technology platforms, genomic databases, or experiments performed on various species. Our aim was to develop a software program that could facilitate the combining of data from heterogeneous sources, and thus allow researchers to perform genomic cross-platform/cross-species studies and to use existing experimental data for compendium studies. RESULTS: We have developed a web-based software resource, called CROPPER that uses the latest genomic information concerning different data identifiers and orthologous genes from the Ensembl database. CROPPER can be used to combine genomic data from different heterogeneous sources, allowing researchers to perform cross-platform/cross-species compendium studies without the need for complex computational tools or the requirement of setting up one's own in-house database. We also present an example of a simple cross-platform/cross-species compendium study based on publicly available Parkinson's disease data derived from different sources. CONCLUSION: CROPPER is a user-friendly and freely available web-based software resource that can be successfully used for cross-species/cross-platform compendium studies.

Animals↗

iProLINK: an integrated protein resource for literature mining.

The exponential growth of large-scale molecular sequence data and of the PubMed scientific literature has prompted active research in biological literature mining and information extraction to facilitate genome/proteome annotation and improve the quality of biological databases. Motivated by the promise of text mining methodologies, but at the same time, the lack of adequate curated data for training and benchmarking, the Protein Information Resource (PIR) has developed a resource for protein literature mining--iProLINK (integrated Protein Literature INformation and Knowledge). As PIR focuses its effort on the curation of the UniProt protein sequence database, the goal of iProLINK is to provide curated data sources that can be utilized for text mining research in the areas of bibliography mapping, annotation extraction, protein named entity recognition, and protein ontology development. The data sources for bibliography mapping and annotation extraction include mapped citations (PubMed ID to protein entry and feature line mapping) and annotation-tagged literature corpora. The latter includes several hundred abstracts and full-text articles tagged with experimentally validated post-translational modifications (PTMs) annotated in the PIR protein sequence database. The data sources for entity recognition and ontology development include a protein name dictionary, word token dictionaries, protein name-tagged literature corpora along with tagging guidelines, as well as a protein ontology based on PIRSF protein family names. iProLINK is freely accessible at http://pir.georgetown.edu/iprolink, with hypertext links for all downloadable files.

Computational Biology↗

Online genomics facilities in the new millennium.

The review begins by providing a brief typology of biological databases on the Internet, illustrated by examples of the most influential resources of each kind. We then take an insider look at one typical on-line genomic resource -- the yeast genome database hosted at the Munich Information Center for Protein Sequences (MIPS) -- and explain how and why it has evolved from a basic sequence repository to a multidomain knowledge base. The role of community efforts in curating and annotating genome data is discussed. The crucial role of data integration and interoperability in developing next-generation genomic facilities is underscored.

Animals↗

Identification of brassinosteroid-related genes by means of transcript co-response analyses.

The comprehensive systems-biology database (CSB.DB) was used to reveal brassinosteroid (BR)-related genes from expression profiles based on co-response analyses. Genes exhibiting simultaneous changes in transcript levels are candidates of common transcriptional regulation. Combining numerous different experiments in data matrices allows ruling out outliers and conditional changes of transcript levels. CSB.DB was queried for transcriptional co-responses with the BR-signalling components BRI1 and BAK1: 301 out of 9694 genes represented in the nasc0271 database showed co-responses with both genes. As expected, these genes comprised pathway-involved genes (e.g. 72 BR-induced genes), because the BRI1 and BAK1 proteins are required for BR-responses. But transcript co-response takes the analysis a step further compared with direct approaches because BR-related non BR-responsive genes were identified. Insights into networks and the functional context of genes are provided, because factors determining expression patterns are reflected in correlations. Our findings demonstrate that transcript co-response analysis presents a valuable resource to uncover common regulatory patterns of genes. Different data matrices in CSB.DB allow examination of specific biological questions. All matrices are publicly available through CSB.DB. This work presents one possible roadmap to use the CSB.DB resources.

Arabidopsis↗

DiscoverySpace: an interactive data analysis application.

DiscoverySpace is a graphical application for bioinformatics data analysis. Users can seamlessly traverse references between biological databases and draw together annotations in an intuitive tabular interface. Datasets can be compared using a suite of novel tools to aid in the identification of significant patterns. DiscoverySpace is of broad utility and its particular strength is in the analysis of serial analysis of gene expression (SAGE) data. The application is freely available online.

Animals↗

An agent-based system for re-annotation of genomes.

Genome annotation projects can produce incorrect results if they are based on obsolete data or inappropriate models. We have developed an automatic re-annotation system that uses agents to perform repetitive tasks and reports the results to the user. These tasks involve BLAST searches on biological databases (GenBank) and the use of detection tools (Genemark and Glimmer) to identify new open reading frames. Several agents execute these tools and combine their results to produce a list of open reading frames that is sent back to the user. Our goal was to reduce the manual work, executing most tasks automatically by computational tools. A prototype was implemented and validated using Mycoplasma pneumoniae and Haemophilus influenzae original annotated genomes. The results reported by the system identify most of new features present in the re-annotated versions of these genomes.

Computational Biology↗

Public services from the European Bioinformatics Institute.

The European Bioinformatics Institute (EBI) provides numerous free-of-charge, publicly available bioinformatics services that can be divided into the following categories: ftp downloads; data submissions processing and biological database production; access to query; analysis and retrieval systems and tools; user support; training and education and industry support through EBI's SME program. These services are all available at the website. It is imperative that EBI's data as well as the tools to analyse it efficiently are made available in a free and unambiguous way to the scientific community. An important part of the EBI's mission is to make this happen in a fast, reliable and efficient manner. This paper serves as a brief introduction to each of these services.

Computational Biology↗

Get ready to GO! A biologist's guide to the Gene Ontology.

The Gene Ontology (GO) project provides a controlled vocabulary to facilitate high-quality functional gene annotation for all species. Genes in biological databases are linked to GO terms, allowing biologists to ask questions about gene function in a manner independent of species. This tutorial provides an introduction for biologists to the GO resources and covers three of the most common methods of querying GO: by individual gene, by gene function and by using a list of genes. [For the sake of brevity, the term 'gene' is used throughout this paper to refer to genes and their products (proteins and RNAs). GO annotations are always based on the characteristics of gene products, even though it may be the gene that is cited in the annotation.].

Abstracting and Indexing↗

Access to DNA and protein databases on the Internet.

During the past year, the number of biological databases that can be queried via Internet has dramatically increased. This increase has resulted from the introduction of networking tools, such as Gopher and WAIS, that make it easy for research workers to index databases and make them available for on-line browsing. Biocomputing in the nineties will see the advent of more client/server options for the solution of problems in bioinformatics.

Amino Acid Sequence↗

Physiological and phylogenetic diversity of bacteria growing on resin acids.

Resin acids are tricyclic diterpenes which are synthesized by trees and are a major cause of toxicity of pulp mill effluents. Bacterial strains isolated from three different sources and which grow on resin acids were physiologically characterized. Eleven strains, representating distinct groups, were further characterized physiologically and phylogenetically. The isolates had distinct specificities for use, as growth substrates, of the different resin acids tested. The isolates also used fatty acids but were generally limited in use of other diverse substrates tested. According to their 16S rDNA sequences, the representative isolates are related to members of the genera, Sphingomonas, Zoogloea, Ralstonia, Burkholderia, Pseudomonas and Mycobacterium. Analysis of whole-cell fatty acid profiles generally supported those phylogenetic relationships. However, most of the isolated did not have high similarities to reference strains in the Microbial Identification System database of fatty acid profiles or in the Biolog database of substrate oxidation patterns. Described species of Sphingomonas, Zoolgoea, Burkholderia Pseudomonas, most closely related to the isolates we characterized, failed to grow on, or degrade, resin acids. We propose recognition of Zoogloea resiniphila sp. nov., Pseudomonas vancouverensis sp. nov., P. abietaniphila sp. nov. and P. multiresinivorans sp. nov.

Bacteria, Aerobic↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/), maintained at the European Bioinformatics Institute (EBI), incorporates, organizes and distributes nucleotide sequences from public sources. The database is a part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. The web-based tool, Webin, is the preferred system for individual submission of nucleotide sequences, including Third Party Annotation (TPA) and alignment data. Automatic submission procedures are used for submission of data from large-scale genome sequencing centres and from the European Patent Office. Database releases are produced quarterly. The latest data collection can be accessed via FTP, email and WWW interfaces. The EBI's Sequence Retrieval System (SRS) integrates and links the main nucleotide and protein databases as well as many other specialist molecular biology databases. For sequence similarity searching, a variety of tools (e.g. FASTA and BLAST) are available that allow external users to compare their own sequences against the data in the EMBL Nucleotide Sequence Database, the complete genomic component subsection of the database, the WGS data sets and other databases. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗

The Saccharomyces Genome Database-a history of ideas and accomplishments, 1994-2026.

The Saccharomyces Genome Database (SGD) is one of the longest-running and most consequential biological databases in the world. Founded in the early 1990s at Stanford University under the visionary leadership of David Botstein and developed under the long-term technical direction of J. Michael Cherry, SGD has served for more than three decades not only as the authoritative knowledge center for the budding yeast Saccharomyces cerevisiae, but also as the source for much of the fundamentals of eukaryotic biology. This history traces the arc of a remarkable intellectual and scientific project: beginning with the challenge of building the very first integrated eukaryotic genome database and evolving across 30 years into a global knowledge hub for genetics, functional genomics, and human disease research. The history is organized chronologically, with each section highlighting the central ideas, technical developments, and concrete accomplishments of that period.

Databases, Genetic↗

Cardio: a web-based knowledge resource of genes and proteins related to cardiovascular disease.

BACKGROUND: Cardiovascular disease (CVD) is the leading killer for human. In order to understand the linkage between cardiovascular diseases and genes or proteins, it is essential to construct a database to organize the body of knowledge. While the existing molecular biological databases focus on the sequence and structural aspects of biological macromolecules, i.e. DNAs, RNAs and proteins, Cardio is the web-based system we built to provide a knowledge environment with visual interface to integrate information about major cardiovascular diseases in relation to genes and proteins. METHODS: We collected the information from the web by using a group of software we developed and used a relational database management system to manage the information of these data. RESULTS: Cardio consists of six sections: GENE, PROTEIN, DISEASE, DRUG, LINKS, and REFERENCE. Each section contains relevant information about the topic. Using a phenotype-driven approach, we can identify genetic mechanisms underlying the physiology and pathophysiology of specific cardiovascular disease, such as atherosclerosis, hypertension, heart failure, and stroke. CONCLUSION: The website titled "Database of Genes and Proteins Related to Cardiovascular Disease (Cardio)" is available at along with additional information on cardiovascular disease, supplementary materials and related figures.

Cardiovascular Diseases↗

A combined approach to data mining of textual and structured data to identify cancer-related targets.

BACKGROUND: We present an effective, rapid, systematic data mining approach for identifying genes or proteins related to a particular interest. A selected combination of programs exploring PubMed abstracts, universal gene/protein databases (UniProt, InterPro, NCBI Entrez), and state-of-the-art pathway knowledge bases (LSGraph and Ingenuity Pathway Analysis) was assembled to distinguish enzymes with hydrolytic activities that are expressed in the extracellular space of cancer cells. Proteins were identified with respect to six types of cancer occurring in the prostate, breast, lung, colon, ovary, and pancreas. RESULTS: The data mining method identified previously undetected targets. Our combined strategy applied to each cancer type identified a minimum of 375 proteins expressed within the extracellular space and/or attached to the plasma membrane. The method led to the recognition of human cancer-related hydrolases (on average, approximately 35 per cancer type), among which were prostatic acid phosphatase, prostate-specific antigen, and sulfatase 1. CONCLUSION: The combined data mining of several databases overcame many of the limitations of querying a single database and enabled the facile identification of gene products. In the case of cancer-related targets, it produced a list of putative extracellular, hydrolytic enzymes that merit additional study as candidates for cancer radioimaging and radiotherapy. The proposed data mining strategy is of a general nature and can be applied to other biological databases for understanding biological functions and diseases.

Biomarkers, Tumor↗

Design and implementation of a qualitative simulation model of lambda phage infection.

MOTIVATION: Molecular biology databases hold a large number of empirical facts about many different aspects of biological entities. That data is static in the sense that one cannot ask a database 'What effect has protein A on gene B?' or 'Do gene A and gene B interact, and if so, how?'. Those questions require an explicit model of the target organism. Traditionally, biochemical systems are modelled using kinetics and differential equations in a quantitative simulator. For many biological processes however, detailed quantitative information is not available, only qualitative or fuzzy statements about the nature of interactions. RESULTS: We designed and implemented a qualitative simulation model of lambda phage growth control in Escherichia coli based on the existing simulation environment QSim. Qualitative reasoning can serve as the basis for automatic transformation of contents of genomic databases into interactive modelling systems that can reason about the relations and interactions of biological entities.

Algorithms↗

An evaluation of GO annotation retrieval for BioCreAtIvE and GOA.

BACKGROUND: The Gene Ontology Annotation (GOA) database http://www.ebi.ac.uk/GOA aims to provide high-quality supplementary GO annotation to proteins in the UniProt Knowledgebase. Like many other biological databases, GOA gathers much of its content from the careful manual curation of literature. However, as both the volume of literature and of proteins requiring characterization increases, the manual processing capability can become overloaded. Consequently, semi-automated aids are often employed to expedite the curation process. Traditionally, electronic techniques in GOA depend largely on exploiting the knowledge in existing resources such as InterPro. However, in recent years, text mining has been hailed as a potentially useful tool to aid the curation process. To encourage the development of such tools, the GOA team at EBI agreed to take part in the functional annotation task of the BioCreAtIvE (Critical Assessment of Information Extraction systems in Biology) challenge. BioCreAtIvE task 2 was an experiment to test if automatically derived classification using information retrieval and extraction could assist expert biologists in the annotation of the GO vocabulary to the proteins in the UniProt Knowledgebase. GOA provided the training corpus of over 9000 manual GO annotations extracted from the literature. For the test set, we provided a corpus of 200 new Journal of Biological Chemistry articles used to annotate 286 human proteins with GO terms. A team of experts manually evaluated the results of 9 participating groups, each of which provided highlighted sentences to support their GO and protein annotation predictions. Here, we give a biological perspective on the evaluation, explain how we annotate GO using literature and offer some suggestions to improve the precision of future text-retrieval and extraction techniques. Finally, we provide the results of the first inter-annotator agreement study for manual GO curation, as well as an assessment of our current electronic GO annotation strategies. RESULTS: The GOA database currently extracts GO annotation from the literature with 91 to 100% precision, and at least 72% recall. This creates a particularly high threshold for text mining systems which in BioCreAtIvE task 2 (GO annotation extraction and retrieval) initial results precisely predicted GO terms only 10 to 20% of the time. CONCLUSION: Improvements in the performance and accuracy of text mining for GO terms should be expected in the next BioCreAtIvE challenge. In the meantime the manual and electronic GO annotation strategies already employed by GOA will provide high quality annotations.

Animals↗