PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

BLAST: improvements for better sequence analysis.

Basic local alignment search tool (BLAST) is a sequence similarity search program. The National Center for Biotechnology Information (NCBI) maintains a BLAST server with a home page at http://www.ncbi.nlm.nih.gov/BLAST/. We report here on recent enhancements to the results produced by the BLAST server at the NCBI. These include features to highlight mismatches between similar sequences, show where the query was masked for low-complexity sequence, and integrate information about the database sequences from the NCBI Entrez system into the BLAST display. Changes to how the database sequences are fetched have also improved the speed of the report generator.

Computer Graphics↗

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

PROSPECT II: protein structure prediction program for genome-scale applications.

A new method for fold recognition is developed and added to the general protein structure prediction package PROSPECT (http://compbio.ornl.gov/PROSPECT/). The new method (PROSPECT II) has four key features. (i) We have developed an efficient way to utilize the evolutionary information for evaluating the threading potentials including singleton and pairwise energies. (ii) We have developed a two-stage threading strategy: (a) threading using dynamic programming without considering the pairwise energy and (b) fold recognition considering all the energy terms, including the pairwise energy calculated from the dynamic programming threading alignments. (iii) We have developed a combined z-score scheme for fold recognition, which takes into consideration the z-scores of each energy term. (iv) Based on the z-scores, we have developed a confidence index, which measures the reliability of a prediction and a possible structure-function relationship based on a statistical analysis of a large data set consisting of threadings of 600 query proteins against the entire FSSP templates. Tests on several benchmark sets indicate that the evolutionary information and other new features of PROSPECT II greatly improve the alignment accuracy. We also demonstrate that the performance of PROSPECT II on fold recognition is significantly better than any other method available at all levels of similarity. Improvement in the sensitivity of the fold recognition, especially at the superfamily and fold levels, makes PROSPECT II a reliable and fully automated protein structure and function prediction program for genome-scale applications.

Algorithms↗

Comparison of capsid sequences from human and animal astroviruses.

We have sequenced the genomic 3'-end, including the structural gene, of human astrovirus (HAstV) serotype 7 and morphologically related viruses infecting pig (PAstV), sheep (OAstV) and turkey (TAstV-1). These sequences were compared with corresponding astrovirus sequences available in the nucleic acid databases, including sequences of the seven other HAstV serotypes, two other avian astroviruses (TAstV-2 and avian nephritis virus) and astrovirus from cat (FAstV). A 35 nt stem-loop motif near the 3'-end of the genome, previously described as being highly conserved, was present in all of the astroviruses except TAstV-2. In the N-terminal half of the capsid precursor protein, there were several short conserved peptide motifs. Otherwise the capsid proteins of astroviruses infecting different hosts were highly divergent. Calculation of genetic distances revealed that the distance between FAstV and HAstV is comparable to the largest distances between different HAstV serotypes. Higher similarities between the HAstV, FAstV and PAstV capsid sequences suggest interspecies transmissions involving humans, cats and pigs relatively recently in the evolutionary history of astroviruses.

3' Untranslated Regions↗

Databases and software for the comparison of prokaryotic genomes.

The explosion in the number of complete genomes over the past decade has spawned a new and exciting discipline, that of comparative genomics. To exploit the full potential of this approach requires the development of novel algorithms, databases and software which are sophisticated enough to draw meaningful comparisons between complete genome sequences and are widely accessible to the scientific community at large. This article reviews progress towards the development of computational tools and databases for organizing and extracting biological meaning from the comparison of large collections of genomes.

Computational Biology↗

A human genomic library enriched in transcriptionally active sequences (aDNA library).

Core histone hyperacetylation, in particular of H4, is concentrated in the promoter-upstream regions of active genes and in certain cases is locuswide. Antibodies to hyperacetylated H4 were used to immunoprecipitate dinucleosomal chromatin derived from K562 human erythroleukemic cells by micrococcal nuclease digestion. The extracted DNA was made into a genomic library and was expected to contain sequences from genes active in K562 cells (an active, 'aDNA' library). Clones (180) were randomly selected from the library; 24 of 103 tested (23%) contained highly repeated sequences, as determined by their hybridization to total genomic DNA, and were not analyzed further. An additional 10 clones (6%) were shown to contain no insert DNA. The remaining 146 were sequenced and compared with the nucleic acid databases and in all six frames to the protein databases: Sixeen clones could be assigned to known genes, the majority of which (12) were tissue specific. All but 2 of these 16 corresponded to segments 5' of the coding sequences, as expected if H4 acetylation is concentrated at promoter regions. Thirty-three clones (23%) displayed high sequence identity to cDNAs in the expressed sequence tag database (dbEST). Northern blots and reverse transcription (RT)-PCR were used to determine the proportion of clones representing sequences expressed in K562 cells: Although only 1 of 34 tested clones showed a band in Northern hybridization, RT-PCR demonstrated that at least 12 of 40 tested clones (30%) were present in the mRNA population. Because a further 8 of these 40 clones were identified as gene fragments by database sequence comparisons, it follows that about half of this subset of 40 clones is derived from genes. The aDNA library is thus very gene rich and not skewed toward the most highly expressed sequences, as in mRNA libraries. The aDNA library is also rich in promoters and could be a valuable source of such sequences, particularly those that lack CpG islands or other features that allow their specific selection.

Blotting, Northern↗

Plant genome resources at the national center for biotechnology information.

The National Center for Biotechnology Information (NCBI) integrates data from more than 20 biological databases through a flexible search and retrieval system called Entrez. A core Entrez database, Entrez Nucleotide, includes GenBank and is tightly linked to the NCBI Taxonomy database, the Entrez Protein database, and the scientific literature in PubMed. A suite of more specialized databases for genomes, genes, gene families, gene expression, gene variation, and protein domains dovetails with the core databases to make Entrez a powerful system for genomic research. Linked to the full range of Entrez databases is the NCBI Map Viewer, which displays aligned genetic, physical, and sequence maps for eukaryotic genomes including those of many plants. A specialized plant query page allow maps from all plant genomes covered by the Map Viewer to be searched in tandem to produce a display of aligned maps from several species. PlantBLAST searches against the sequences shown in the Map Viewer allow BLAST alignments to be viewed within a genomic context. In addition, precomputed sequence similarities, such as those for proteins offered by BLAST Link, enable fluid navigation from unannotated to annotated sequences, quickening the pace of discovery. NCBI Web pages for plants, such as Plant Genome Central, complete the system by providing centralized access to NCBI's genomic resources as well as links to organism-specific Web pages beyond NCBI.

Biotechnology↗

Prospects for building the tree of life from large sequence databases.

We assess the phylogenetic potential of approximately 300,000 protein sequences sampled from Swiss-Prot and GenBank. Although only a small subset of these data was potentially phylogenetically informative, this subset retained a substantial fraction of the original taxonomic diversity. Sampling biases in the databases necessitate building phylogenetic data sets that have large numbers of missing entries. However, an analysis of two "supermatrices" suggests that even data sets with as much as 92% missing data can provide insights into broad sections of the tree of life.

Animals↗

Haemophilia A and haemophilia B: molecular insights.

This review focuses on selected areas that should interest both the scientist and the clinician alike: polymorphisms within the factor VIII and factor IX genes, their linkage, and their ethnic variation; a general assessment of mutations within both genes and a detailed inspection of the molecular pathology of certain mutations to illustrate the diverse cause-effect relations that exist; a summary of current knowledge on molecular aspects of inhibitor production; and an introduction to the new areas of factor VIII and factor IX catabolism. An appendix defining various terms encountered in the molecular genetics of the haemophilias is included, together with an appendix providing accession numbers and locus identification links for accessing gene and sequence information in the international nucleic acid databases.

Factor IX↗

Haemophilia A and haemophilia B: molecular insights.

This review focuses on selected areas that should interest both the scientist and the clinician alike: polymorphisms within the factor VIII and factor IX genes, their linkage, and their ethnic variation; a general assessment of mutations within both genes and a detailed inspection of the molecular pathology of certain mutations to illustrate the diverse cause-effect relations that exist; a summary of current knowledge on molecular aspects of inhibitor production; and an introduction to the new areas of factor VIII and factor IX catabolism. An appendix defining various terms encountered in the molecular genetics of the haemophilias is included, together with an appendix providing accession numbers and locus identification links for accessing gene and sequence information in the international nucleic acid databases.

Blood Coagulation↗

Identification of trace element-containing proteins in genomic databases.

Development of bioinformatics tools provided researchers with the ability to identify full sets of trace element-containing proteins in organisms for which complete genomic sequences are available. Recently, independent bioinformatics methods were used to identify all, or almost all, genes encoding selenocysteine-containing proteins in human, mouse, and Drosophila genomes, characterizing entire selenoproteomes in these organisms. It also should be possible to search for entire sets of other trace element-associated proteins, such as metal-containing proteins, although methods for their identification are still in development.

Animals↗

Gene homology resources on the World Wide Web.

As the amount of information available to biologists increases exponentially, data analysis becomes progressively more challenging. Sequence homology has been a traditional tool in the researchers' armamentarium; it is a very versatile instrument and can be employed to assist in numerous tasks, from establishing the function of a gene to determination of the evolutionary development of an organism. Consequently, numerous specialized tools have been established in the public domain (most commonly, the World Wide Web) to help investigators use sequence homology in their research. These homology databases differ both in techniques they use to compare sequences as well as in the size of the unit of analysis, which can be the whole gene, a domain, or a motif. In this paper, we aim to present a systematic review of the inner details of the most commonly used databases as well as to offer guidelines for their use.

Animals↗

The COG database: an updated version includes eukaryotes.

BACKGROUND: The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. RESULTS: We describe here a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes. CONCLUSION: The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies.

Animals↗

Large scale hierarchical clustering of protein sequences.

BACKGROUND: Searching a biological sequence database with a query sequence looking for homologues has become a routine operation in computational biology. In spite of the high degree of sophistication of currently available search routines it is still virtually impossible to identify quickly and clearly a group of sequences that a given query sequence belongs to. RESULTS: We report on our developments in grouping all known protein sequences hierarchically into superfamily and family clusters. Our graph-based algorithms take into account the topology of the sequence space induced by the data itself to construct a biologically meaningful partitioning. We have applied our clustering procedures to a non-redundant set of about 1,000,000 sequences resulting in a hierarchical clustering which is being made available for querying and browsing at http://systers.molgen.mpg.de/. CONCLUSIONS: Comparisons with other widely used clustering methods on various data sets show the abilities and strengths of our clustering methods in producing a biologically meaningful grouping of protein sequences.

Algorithms↗

Integrating alternative splicing detection into gene prediction.

BACKGROUND: Alternative splicing (AS) is now considered as a major actor in transcriptome/proteome diversity and it cannot be neglected in the annotation process of a new genome. Despite considerable progresses in term of accuracy in computational gene prediction, the ability to reliably predict AS variants when there is local experimental evidence of it remains an open challenge for gene finders. RESULTS: We have used a new integrative approach that allows to incorporate AS detection into ab initio gene prediction. This method relies on the analysis of genomically aligned transcript sequences (ESTs and/or cDNAs), and has been implemented in the dynamic programming algorithm of the graph-based gene finder EuGENE. Given a genomic sequence and a set of aligned transcripts, this new version identifies the set of transcripts carrying evidence of alternative splicing events, and provides, in addition to the classical optimal gene prediction, alternative optimal predictions (among those which are consistent with the AS events detected). This allows for multiple annotations of a single gene in a way such that each predicted variant is supported by a transcript evidence (but not necessarily with a full-length coverage). CONCLUSIONS: This automatic combination of experimental data analysis and ab initio gene finding offers an ideal integration of alternatively spliced gene prediction inside a single annotation pipeline.

Algorithms↗

DSD--an integrated, web-accessible database of Dehydrogenase Enzyme Stereospecificities.

BACKGROUND: Dehydrogenase enzymes belong to the oxidoreductase class and utilise the coenzymes NAD and NADP. Stereo-selectivity is focused on the C4 hydrogen atoms of the nicotinamide ring of NAD(P). Depending upon which hydrogen is transferred at the C4 location, the enzyme is designated as A or B stereospecific. DESCRIPTION: The Dehydrogenase Stereospecificity Database v1.0 (DSD) provides a compilation of enzyme stereochemical data, as sourced from the primary literature, in the form of a web-accessible database. There are two search engines, a menu driven search and a BLAST search. The entries are also linked to several external databases, including the NCBI and the Protein Data Bank, providing wide background information. The database is freely available online at: http://www.jenner.ac.uk/DSD/. CONCLUSION: DSD is a unique compilation available on-line for the first time which provides a key resource for the comparative analysis of reductase hydrogen transfer stereospecificity. As databases increasingly form the backbone of science, largely complete databases such as DSD, are a vital addition.

Computational Biology↗