Protein and Nucleic Acid Sequence Database Systems.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Nucleic acids are generally considered as efficient cation binders. Therefore, the likelihood that negatively charged ions might intrude their first hydration shell is rarely considered. Here, we show on the basis of (i) a survey of the Nucleic Acid Database, (ii) several structures extracted from the Cambridge Structural Database, and (iii) molecular dynamics simulations, that the nucleotide electropositive edges involving mainly amino, imino, and hydroxyl groups can cast specific anion binding sites. These binding sites constitute also good locations for the binding of the negatively charged groups of the Asp and Glu residues or the nucleic acid phosphate groups. Furthermore, it is observed in several instances that anions, like water molecules and cations, do mediate protein/nucleic acid interactions. Thus, anions as well as negatively charged groups are directly involved in specific recognition and folding phenomena involving polyanionic nucleic acids.
Several primer prediction and analysis programs have been developed for diverse applications. However, none of these existing programs can be directly used for the design of primers in protein interaction experiments, since proteins may have transmembrane domains (TMDs) and/or a signal peptide that must be excluded from experiments. Furthermore, it is frequently the case that a short restriction sequences must be added to each primer in order to clone PCR products into a given destination vectors for expression. DePIE, a web-based primer design tool, was developed to address these deficiencies. The program takes as input NCBI protein accession numbers and returns primer information including nucleotide sequences, thermodynamic melting temperature of the nucleotide sequences and the target positions. DePIE is implemented in JAVA, PERL and PHP and has proven to be very efficient in designing primers for our interaction experiments. DePIE services can be accessed at the web site: http://biocore.unl.edu/primer/primerPI.html.
A certain concern exists that the exponential growth of nucleic acid and protein sequence data will saturate the channels of data acquisition, distribution and utilization on the one hand and, on the other hand, that even the actual resources are still not fully and easily accessible to any bench scientist. Despite the stake of the scientific community at large in the fundamental data collected in this field, there has been in past years only a modest effort to discuss the common problems at an international level. Three international meetings were organized in 1987 on this subject: the annual meeting of CODATA Task Group on Coordination of Protein Sequence Data Banks (Nice, France, January 1987), the EMBL/NIH Workshop concerned primarily with nucleic acid databases (Heidelberg, FRG, February 1987) and the CODATA Workshop on Nucleic Acid and Protein Sequencing Data (Gaithersburg, USA, May 1987).
The program SFCHECK [Vaguine et al. (1999), Acta Cryst. D55, 191-205] is used to survey the quality of the structure-factor data and the agreement of those data with the atomic coordinates in 105 nucleic acid crystal structures for which structure-factor amplitudes have been deposited in the Nucleic Acid Database [NDB; Berman et al. (1992), Biophys. J. 63, 751-759]. Nucleic acid structures present a particular challenge for structure-quality evaluations. The majority of these structures, and DNA molecules in particular, have been solved by molecular replacement of the double-helical motif, whose high degree of symmetry can lead to problems in positioning the molecule in the unit cell. In this paper, the overall quality of each structure was evaluated using parameters such as the R factor, the correlation coefficient and various atomic error estimates. In addition, each structure is characterized by the average values of several local quality indicators, which include the atomic displacement, the density correlation, the B factor and the density index. The latter parameter measures the relative electron-density level at the atomic position. In order to assess the quality of the model in specific regions, the same local quality indicators are also surveyed for individual groups of atoms in each structure. Several of the global quality indicators are found to vary linearly with resolution and less than a dozen structures are found to exhibit values significantly different from the mean for these indicators, showing that the quality of the nucleic acid structures tends to be rather uniform. Analysis of the mutual dependence of the values of different local quality indicators, computed for individual residues and atom groups, reveals that these indicators essentially complement each other and are not redundant with the B factor. Using several of these indicators, it was found that the atomic coordinates of the nucleic acid bases tend to be better defined than those of the backbone. One of the local indicators, the density index, is particularly useful in spotting regions of the model that fit poorly in the electron density. Using this parameter, the quality of crystallographic water positions in the analyzed structures was surveyed and it was found that a sizable fraction of these positions have poorly defined electron density and may therefore not be reliable. The possibility that cases of poorly positioned water molecules are symptomatic of more widespread problems with the structure as a whole is also raised.
We describe two novel sequence similarity search algorithms, FASTS and FASTF, that use multiple short peptide sequences to identify homologous sequences in protein or DNA databases. FASTS searches with peptide sequences of unknown order, as obtained by mass spectrometry-based sequencing, evaluating all possible arrangements of the peptides. FASTF searches with mixed peptide sequences, as generated by Edman sequencing of unseparated mixtures of peptides. FASTF deconvolutes the mixture, using a greedy heuristic that allows rapid identification of high scoring alignments while reducing the total number of explored alternatives. Both algorithms use the heuristic FASTA comparison strategy to accelerate the search but use alignment probability, rather than similarity score, as the criterion for alignment optimality. Statistical estimates are calculated using an empirical correction to a theoretical probability. These calculated estimates were accurate within a factor of 10 for FASTS and 1000 for FASTF on our test dataset. FASTS requires only 15-20 total residues in three or four peptides to robustly identify homologues sharing 50% or greater protein sequence identity. FASTF requires about 25% more sequence data than FASTS for equivalent sensitivity, but additional sequence data are usually available from mixed Edman experiments. Thus, both algorithms can identify homologues that diverged 100 to 500 million years ago, allowing proteomic identification from organisms whose genomes have not been sequenced.
Metal ions are essential for the folding of RNA into stable tertiary structures and for the catalytic activity of some RNA enzymes. To aid in the study of the roles of metal ions in RNA structural biology, we have created MeRNA (Metals in RNA), a comprehensive compilation of all metal binding sites identified in RNA 3D structures available from the PDB and Nucleic Acid Database. Currently, our database contains information relating to binding of 9764 metal ions corresponding to 23 distinct elements, in 256 RNA structures. The metal ion locations were confirmed and ligands characterized using original literature references. MeRNA includes eight manually identified metal-ion binding motifs, which are described in the literature. MeRNA is searchable by PDB identifier, metal ion, method of structure determination, resolution and R-values for X-ray structure and distance from metal to any RNA atom or to water. New structures with their respective binding motifs will be added to the database as they become available. The MeRNA database will further our understanding of the roles of metal ions in RNA folding and catalysis and have applications in structural and functional analysis, RNA design and engineering. The MeRNA database is accessible at http://merna.lbl.gov.
It has been noticed that converged conformations of B-DNA oligomers obtained in MD calculations often have very small atom position rmsd values from the canonical B-DNA and all helical parameters close to the standard values, but their minor grooves tend to be somewhat narrower. This apparent bias disappears, however, when C5' rather than phosphorus atoms are used for measuring the groove width. At the origin of this effect is the specific orientation of phosphate groups in the canonical B-DNA model which maximizes their separation across the minor groove. When measured by C5' traces, minor groove profiles of experimental structures available in the Nucleic Acids Database show much less tendency to narrow below the canonical width. Correlation analysis reveals a high degree of correspondence in shapes of minor grooves of calculated and experimental single-crystal structures of B-DNA oligomers.
Protein-DNA interactions are crucial for many biological processes. Attempts to model these interactions have generally taken the form of amino acid-base recognition codes or purely sequence-based profile methods, which depend on the availability of extensive sequence and structural information for specific structural families, neglect side-chain conformational variability, and lack generality beyond the structural family used to train the model. Here, we take advantage of recent advances in rotamer-based protein design and the large number of structurally characterized protein-DNA complexes to develop and parameterize a simple physical model for protein-DNA interactions. The model shows considerable promise for redesigning amino acids at protein-DNA interfaces, as design calculations recover the amino acid residue identities and conformations at these interfaces with accuracies comparable to sequence recovery in globular proteins. The model shows promise also for predicting DNA-binding specificity for fixed protein sequences: native DNA sequences are selected correctly from pools of competing DNA substrates; however, incorporation of backbone movement will likely be required to improve performance in homology modeling applications. Interestingly, optimization of zinc finger protein amino acid sequences for high-affinity binding to specific DNA sequences results in proteins with little or no predicted specificity, suggesting that naturally occurring DNA-binding proteins are optimized for specificity rather than affinity. When combined with algorithms that optimize specificity directly, the simple computational model developed here should be useful for the engineering of proteins with novel DNA-binding specificities.
SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and a high level of integration with other databases. Together with its automatically annotated supplement TrEMBL, it provides a comprehensive and high-quality view of the current state of knowledge about proteins. Ongoing developments include the further improvement of functional and automatic annotation in the databases including evidence attribution with particular emphasis on the human, archaeal and bacterial proteomes and the provision of additional resources such as the International Protein Index (IPI) and XML format of SWISS-PROT and TrEMBL to the user community.
With the advent of automated and high-throughput techniques, the number of patent applications containing biological sequences has been increasing rapidly. However, they have attracted relatively little attention compared to other sequence resources. We have built a database server called Patome, which contains biological sequence data disclosed in patents and published applications, as well as their analysis information. The analysis is divided into two steps. The first is an annotation step in which the disclosed sequences were annotated with RefSeq database. The second is an association step where the sequences were linked to Entrez Gene, OMIM and GO databases, and their results were saved as a gene-patent table. From the analysis, we found that 55% of human genes were associated with patenting. The gene-patent table can be used to identify whether a particular gene or disease is related to patenting. Patome is available at http://www.patome.org/; the information is updated bimonthly.
NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.
This chapter outlines the basic requirements for finding and exploring sequences of interest in public databases, such as GenBank. As such, it is not aimed at experienced sequencers, for whom this will be "second nature," but at the many clinical bacteriologists who rarely have need of DNA sequences in their usual work, and who would like to develop their interest in what can appear to be a daunting area. The topics discussed include finding and retrieving sequences from GenBank, identifying homologous sequences using BLAST searches, resources for accessing microbial genomes, and the Protein Data Bank. Finally, recommendations are made for useful software (freeware) and online sequence manipulation resources.
EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).
In order to scan nucleic acid databases for potentially relevant but as yet unknown signals, we have developed an improved statistical model for pattern analysis of nucleic acid sequences by modifying previous methods based on Markov chains. We demonstrate the importance of selecting the appropriate parameters in order for the method to function at all. The model allows the simultaneous analysis of several short sequences with unequal base frequencies and Markov order k not equal to 0 as is usually the case in databases. As a test of these modifications, we show that in E. coli sequences there is a bias against palindromic hexamers which correspond to known restriction enzyme recognition sites.
The Histone Database is a curated and searchable collection of full-length sequences and structures of histones and nonhistone proteins containing histone-like folds, compiled from major public databases. Several new histone fold-containing proteins have been identified, including the huntingtin-interacting protein HYPM. Additionally, based on the recent crystal structure of the Son of Sevenless protein, an interpretation of the sequence analysis of the histone fold domain is presented. The database contains an updated collection of multiple sequence alignments for the four core histones (H2A, H2B, H3, and H4) and the linker histones (H1/H5) from a total of 975 organisms. The database also contains information on the human histone gene complement and provides links to three-dimensional structures of histone and histone fold-containing proteins. The Histone Database is a comprehensive bioinformatics resource for the study of structure and function of histones and histone fold-containing proteins. The database is available at http://research.nhgri.nih.gov/histones/.
We have compiled the nucleotide sequences and their amino acid translations from a total of 89 Killer Immunoglobulin-like Receptor (KIR) alleles, derived from 17 different KIR genes. The alignments use the KIR3DL2*001 allele as a reference sequence. Each of the KIR sequences included in these alignments has been checked and where discrepancies have arisen between reported sequences, the original authors have been contacted where possible, and necessary amendments to published sequences have been incorporated into this alignment. Future sequencing may identify errors in this list and we would welcome any evidence that helps to maintain the accuracy of this compilation.
Compensated frameshift mutation is a modification of the reading frame of a gene that takes place by way of various molecular events. It appears to be a widespread event that is only observed when homologous amino acid and nucleodotide sequences are compared. To identify these mutation events, the sequence analysis rationale was based on the search for short regions that would have much lower degrees of conservation in protein, but not in DNA, in well-conserved beta-glucosidase families. We have restricted our study to a seed set of sequences of O-glycoside hydrolase families 1 and 3. We found compensated frameshift mutation in the family of 1 beta-glucosidases for the Erwinia herbicola, Cellulomonas fimi, and (non-cyanogenic) Trifolium repens gene sequences, and in the family of 3 beta-glucosidases for the Clostridium thermocellum and Clostridium stercorarium gene sequences. By computational treatment, the observed mutation events in the gene frameshifting sub-sequence have been neutralised. Each nucleotide insertion must be eliminated and each nucleotide deletion must be substituted by the symbol N (any nucleotide). When the frameshifting fragments of the amino acid sequences were substituted by the computationally neutralised subsequences, the beta-glucosidase alignments were improved. We also discuss the structural implications of the compensated frameshift mutations events.