PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Genomic BLAST: custom-defined virtual databases for complete and unfinished genomes.

BLAST (Basic Local Alignment Search Tool) searches against DNA and protein sequence databases have become an indispensable tool for biomedical research. The proliferation of the genome sequencing projects is steadily increasing the fraction of genome-derived sequences in the public databases and their importance as a public resource. We report here the availability of Genomic BLAST, a novel graphical tool for simplifying BLAST searches against complete and unfinished genome sequences. This tool allows the user to compare the query sequence against a virtual database of DNA and/or protein sequences from a selected group of organisms with finished or unfinished genomes. The organisms for such a database can be selected using either a graphic taxonomy-based tree or an alphabetical list of organism-specific sequences. The first option is designed to help explore the evolutionary relationships among organisms within a certain taxonomy group when performing BLAST searches. The use of an alphabetical list allows the user to perform a more elaborate set of selections, assembling any given number of organism-specific databases from unfinished or complete genomes. This tool, available at the NCBI web site http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/genom_table_cgi, currently provides access to over 170 bacterial and archaeal genomes and over 40 eukaryotic genomes.

Amino Acid Sequence↗

Study of transcripts from AC010088, a 199,485 bp fragment of the human Y chromosome located in the azoospermia factor region c.

Deletions on the long arm of the human Y chromosome are associated with male infertility. In this work, we studied transcripts of a 199,485 bp long fragment of the Yq11 region (GenBank accession number, AC010088) located in the AZFc (azoospermia factor region c), and characterized their gene structures. After masking repetitive elements, we searched human mRNA Refseqs (reference sequences), a dbEST (database of expressed sequence tags) and a non-redundant nucleic acid database for the mRNAs and ESTs corresponding to the AC010088 using the BLAST programs at the NCBI (National Center for Biotechnology Information) site. Our findings are summarized as follows: i) BPY2 (testis basic protein on Y, 2), DAZ1 (deleted in azoospermia 1), TTY4 (testis transcript Y 4) mRNAs and 23 ESTs were found; ii) Eighteen of 23 ESTs were transcripts of the DAZ gene(s), one EST was a transcript of TTY4 gene, and the remaining 4 probably corresponded to 4 different pseudogenes; iii) DAZ gene(s) were expressed not only in testis, but also in lung carcinoma cells, stomach and Ewing's sarcoma cells; iv) beta-satellite clusters were present around and within the BPY2 and TTY4 gene region; v) In this study, TTY4, BPY2 and DAZ1 genes were mapped precisely to the AC010088 region.

Chromosome Mapping↗

Generation of a database containing discordant intron positions in eukaryotic genes (MIDB).

MOTIVATION: Intron sliding is the relocation of intron-exon boundaries over short distances and is often also referred to as intron slippage or intron migration or intron drift. We have generated a database containing discordant intron positions in homologous genes (MIDB--Mismatched Intron DataBase). Discordant intron positions are those that are either closely located in homologous genes (within a window of 10 nucleotides) or an intron position that is present in one gene but not in any of its homologs. The MIDB database aims at systematically collecting information about mismatched introns in the genes from GenBank and organizing it into a form useful for understanding the genomics and dynamics of introns thereby helping understand the evolution of genes. RESULTS: Intron displacement or sliding is critically important for explaining the present distribution of introns among orthologous and paralogous genes. MIDB allows examining of intron movements and allows mapping of intron positions from homologous proteins onto a single sequence. The database is of potential use for molecular biologists in general and for researchers who are interested in gene evolution and eukaryotic gene structure. Partial analysis of this database allowed us to identify a few putative cases of intron sliding. AVAILABILITY: http://intron.bic.nus.edu.sg/midb/midb.html

Amino Acid Sequence↗

PROSIT: pseudo-rotational online service and interactive tool, applied to a conformational survey of nucleosides and nucleotides.

A Pseudo-Rotational Online Service and Interactive Tool (PROSIT) designed to perform complete pseudorotational analysis of nucleosides and nucleotides is described. This service is freely available at http://cactus.nci.nih.gov/prosit/. Files containing nucleosides/nucleotides or DNA/RNA segments, isolated or bound to other molecules (e.g., a protein) can be uploaded to be processed by PROSIT. The service outputs the pseudorotational phase angle P, puckering amplitude numax, and other related information for each nucleoside/nucleotide detected. The service was implemented using the chemoinformatics toolkit CACTVS. PROSIT was used for a survey of nucleosides contained in the Cambridge Structural Database and nucleotides in high-resolution crystal structures from the Nucleic Acid Database. Special cases discussed include nucleosides having constrained sugar moieties with extreme puckering amplitudes, and several specific DNA/RNA helices and protein-bound DNA oligonucleotides (Dickerson-Drew dodecamer, RNA/DNA hybrid viral polypurine tract, Z-DNA enantiomers, B-DNA containing (L)-alpha-threofuranosyl nucleotides, TATA-box binding protein/TATA-box complex, and DNA (cytosine C5)-methyltransferase complexed with an oligodeoxyribonucleotide containing transition state analogue 5,6-dihydro-5-azacytosine). When the puckering amplitude decreases to a small value, the sugar becomes increasingly planar, thus reducing the significance of the phase angle P. We introduce the term "central conformation" to describe this part of the pseudorotational hyperspace in contrast to the conventional north and south conformations.

Nucleic Acid Conformation↗

Improving the accuracy of NMR structures of DNA by means of a database potential of mean force describing base-base positional interactions.

NMR structure determination of nucleic acids presents an intrinsically difficult problem since the density of short interproton distance contacts is relatively low and limited to adjacent base pairs. Although residual dipolar couplings provide orientational information that is clearly helpful, they do not provide translational information of either a short-range (with the exception of proton-proton dipolar couplings) or long-range nature. As a consequence, the description of the nonbonded contacts has a major impact on the structures of nucleic acids generated from NMR data. In this paper, we describe the derivation of a potential of mean force derived from all high-resolution (2 A or better) DNA crystal structures available in the Nucleic Acid Database (NDB) as of May 2000 that provides a statistical description, in simple geometric terms, of the relative positions of pairs of neighboring bases (both intra- and interstrand) in Cartesian space. The purpose of this pseudopotential, which we term a DELPHIC base-base positioning potential, is to bias sampling during simulated annealing refinement to physically reasonable regions of conformational space within the range of possibilities that are consistent with the experimental NMR restraints. We illustrate the application of the DELPHIC base-base positioning potential to the structure refinement of a DNA dodecamer, d(CGCGAATTCGCG)(2), for which NOE and dipolar coupling data have been measured in solution and for which crystal structures have been determined. We demonstrate by cross-validation against independent NMR observables (that is, both residual dipolar couplings and NOE-derived intereproton distance restraints) that the DELPHIC base-base positioning potential results in a significant increase in accuracy and obviates artifactual distortions in the structures arising from the limitations of conventional descriptions of the nonbonded contacts in terms of either Lennard-Jones van der Waals and electrostatic potentials or a simple van der Waals repulsion potential. We also demonstrate, using experimental NMR data for a complex of the male sex determining factor SRY with a duplex DNA 14mer, which includes a region of highly unusual and distorted DNA, that the DELPHIC base-base positioning potential does not in any way hinder unusual interactions and conformations from being satisfactorily sampled and reproduced. We expect that the methodology described in this paper for DNA can be equally applied to RNA, as well as side chain-side chain interactions in proteins and protein-protein complexes, and side chain-nucleic acid interactions in protein-nucleic acid complexes. Further, this approach should be useful not only for NMR structure determination but also for refinement of low-resolution (3-3.5 A) X-ray data.

DNA↗

EMBL-Align: a new public nucleotide and amino acid multiple sequence alignment database.

UNLABELLED: The submission of multiple sequence alignment data to EMBL has grown 30-fold in the past 10 years, creating a problem of archiving them. The EBI has developed a new public database of multiple sequence alignments called EMBL-Align. It has a dedicated web-based submission tool, Webin-Align. Together they represent a comprehensive data management solution for alignment data. Webin-Align accepts all the common alignment formats and can display data in CLUSTALW format as well as a new standard EMBL-Align flat file format. The alignments are stored in the EMBL-Align database and can be queried from the EBI SRS (Sequence Retrieval System) server. AVAILABILITY: Webin-Align: http://www.ebi.ac.uk/embl/Submission/align_top.html, EMBL-Align: ftp://ftp.ebi.ac.uk/pub/databases/embl/align, http://srs.ebi.ac.uk/

Amino Acid Sequence↗

Pandit: a database of protein and associated nucleotide domains with inferred trees.

MOTIVATION: A large, high-quality database of homologous sequence alignments with good estimates of their corresponding phylogenetic trees will be a valuable resource to those studying phylogenetics. It will allow researchers to compare current and new models of sequence evolution across a large variety of sequences. The large quantity of data may provide inspiration for new models and methodology to study sequence evolution and may allow general statements about the relative effect of different molecular processes on evolution. RESULTS: The Pandit 7.6 database contains 4341 families of sequences derived from the seed alignments of the Pfam database of amino acid alignments of families of homologous protein domains (Bateman et al., 2002). Each family in Pandit includes an alignment of amino acid sequences that matches the corresponding Pfam family seed alignment, an alignment of DNA sequences that contain the coding sequence of the Pfam alignment when they can be recovered (overall, 82.9% of sequences taken from Pfam) and the alignment of amino acid sequences restricted to only those sequences for which a DNA sequence could be recovered. Each of the alignments has an estimate of the phylogenetic tree associated with it. The tree topologies were obtained using the neighbor joining method based on maximum likelihood estimates of the evolutionary distances, with branch lengths then calculated using a standard maximum likelihood approach.

Algorithms↗

Exploring the sequence-structure protein landscape in the glycosyltransferase family.

To understand the molecular basis of glycosyltransferases' (GTFs) catalytic mechanism, extensive structural information is required. Here, fold recognition methods were employed to assign 3D protein shapes (folds) to the currently known GTF sequences, available in public databases such as GenBank and Swissprot. First, GTF sequences were retrieved and classified into clusters, based on sequence similarity only. Intracluster sequence similarity was chosen sufficiently high to ensure that the same fold is found within a given cluster. Then, a representative sequence from each cluster was selected to compose a subset of GTF sequences. The members of this reduced set were processed by three different fold recognition methods: 3D-PSSM, FUGUE, and GeneFold. Finally, the results from different fold recognition methods were analyzed and compared to sequence-similarity search methods (i.e., BLAST and PSI-BLAST). It was established that the folds of about 70% of all currently known GTF sequences can be confidently assigned by fold recognition methods, a value which is higher than the fold identification rate based on sequence comparison alone (48% for BLAST and 64% for PSI-BLAST). The identified folds were submitted to 3D clustering, and we found that most of the GTF sequences adopt the typical GTF A or GTF B folds. Our results indicate a lack of evidence that new GTF folds (i.e., folds other than GTF A and B) exist. Based on cases where fold identification was not possible, we suggest several sequences as the most promising targets for a structural genomics initiative focused on the GTF protein family.

Algorithms↗

Usability of BioCon2 for nucleic acid structures in database.

With the specific three-dimensional structure, DNA/RNA molecules express the specific structural and catalytic functions. The knowledge integration based on three dimensional structures can be used to explain biochemical observations, to predict biological functions and to design drugs specific to a given complex system. In the large RNAs, the helical stems close to each other with specific arrangement. The polymorphic nature of multiple stranded helical structure of nucleic acid relates a promising technique of achieving base sequence specific recognition and hence artificial gene regulation. The database cataloging the interaction motifs of nucleic acid moieties has been developed. Skew matrix, a kind of kink parameter, affords the structural description between the adjacent moieties. The program BioCon2 for estimating these skew matrix and cataloguing for multiple stranded helices has been developed. New version allows for the good properties, i.e., impregnable and flexible presentation of multiple helices. The tentative database including these parameters with the species and physical properties of the surrounding nucleic acid components and amino acids has been constructed.

Databases, Nucleic Acid↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

HUGE: a database for human large proteins identified in the Kazusa cDNA sequencing project.

We have been developing a HUGE database to summarize results from the sequence analysis of human novel large (>4 kb) cDNAs identified in the Kazusa cDNA sequencing project, systematically designated KIAA plus a four-digit number. HUGE currently contains nearly 2000 gene/protein characteristic tables harboring the results of the computer-assisted analysis of the cDNA and the predicted protein sequences together with those of expression profiling and chromosomal mapping. In the updated version of HUGE, we made it possible to compare each KIAA cDNA sequence with the corresponding entry in the human draft genome sequence that was published recently. Approximately 90% of KIAA cDNAs in HUGE can be localized along the human genome for at least half or more of the cDNA's length. Any nucleotide differences between the cDNA and the corresponding genomic sequences are also presented in detail. This new version of HUGE greatly helps us evaluate the completeness of cDNA clones and the accuracy of cDNA/genomic sequences. More interestingly, in some cases, the ability to compare cDNA with genomic sequences allows us to identify candidate sites of RNA editing. HUGE is available on the World Wide Web at http://www.kazusa.or.jp/huge.

Amino Acid Sequence↗

Prediction of protein function from protein sequence and structure.

The sequence of a genome contains the plans of the possible life of an organism, but implementation of genetic information depends on the functions of the proteins and nucleic acids that it encodes. Many individual proteins of known sequence and structure present challenges to the understanding of their function. In particular, a number of genes responsible for diseases have been identified but their specific functions are unknown. Whole-genome sequencing projects are a major source of proteins of unknown function. Annotation of a genome involves assignment of functions to gene products, in most cases on the basis of amino-acid sequence alone. 3D structure can aid the assignment of function, motivating the challenge of structural genomics projects to make structural information available for novel uncharacterized proteins. Structure-based identification of homologues often succeeds where sequence-alone-based methods fail, because in many cases evolution retains the folding pattern long after sequence similarity becomes undetectable. Nevertheless, prediction of protein function from sequence and structure is a difficult problem, because homologous proteins often have different functions. Many methods of function prediction rely on identifying similarity in sequence and/or structure between a protein of unknown function and one or more well-understood proteins. Alternative methods include inferring conservation patterns in members of a functionally uncharacterized family for which many sequences and structures are known. However, these inferences are tenuous. Such methods provide reasonable guesses at function, but are far from foolproof. It is therefore fortunate that the development of whole-organism approaches and comparative genomics permits other approaches to function prediction when the data are available. These include the use of protein-protein interaction patterns, and correlations between occurrences of related proteins in different organisms, as indicators of functional properties. Even if it is possible to ascribe a particular function to a gene product, the protein may have multiple functions. A fundamental problem is that function is in many cases an ill-defined concept. In this article we review the state of the art in function prediction and describe some of the underlying difficulties and successes.

Amino Acid Sequence↗

Interrogating the human genome using uninterpreted mass spectrometry data.

The public availability of a draft assembly of the human genome has enabled us to demonstrate, for the first time, the feasibility of searching a complete, unmasked eukaryotic genome using uninterpreted mass spectrometry data. A complex LC-MS/MS data set, containing peptides from at least 22 human proteins, was searched against a comprehensive, nonidentical protein database, an expressed sequence tag (EST) database, and the International Human Genome Project draft assembly of the human genome. The results from the three searches are compared in detail, and the merits of the different databases for this application are discussed. In the case of the EST database, the UniGene index provided a method of simplifying and summarising the search results. In the case of the genomic DNA, the presence of introns prevented matching of roughly one quarter of the spectra, but the technique can provide primary experimental verification of predicted coding sequences, and has the potential to identify novel coding sequences.

Algorithms↗

Genome sequencing and annotation: an overview.

Many microbial genome sequences have been determined, and more new genome projects are ongoing. Shotgun sequencing of randomly cloned short pieces of genomic DNA can provide a simple way of determining whole genome sequences. This process requires sequencing of many fragments, compilation of the separate sequences into one contiguous sequence, and careful editing of the assembled sequence. The genes present on the microbial genome are then predicted using clues derived from typical gene features, such as codon usage, ribosomal binding sequences, and bacterial initiation codons. Function of genes is predicted by homology searches performed against either public or well-established protein databases. This chapter discusses each of these stages in a genome-sequencing project.

Amino Acid Sequence↗

Paper2sequences: retrieval of sequences listed in a publication.

Our web-based tool simplifies the often laborious procedure of retrieving a set of biosequences in a publication or webpage. As a front-end to the Bioperl toolkit, it accepts as an input a list of identifiers. They are specified in an ASCII table (copy-pasted from the publication's PDF or HTML page) and give rise to queries in multiple databases for the protein/nucleic acid data specified. Currently, GenBank, PIR (Protein Information Resource) and Swiss-Prot are supported. For any sequence accession code listed, the database can be specified and, if retrieval fails, automatic lookup for the same code in other databases can be requested. Sequence length information (if specified) and heuristic rules are used to drive the lookup if multiple protein coding sequences (CDS) are part of a single accession. Warnings are issued in cases of ambiguities and inconsistencies. An advanced option enables the user to format the output in whatever format they wish.

Amino Acid Sequence↗

Standard atomic volumes in double-stranded DNA and packing in protein--DNA interfaces.

Standard volumes for atoms in double-stranded B-DNA are derived using high resolution crystal structures from the Nucleic Acid Database (NDB) and compared with corresponding values derived from crystal structures of small organic compounds in the Cambridge Structural Database (CSD). Two different methods are used to compute these volumes: the classical Voronoi method, which does not depend on the size of atoms, and the related Radical Planes method which does. Results show that atomic groups buried in the interior of double-stranded DNA are, on average, more tightly packed than in related small molecules in the CSD. The packing efficiency of DNA atoms at the interfaces of 25 high resolution protein-DNA complexes is determined by computing the ratios between the volumes of interfacial DNA atoms and the corresponding standard volumes. These ratios are found to be close to unity, indicating that the DNA atoms at protein-DNA interfaces are as closely packed as in crystals of B-DNA. Analogous volume ratios, computed for buried protein atoms, are also near unity, confirming our earlier conclusions that the packing efficiency of these atoms is similar to that in the protein interior. In addition, we examine the number, volume and solvent occupation of cavities located at the protein-DNA interfaces and compared them with those in the protein interior. Cavities are found to be ubiquitous in the interfaces as well as inside the protein moieties. The frequency of solvent occupation of cavities is however higher in the interfaces, indicating that those are more hydrated than protein interiors. Lastly, we compare our results with those obtained using two different measures of shape complementarity of the analysed interfaces, and find that the correlation between our volume ratios and these measures, as well as between the measures themselves, is weak. Our results indicate that a tightly packed environment made up of DNA, protein and solvent atoms plays a significant role in protein-DNA recognition.

Animals↗