PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

GATA: a graphic alignment tool for comparative sequence analysis.

BACKGROUND: Several problems exist with current methods used to align DNA sequences for comparative sequence analysis. Most dynamic programming algorithms assume that conserved sequence elements are collinear. This assumption appears valid when comparing orthologous protein coding sequences. Functional constraints on proteins provide strong selective pressure against sequence inversions, and minimize sequence duplications and feature shuffling. For non-coding sequences this collinearity assumption is often invalid. For example, enhancers contain clusters of transcription factor binding sites that change in number, orientation, and spacing during evolution yet the enhancer retains its activity. Dot plot analysis is often used to estimate non-coding sequence relatedness. Yet dot plots do not actually align sequences and thus cannot account well for base insertions or deletions. Moreover, they lack an adequate statistical framework for comparing sequence relatedness and are limited to pairwise comparisons. Lastly, dot plots and dynamic programming text outputs fail to provide an intuitive means for visualizing DNA alignments. RESULTS: To address some of these issues, we created a stand alone, platform independent, graphic alignment tool for comparative sequence analysis (GATA http://gata.sourceforge.net/). GATA uses the NCBI-BLASTN program and extensive post-processing to identify all small sub-alignments above a low cut-off score. These are graphed as two shaded boxes, one for each sequence, connected by a line using the coordinate system of their parent sequence. Shading and colour are used to indicate score and orientation. A variety of options exist for querying, modifying and retrieving conserved sequence elements. Extensive gene annotation can be added to both sequences using a standardized General Feature Format (GFF) file. CONCLUSIONS: GATA uses the NCBI-BLASTN program in conjunction with post-processing to exhaustively align two DNA sequences. It provides researchers with a fine-grained alignment and visualization tool aptly suited for non-coding, 0-200 kb, pairwise, sequence analysis. It functions independent of sequence feature ordering or orientation, and readily visualizes both large and small sequence inversions, duplications, and segment shuffling. Since the alignment is visual and does not contain gaps, gene annotation can be added to both sequences to create a thoroughly descriptive picture of DNA conservation that is well suited for comparative sequence analysis.

Algorithms↗

A systematic method for studying the spatial distribution of water molecules around nucleic acid bases.

A new method to analyze the distribution of water molecules around the bases in DNA is presented. This method relies on the notion of a "hydrated building block," which represents the joint observed hydration around all bases of a particular type, in structures of a particular conformation type. The hydrated building blocks were constructed using atomic coordinates from 40 structures contained in the Nucleic Acid Database. Pseudoelectron densities were calculated for water molecules in each hydrated building block using standard crystallographic procedures. The electron densities were fitted to obtain "average building blocks," which represent bases with waters only at average or probable positions. Both types of building blocks were used to construct models of hydrated DNA oligomers. The essential features of the solvent structure around d(CGCGAATTCGCG)2 in the B form and d(CGCGCG)2 in the Z form were reproduced.

Adenine↗

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

PARALIGN: rapid and sensitive sequence similarity searches powered by parallel computing technology.

PARALIGN is a rapid and sensitive similarity search tool for the identification of distantly related sequences in both nucleotide and amino acid sequence databases. Two algorithms are implemented, accelerated Smith-Waterman and ParAlign. The ParAlign algorithm is similar to Smith-Waterman in sensitivity, while as quick as BLAST for protein searches. A form of parallel computing technology known as multimedia technology that is available in modern processors, but rarely used by other bioinformatics software, has been exploited to achieve the high speed. The software is also designed to run efficiently on computer clusters using the message-passing interface standard. A public search service powered by a large computer cluster has been set-up and is freely available at www.paralign.org, where the major public databases can be searched. The software can also be downloaded free of charge for academic use.

Algorithms↗

Protein-DNA interactions: A structural analysis.

A detailed analysis of the DNA-binding sites of 26 proteins is presented using data from the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB). Chemical and physical properties of the protein-DNA interface, such as polarity, size, shape, and packing, were analysed. The DNA-binding sites shared common features, comprising many discontinuous sequence segments forming hydrophilic surfaces capable of direct and water-mediated hydrogen bonds. These interface sites were compared to those of protein-protein binding sites, revealing them to be more polar, with many more intermolecular hydrogen bonds and buried water molecules than the protein-protein interface sites. By looking at the number and positioning of protein residue-DNA base interactions in a series of interaction footprints, three modes of DNA binding were identified (single-headed, double-headed and enveloping). Six of the eight enzymes in the data set bound in the enveloping mode, with the protein presenting a large interface area effectively wrapped around the DNA.A comparison of structural parameters of the DNA revealed that some values for the bound DNA (including twist, slide and roll) were intermediate of those observed for the unbound B-DNA and A-DNA. The distortion of bound DNA was evaluated by calculating a root-mean-square deviation on fitting to a canonical B-DNA structure. Major distortions were commonly caused by specific kinks in the DNA sequence, some resulting in the overall bending of the helix. The helix bending affected the dimensions of the grooves in the DNA, allowing the binding of protein elements that would otherwise be unable to make contact. From this structural analysis a preliminary set of rules that govern the bending of the DNA in protein-DNA complexes, are proposed.

Binding Sites↗

TranScout: prediction of gene expression regulatory proteins from their sequences.

MOTIVATION: The advent of genomics yields thousands of reading frames in search of function. Identification of conserved functional motifs in protein sequences can be helpful for function prediction. RESULTS: A database and a classification of reported DNA-binding protein motifs has been designed. A program ('TranScout') has been developed for the detection and evaluation of conserved motifs in prokaryotic and eukaryotic sequences of proteins with a gene regulatory function. The efficiency of the program is shown in a benchmark against a database obtained from SWISS-PROT without the protein sequences used to train the program. All motifs were detected with a mean average sensitivity of 0.98 and a mean average specificity of 0.92. AVAILABILITY: The program is freely available for use on the internet at http://luz.uab.es/transcout/. The user can find additional information at this site.

Algorithms↗

Preparing a human membrane and secreted protein-enriched cDNA library using PCR primers derived from a genomic database.

We describe here a strategy for preparing a human membrane and secreted protein (MSP)-enriched cDNA library based on human MSP- and non-MSP-encoding cDNA sequences in the databases. The signal peptide parts of the MSP-encoding cDNA sequences, which currently comprise about half of the estimated total number in humans, were analyzed for common patterns. These patterns form a 'minimal' set of polymerase chain reaction primer candidates of length varying from 9 to 21 nt. The products stemming from each primer candidate were determined and the results allowed us to obtain an 'optimal' mixed-length primer set. Ninety-six percent of the primers in this set were predicted to yield </=10% undesired products, and the desired MSP-cDNA products could be easily separated by gel electrophoresis. The present analysis establishes a methodology for preparing a cDNA library that enables the analysis of individual MSPs. This methodology may also help identify new MSPs. As many cell regulatory processes are mediated by secreted proteins and their membrane-bound receptors, the preparation of a MSP-enriched cDNA library should benefit research on MSPs.

Amino Acid Sequence↗

SWISS-PROT: connecting biomolecular knowledge via a protein database.

With the explosive growth of biological data, the development of new means of data storage was needed. More and more often biological information is no longer published in the conventional way via a publication in a scientific journal, but only deposited into a database. In the last two decades these databases have become essential tools for researchers in biological sciences. Biological databases can be classified according to the type of information they contain. There are basically three types of sequence-related databases (nucleic acid sequences, protein sequences and protein tertiary structures) as well as various specialized data collections. It is important to provide the users of biomolecular databases with a degree of integration between these databases as by nature all of these databases are connected in a scientific sense and each one of them is an important piece to biological complexity. In this review we will highlight our effort in connecting biological information as demonstrated in the SWISS-PROT protein database.

Amino Acid Sequence↗

[Strategy for the protein identification of human proteome expression profile: selection of searching database].

Widely used method of protein identification for high-throughout proteome expression profile studies was database-dependent, so the selection of databases for the protein identification was very important. Despite the deficiency of available human protein databases, the complementarity of human proteins could be got mainly from human genome but not from the protein databases of other organisms. According to the comparison of the current protein databases from different aspects, IPI was recommended for the basic identification for the studies of human proteome expression profile, and other human protein or nucleic acid databases were needed for the complementary identification and novel protein mining.

Animals↗

Bioinformatic tools for DNA/protein sequence analysis, functional assignment of genes and protein classification.

The development of efficient DNA sequencing methods has led to the achievement of the DNA sequence of entire genomes from (to date) 55 prokaryotes, 5 eukaryotic organisms and 10 eukaryotic chromosomes. Thus, an enormous amount of DNA sequence data is available and even more will be forthcoming in the near future. Analysis of this overwhelming amount of data requires bioinformatic tools in order to identify genes that encode functional proteins or RNA. This is an important task, considering that even in the well-studied Escherichia coli more than 30% of the identified open reading frames are hypothetical genes. Future challenges of genome sequence analysis will include the understanding of gene regulation and metabolic pathway reconstruction including DNA chip technology, which holds tremendous potential for biomedicine and the biotechnological production of valuable compounds. The overwhelming volume of information often confuses scientists. This review intends to provide a guide to choosing the most efficient way to analyze a new sequence or to collect information on a gene or protein of interest by applying current publicly available databases and Web services. Recently developed tools that allow functional assignment of genes, mainly based on sequence similarity of the deduced amino acid sequence, using the currently available and increasing biological databases will be discussed.

Computational Biology↗

A computational study of Shewanella oneidensis MR-1: structural prediction and functional inference of hypothetical proteins.

The genomes of many organisms have been sequenced in the last 5 years. Typically about 30% of predicted genes from a newly sequenced genome cannot be given functional assignments using sequence comparison methods. In these situations three-dimensional structural predictions combined with a suite of computational tools can suggest possible functions for these hypothetical proteins. Suggesting functions may allow better interpretation of experimental data (e.g., microarray data and mass spectroscopy data) and help experimentalists design new experiments. In this paper, we focus on three hypothetical proteins of Shewanella oneidensis MR-1 that are potentially related to iron transport/metabolism based on microarray experiments. The threading program PROSPECT was used for protein structural predictions and functional annotation, in conjunction with literature search and other computational tools. Computational tools were used to perform transmembrane domain predictions, coiled coil predictions, signal peptide predictions, sub-cellular localization predictions, motif prediction, and operon structure evaluations. Combined computational results from all tools were used to predict roles for the hypothetical proteins. This method, which uses a suite of computational tools that are freely available to academic users, can be used to annotate hypothetical proteins in general.

ATP-Binding Cassette Transporters↗

BlastXtract--a new way of exploring translated searches.

SUMMARY: Searches of translated, unannotated genomic DNA sequences against protein databases is a useful early-stage method for discovering protein homologues encoded by the sequence, but generates huge amounts of output data that quickly become impregnable. BlastXtract is a web-based tool for managing and visualizing results from large translated BLAST and FastA searches. It combines the speed and storage benefits of relational database management systems with an easy-to-use graphical navigation map, and greatly facilitates the early exploration of genomic sequence. AVAILABILITY: BlastXtract can be downloaded from http://bioinfo.ucc.ie/blastxtract/.

Amino Acid Sequence↗

PANDIT: an evolution-centric database of protein and associated nucleotide domains with inferred trees.

PANDIT is a database of homologous sequence alignments accompanied by estimates of their corresponding phylogenetic trees. It provides a valuable resource to those studying phylogenetic methodology and the evolution of coding-DNA and protein sequences. Currently in version 17.0, PANDIT comprises 7738 families of homologous protein domains; for each family, DNA and corresponding amino acid sequence multiple alignments are available together with high quality phylogenetic tree estimates. Recent improvements include expanded methods for phylogenetic tree inference, assessment of alignment quality and a redesigned web interface, available at the URL http://www.ebi.ac.uk/goldman-srv/pandit.

Databases, Nucleic Acid↗

SWAPSC: sliding window analysis procedure to detect selective constraints.

UNLABELLED: Sliding-window analysis procedure to detect selective constraints (SWAPSC) is a software system to dissect the constraints on the evolution of protein-coding genes. The program estimates rates of nucleotide substitutions at specific codon regions in each branch of a phylogenetic tree. The program uses several sets of simulated sequence alignments to estimate the probability of synonymous and non-synonymous nucleotide substitutions. Thereafter, a statistical analysis is conducted to determine the optimum window size to detect selective constraints. Finally, the optimum window size is slid along the real alignment and a test for significance of the estimated number of synonymous and non-synonymous nucleotide substitutions in each sliding step is conducted. A number of friendly useful output files is generated. AVAILABILITY: SWAPSC is available at http://www.may.ie/academic/biology/staff/mfmolecevolandbioinf.shtml distribution versions for both Linux and Windows operating systems are available, including manual and example files.

Algorithms↗

Computational analysis of evolution and conservation in a protein superfamily.

Many gene superfamilies have hundreds or thousands of members and hence pose a significant challenge when performing a large-scale phylogenetic analysis. Derivation of the most accurate alignment possible and inference of evolutionary relationships (with an appropriate measure of confidence) are significant "bottlenecks" in the process. A generally applicable strategy is outlined for identifying and aligning sequences, performing simple analysis of the resulting alignment, and inferring evolutionary relationships. Reference is made to the serpin superfamily. The 'partition cluster' method, a relatively rapid technique for extracting underlying associations from phylogenetic bootstrap trees, is also presented.

Computational Biology↗

From sequences to a functional unit.

Functional insights at the gene product level would help the drug discovery industry to effectively tap targets for therapeutics and biomedical applications. A complete functional unit can be multidomain, and it is the co-occurrence and interaction of these multiple domains that determine the function and functional diversity of their gene products. With at least 10% of genes from complete genomes existing in fused form, identifying gene fusion events helps us categorize the protein universe into distinct functional units with only sequence information.

Animals↗

Protein-directed DNA structure II. Raman spectroscopy of a leucine zipper bZIP complex.

Mechanisms of transcription may involve protein-directed changes in DNA structure and DNA-directed changes in protein structure. We have employed Raman spectroscopy to characterize vibrational signatures associated with such induced molecular fitting for two classes of transcription factors-the basic leucine-zipper (bZIP) motif and the high-mobility-group (HMG) box-each with a DNA target site. Results for bZIP are described here; findings for the HMG-box are reported in the preceding paper in this issue [Benevides, J. M., Chan, G., Lu, X.-J., Olson, W. K., Weiss, M. A., and Thomas, G. J., Jr. (2000) Biochemistry 39, 537-547]. The yeast activator GCN4 provides a well-studied example of bZIP recognition, wherein B-DNA serves essentially as a template for protein folding. Analysis of Raman spectra of the 57-residue GCN4 bZIP domain, its AP-1 binding site, and their specific complex confirms a DNA-induced increase in alpha-helicity, attributable to folding of GCN4 basic arms with virtually no change in B-DNA structure, consistent with previous X-ray and NMR structure determinations. The absence of DNA perturbations in the bZIP model contrasts sharply with the HMG box, where DNA structure perturbations predominate. The bZIP and HMG-box models represent two opposing extremes in a range of induced fits identifiable by Raman spectroscopy. Previously characterized lambda repressor/operator complexes [Benevides, J. M., Weiss, M. A., and Thomas, G. J. (1994) J. Biol. Chem. 269, 10869-10878] occupy an intermediate position within this range. A comprehensive tabulation of Raman markers proposed as diagnostic of different protein/DNA recognition motifs is presented. The results are analyzed in terms of available DNA crystal structures (Nucleic Acid Database) to identify details of DNA conformation that correlate with specific Raman recognition markers.

Base Sequence↗