PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Prediction of an inter-residue interaction in the chaperonin GroEL from multiple sequence alignment is confirmed by double-mutant cycle analysis.

A search for co-ordinated amino acid changes in the hsp60 family of chaperonins suggested that cysteine residues at positions 137 and 518 in the Escherichia coli chaperonin GroEL may interact with each other. In order to determine whether this interaction indeed exists we constructed a double-mutant cycle comprising wild-type GroEL, the single mutants Cys137-->Ser and Cys518-->Ser and the corresponding double mutant. The effects of the two mutations on the function of GroEL, in assisting the refolding of a non-folded protein substrate (rhodanese), are shown to be non-additive. It is also shown that ADP by itself specifically destabilizes the Cys518-->Ser mutant GroEL particle with this effect being suppressed in the double mutant. The observed pattern of co-ordinated mutations in the hsp60 family of chaperonins is thus shown to reflect a real interaction, though most likely indirect, between Cys137 and Cys518 in GroEL. Our study demonstrates that patterns of co-ordinated mutations combined with double-mutant cycle analysis can provide structural information on interactions in a protein without an available three-dimensional structure at atomic resolution.

Bacterial Proteins↗

Gap mapping: a paradigm for aligning two sequences.

Pairwise sequence alignment is one of the most essential tools in comparative genomic sequence analysis. It is used to compare the sequences of genes and proteins with the aim of inferring structural, functional and evolutionary relationships. However, current 'mainstream' alignment algorithms have optimisation criteria based primarily on computational efficiency using parameters such as gap penalties, which are not biologically motivated. In addition, current alignment algorithms such as the Smith and Waterman technique provide a single alignment that could be sensitive to rather arbitrary choices in parameters such as gap penalties. This paper explores the range of properties resulting from posing the alignment problem more as a 'mapping gaps in sequences' exercise. We argue that this approach is intuitive and provides greater control over the number of gaps placed within an alignment. This type of approach was proposed by Sankoff (1972), but unfortunately has not received much attention. We report and discuss our findings by comparing this approach to other techniques using structurally confirmed aligned sequences from a benchmark alignment database. Interestingly, this approach consistently provides optimal and near optimal alignments and is thus a viable approach to sequence alignment.

Algorithms↗

Predicting reliable regions in protein alignments from sequence profiles.

For applications such as comparative modelling one major issue is the reliability of sequence alignments. Reliable regions in alignments can be predicted using sub-optimal alignments of the same pair of sequences. Here we show that reliable regions in alignments can also be predicted from multiple sequence profile information alone. Alignments were created for a set of remotely related pairs of proteins using five different test methods. Structural alignments were used to assess the quality of the alignments and the aligned positions were scored using information from the observed frequencies of amino acid residues in sequence profiles pre-generated for each template structure. High-scoring regions of these profile-derived alignment scores were a good predictor of reliably aligned regions. These profile-derived alignment scores are easy to obtain and are applicable to any alignment method. They can be used to detect those regions of alignments that are reliably aligned and to help predict the quality of an alignment. For those residues within secondary structure elements, the regions predicted as reliably aligned agreed with the structural alignments for between 92% and 97.4% of the residues. In loop regions just under 92% of the residues predicted to be reliable agreed with the structural alignments. The percentage of residues predicted as reliable ranged from 32.1% for helix residues to 52.8% for strand residues. This information could also be used to help predict conserved binding sites from sequence alignments. Residues in the template that were identified as binding sites, that aligned to an identical amino acid residue and where the sequence alignment agreed with the structural alignment were in highly conserved, high scoring regions over 80% of the time. This suggests that many binding sites that are present in both target and template sequences are in sequence-conserved regions and that there is the possibility of translating reliability to binding site prediction.

Algorithms↗

ALIGNMENT SERVICE: creation and processing of alignments of sequences of unlimited length.

A package for the creation and processing of multiple sequence alignment is described. There is no limit on the lengths of the processed nucleotide or amino acid sequences, and the number of sequences in the alignment is also unlimited. The main groups of functions are: a semiautomatic alignment editor; a wide set of functions for technical processing of alignments; nucleotide alignment mapping and translation; and similarity search functions. A user-friendly interface and a set of generally used file actions provide a special operational subsystem for everyday tasks.

Amino Acid Sequence↗

An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.

The sensitivity of most protein sequence alignment methods depends strongly on the quality of the comparison matrices used. These matrices, which assign weights or similarity scores to every possible amino acid substitution pair, are utilized to differentiate amongst the various possible alignments of two or more sequences. There are many ways to generate these exchange weights and new matrices are constantly published. There has been no overall assessment of these various matrices when applied in different alignment techniques and over many protein folds and families, both close and distant and with the use of several gap penalty values. In this work, a set of amino acid sequences matched by superposition of known protein tertiary topologies is used to test the alignment accuracy of the different method/matrix/penalty combinations. The comparisons show relatively similar results for the top scoring matrices, a preference for the global alignment method of Needleman and Wunsch, and the importance of matrix modification and optimized gap penalties. The relationship between the percentage identity in a resulting alignment and the level of correctness to be expected are given for the top-performing matrix, resulting in a better definition of the so-called "twilight zone". Estimates are made for the probability that two sequences, aligned at a certain level of residue percentage identity, are in fact unrelated.

Amino Acid Sequence↗

Parallel hardware for sequence comparison and alignment.

Sequence comparison, a vital research tool in computational biology, is based on a simple O(n2) algorithm that easily maps to a linear array of processors. This paper reviews and compares high-performance sequence analysis on general-purpose supercomputers and single-purpose reconfigurable, and programmable co-processors. The difficulty of comparing hardware from published performance figures is also noted.

Algorithms↗

PdbAlign, PdbDist and DistAlign: tools to aid in relating sequence variability to structure.

Many sequence analysis problems involve consideration of a multiple sequence alignment where the 3-dimensional structure of one (or more) of the aligned sequences is known. In such cases, it is useful to map the sequence variability onto the atomic co-ordinates of known structure. If the structure also includes a bound ligand (or the location of the active site is known), each column position in the multiple sequence alignment may be annotated with its 'distance' from the binding site. These annotations, together with a measure of sequence variability, provide additional insights into drug specificity, for example among viral mutants. This paper describes several useful programs that automate this analysis.

Amino Acid Sequence↗

Conserved sequence motifs, alignment, and secondary structure for the third domain of animal 12S rRNA.

Secondary structure models are an important step for aligning sequences, understanding probabilities of nucleotide substitutions, and evaluating the reliability of phylogenetic reconstructions. A set of conserved sequence motifs is derived from comparative sequence analysis of 184 invertebrate and vertebrate taxa (including many taxa from the same genera, families, and orders) with reference to a secondary structure model for domain III of animal mitochondrial small subunit (12S) ribosomal RNA. A template is presented to assist with secondary structure drawing. Our model is similar to previous models but is more specific to mitochondrial DNA, fitting both invertebrate and vertebrate groups, including taxa with markedly different nucleotide compositions. The second half of the domain III sequence can be difficult to align precisely, even when secondary structure information is considered. This is especially true for comparisons of anciently diverged taxa, but well-conserved motifs assist in determining biologically meaningful alignments. Patterns of conservation and variability in both paired and unpaired regions make differential phylogenetic weighting in terms of "stems" and "loops" unsatisfactory. We emphasize looking carefully at the sequence data before and during analyses, and advocate the use of conserved motifs and other secondary structure information for assessing sequencing fidelity.

Animals↗

Self-consistently optimized statistical mechanical energy functions for sequence structure alignment.

A quantitative form of the principle of minimal frustration is used to obtain from a database analysis statistical mechanical energy functions and gap parameters for aligning sequences to three-dimensional structures. The analysis that partially takes into account correlations in the energy landscape improves upon the previous approximations of Goldstein et al. (1994, 1995) (Goldstein R, Luthey-Schulten Z, Wolynes P, 1994, Proceedings of the 27th Hawaii International Conference on System Sciences. Los Alamitos, California: IEEE Computer Society Press. pp 306-315; Goldstein R, Luthey-Schulten Z, Wolynes P, 1995, In: Elber R, ed. New developments in theoretical studies of proteins. Singapore: World Scientific). The energy function allows for ordering of alignments based on the compatibility of a sequence to be in a given structure (i.e., lowest energy) and therefore removes the necessity of using percent identity or similarity as scoring parameters. The alignments produced by the energy function on distant homologues with low percent identity (less than 21%) are generally better than those generated with evolutionary information. The lowest energy alignment generated with the energy function for sequences containing prosite signatures but unknown structures is a structure containing the same prosite signature, providing a check on the robustness of the algorithm. Finally, the energy function can make use of known experimental evidence as constraints within the alignment algorithm to aid in finding the correct structural alignment.

Data Interpretation, Statistical↗

Sequence divergence analysis for the prediction of seven-helix membrane protein structures: II. A 3-D model of human rhodopsin.

A three-dimensional (3-D) model of the transmembrane domain of human rhodopsin was predicted from the sequence divergence analysis of 42 sequences of rhodopsins and visual pigments without a template. The prediction steps include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, their secondary structure and packing shape in a helix bundle, prediction of side-chain conformations and structure refinement. The identification of the retinal binding site was assisted by its known covalent linkage with K296. The structural features of the predicted 3-D model are in good agreement with a low resolution electron density map of bovine rhodopsin and with residues in contact with retinal as determined experimentally.

Amino Acid Sequence↗

Genetic relatedness of human DNA polymerase beta and terminal deoxynucleotidyltransferase.

The Protein Identification Resource (PIR) protein sequence data bank was searched for sequence similarity between known proteins and human DNA polymerase beta (Pol beta) or human terminal deoxynucleotidyltransferase (TdT). Pol beta and TdT were found to exhibit amino acid sequence similarity only with each other and not with any other of the 4750 entries in release 12.0 of the PIR data bank. Optimal amino acid sequence alignment of the entire 39-kDa Pol beta polypeptide with the C-terminal two thirds of TdT revealed 24% identical aa residues and 21% conservative aa substitutions. The Monte Carlo score of 12.6 for the entire aligned sequences indicates highly significant aa sequence homology. The hydropathicity profiles of the aligned aa sequences were remarkably similar throughout, suggesting structural similarity of the polypeptides. The most significant regions of homology are aa residues 39-224 and 311-333 of Pol beta vs. aa residues 191-374 and 484-506 of TdT. In addition, weaker homology was seen between a large portion of the 'nonessential' N-terminal end of TdT (aa residues 33-130) and the first region of strong homology between the two proteins (aa residues 31-128 of Pol beta and aa residues 183-280 of TdT), suggestive of genetic duplication within the ancestral gene. On the basis of nucleotide differences between conserved regions of Pol beta and TdT genes (aligned according to optimally aligned aa sequences) it was estimated that Pol beta and TdT diverged on the order of 250 million years ago, corresponding roughly to a time before radiation of mammals and birds.

Amino Acid Sequence↗

Recombination Analysis Tool (RAT): a program for the high-throughput detection of recombination.

MOTIVATION: Recombination can be a prevailing drive in shaping genome evolution. RAT (Recombination Analysis Tool) is a Java-based tool for investigating recombination events in any number of aligned sequences (protein or DNA) of any length (short viral sequences to full genomes). It is an uncomplicated and intuitive application and allows the user to view only the regions of sequence alignments they are interested in. RESULTS: RAT was applied to viral sequences. Its utility was demonstrated through the detection of a known recombinant of HIV and a detailed analysis of Noroviruses, the most common cause of viral gastroenteritis in humans. AVAILABILITY: RAT, along with a user's guide, is freely available from http://jic-bioinfo.bbsrc.ac.uk/bioinformatics-research/staff/graham_etherington/RAT.htm.

DNA, Viral↗

An approach to improving multiple alignments of protein sequences using predicted secondary structure.

The object of this work was to improve multiple sequence alignments using public-domain software and methods as far as possible. A method is described where the secondary structure of proteins is predicted and this information, coupled with a simplified description of the amino acids, is used to produce multiple sequence alignments. This method improved the accuracy of the resulting alignments by between 5 and 14% when compared with full sequence profile alignments (as scored against structural alignments). These improved alignments were used to predict the secondary structure of the sequences they contain. The resultant predictions were more accurate than those produced from less optimal alignments. An improvement of 6% for a three-state (helix, sheet and coil) prediction was observed when using the best alignment from the method presented here and the alignment obtained using sequence only. The method makes use of public domain software and all the associated files required to repeat the work are available from the primary author.

Algorithms↗

Multiple alignment of protein sequences with repeats and rearrangements.

Multiple sequence alignments are the usual starting point for analyses of protein structure and evolution. For proteins with repeated, shuffled and missing domains, however, traditional multiple sequence alignment algorithms fail to provide an accurate view of homology between related proteins, because they either assume that the input sequences are globally alignable or require locally alignable regions to appear in the same order in all sequences. In this paper, we present ProDA, a novel system for automated detection and alignment of homologous regions in collections of proteins with arbitrary domain architectures. Given an input set of unaligned sequences, ProDA identifies all homologous regions appearing in one or more sequences, and returns a collection of local multiple alignments for these regions. On a subset of the BAliBASE benchmarking suite containing curated alignments of proteins with complicated domain architectures, ProDA performs well in detecting conserved domain boundaries and clustering domain segments, achieving the highest accuracy to date for this task. We conclude that ProDA is a practical tool for automated alignment of protein sequences with repeats and rearrangements in their domain architecture.

Algorithms↗

A greedy algorithm for aligning DNA sequences.

For aligning DNA sequences that differ only by sequencing errors, or by equivalent errors from other sources, a greedy algorithm can be much faster than traditional dynamic programming approaches and yet produce an alignment that is guaranteed to be theoretically optimal. We introduce a new greedy alignment algorithm with particularly good performance and show that it computes the same alignment as does a certain dynamic programming algorithm, while executing over 10 times faster on appropriate data. An implementation of this algorithm is currently used in a program that assembles the UniGene database at the National Center for Biotechnology Information.

Algorithms↗

Alignment of molecular sequences seen as random path analysis.

We propose a generating functional method--random path analysis (RPA)--that generalizes the classical dynamic programming (DP) method widely used in sequence alignments. For a given cost function, DP is a deterministic method that finds an optimal alignment by minimizing the total cost function for all possible alignments. By allowing uncertainty, RPA is a statistical method that weights fluctuating alignments by probabilities. Therefore, DP maybe thought of as the deterministic limit of RPA when the fluctuations approach zero. DP is the method of choice if one is only interested in optimal alignment. But we argue that, when information beyond the optimal alignment is desired, RPA gives a natural extension of DP for biological applications. As an algebraic approach, RPA is computationally intensive for long sequences, but it can provide better parametric control for developing analytical or perturbational results and it is more informative and biologically relevant. The idea of RPA opens up new opportunities for simulational approaches and more importantly it suggests a novel hardware implementation that has the potential of improving the way a sequence alignment is done. Here we focus on deriving a mathematically rigorous solution to RPA both in its combinatorial form and in its graphical representation; this puts DP in logical perspective under a more general conceptual framework.

Animals↗

Statistical potential-based amino acid similarity matrices for aligning distantly related protein sequences.

Aligning distantly related protein sequences is a long-standing problem in bioinformatics, and a key for successful protein structure prediction. Its importance is increasing recently in the context of structural genomics projects because more and more experimentally solved structures are available as templates for protein structure modeling. Toward this end, recent structure prediction methods employ profile-profile alignments, and various ways of aligning two profiles have been developed. More fundamentally, a better amino acid similarity matrix can improve a profile itself; thereby resulting in more accurate profile-profile alignments. Here we have developed novel amino acid similarity matrices from knowledge-based amino acid contact potentials. Contact potentials are used because the contact propensity to the other amino acids would be one of the most conserved features of each position of a protein structure. The derived amino acid similarity matrices are tested on benchmark alignments at three different levels, namely, the family, the superfamily, and the fold level. Compared to BLOSUM45 and the other existing matrices, the contact potential-based matrices perform comparably in the family level alignments, but clearly outperform in the fold level alignments. The contact potential-based matrices perform even better when suboptimal alignments are considered. Comparing the matrices themselves with each other revealed that the contact potential-based matrices are very different from BLOSUM45 and the other matrices, indicating that they are located in a different basin in the amino acid similarity matrix space.

Amino Acids↗