PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Sequence alignments, variabilities, and vagaries.

It seems as if the algorithms and weighting matrices for multiple sequence alignments of the highly divergent members of the P450 gene superfamily have advanced to the point that unknown proteins can be aligned to structurally known members with reasonable accuracy. As stated earlier, the alignment tends to break down at gaps in the sequence alignments, but these regions can be improved manually. This type of alignment and analysis is especially useful for extracting and analyzing the various genome databases. Variations of the conservation analysis can be used to identify charged and uncharged residues that may be important in domain/domain interactions with redox partners or effector molecules (e.g., cytochrome b5). From these alignments and with comparative analysis within families and across P450 families, one can readily obtain an estimation of those residues that might be involved in substrate binding, in redox partner interaction, and in the catalytic mechanism.

Amino Acid Sequence↗

Efficient methods for multiple sequence alignment with guaranteed error bounds.

Multiple string (sequence) alignment is a difficult and important problem in computational biology, where it is central in two related tasks: finding highly conserved subregions or embedded patterns of a set of biological sequences (strings of DNA, RNA or amino acids), and inferring the evolutionary history of a set of taxa from their associated biological sequences. Several precise measures have been proposed for evaluating the goodness of a multiple alignment, but no efficient methods are known which compute the optimal alignment for any of these measures in any but small cases. In this paper, we consider two previously proposed measures, and give two computationaly efficient multiple alignment methods (one for each measure) whose deviation from the optimal value is guaranteed to be less than a factor of two. This is the novel feature of thse methods. but the methods have additional virtues as well. For both methods, the guaranteed bounds are much smaller than two when the number of strings is small (1.33 for three strings of any length); for one of the methods we give a related randomized method which is much faster and which gives, with high probability, multiple alignments with fairly small error bounds; and for the other measure, the method given yields a non-obvious lower bound on the optimal alignment.

Amino Acid Sequence↗

RDP2: recombination detection and analysis from sequence alignments.

UNLABELLED: RDP2 is a Windows 95/XP program that examines nucleotide sequence alignments and attempts to identify recombinant sequences and recombination breakpoints using 10 published recombination detection methods, including GENECONV, BOOTSCAN, MAXIMUM chi(2), CHIMAERA and SISTER SCANNING. The program enables fast automated analysis of large alignments (up to 300 sequences containing 13 000 sites), and interactive exploration, management and verification of results with different recombination detection and tree drawing methods. AVAILABILITY: RDP2 is available free from the RDP2 website (http://darwin.uvigo.es/rdp/rdp.html) CONTACT: darren@science.uct.ac.za SUPPLEMENTARY INFORMATION: Detailed descriptions of RDP2 and the methods it implements are included in the program manual, which can be downloaded from the RDP2 website.

Algorithms↗

Simulated annealing algorithm for the multiple sequence alignment problem: the approach of polymers in a random medium.

We propose a probabilistic algorithm to solve the multiple sequence alignment problem. The algorithm is a simulated annealing that exploits the representation of the multiple alignment between D sequences as a directed polymer in D dimensions. Within this representation we can easily track the evolution of the alignment through local moves of low computational cost. In contrast with other probabilistic algorithms proposed to solve this problem, our approach allows the creation and deletion of gaps without extra computational cost. The algorithm was tested by aligning proteins from the kinase family. When D=3 the results are consistent with those obtained using a complete algorithm. For D>3 where the complete algorithm fails, we show that our algorithm still converges to reasonable alignments. We also study the space of solutions obtained and show that depending on the number of sequences aligned the solutions are organized in different ways, suggesting a possible source of errors for progressive algorithms. Finally, we test our algorithm in artificially generated sequences and prove that it may perform better than progressive algorithms. Moreover, in those cases in which a progressive algorithm works better, its solution may be taken as an initial condition of our algorithm and, again, we obtain alignments with lower scores and more relevant from the biological point of view.

Algorithms↗

LALNVIEW: a graphical viewer for pairwise sequence alignments.

LALNVIEW is a graphical program for visualising local alignments between two sequences (protein or nucleic acids). Sequences are represented by coloured rectangles to give an overall picture of their similarities. LALNVIEW can display sequence features (exon, intron, active site, domain, propeptide, etc.) along with the alignment. When using LALNVIEW through our Web servers, sequence features are automatically extracted from database annotations (SWISS-PROT, GenBank, EMBL or HOVERGEN) and displayed with the alignment. LALNVIEW is a useful tool for analysing pairwise sequence alignments and for making the link between sequence homology and what is known about the structure or function of sequences. LALNVIEW executables for UNIX, Macintosh and PC computers are freely available from our server (http:// expasy.hcuge.ch/sprot/lalnview.html).

Acyltransferases↗

Discriminating between rate heterogeneity and interspecific recombination in DNA sequence alignments with phylogenetic factorial hidden Markov models.

MOTIVATION: A recently proposed method for detecting recombination in DNA sequence alignments is based on the combination of hidden Markov models (HMMs) with phylogenetic trees. Although this method was found to detect breakpoints of recombinant regions more accurately than most existing techniques, it inherently fails to distinguish between recombination and rate variation. In the present paper, we propose to marry the phylogenetic tree to a factorial HMM (FHMM). The states of the first hidden chain represent tree topologies, whereas the states of the second independent hidden chain represent different global scaling factors of the branch lengths. Inference is done in terms of a hierarchical Bayesian model, where parameters and hidden states are sampled from the posterior distribution with Gibbs sampling. RESULTS: We have tested the proposed model on various synthetic and real-world DNA sequence alignments. The simulation results suggest that as opposed to the standard phylogenetic HMM, the phylogenetic FHMM clearly distinguishes between recombination and rate heterogeneity and thereby avoids the prediction of spurious recombinant regions. AVAILABILITY: The proposed method has been implemented in a MATLAB package that extends Kevin Murphy's HMM toolbox. Software and data used in our study are available from http://www.bioss.sari.ac.uk/~dirk/Supplements

Algorithms↗

Multiple sequence alignment by a pairwise algorithm.

An algorithm is described that processes the results of a conventional pairwise sequence alignment program to automatically produce an unambiguous multiple alignment of many sequences. Unlike other, more complex, multiple alignment programs, the method described here is fast enough to be used on almost any multiple sequence alignment problem.

Algorithms↗

Iterative sequence/secondary structure search for protein homologs: comparison with amino acid sequence alignments and application to fold recognition in genome databases.

MOTIVATION: Sequence alignment techniques have been developed into extremely powerful tools for identifying the folding families and function of proteins in newly sequenced genomes. For a sufficiently low sequence identity it is necessary to incorporate additional structural information to positively detect homologous proteins. We have carried out an extensive analysis of the effectiveness of incorporating secondary structure information directly into the alignments for fold recognition and identification of distant protein homologs. A secondary structure similarity matrix based on a database of three-dimensionally aligned proteins was first constructed. An iterative application of dynamic programming was used which incorporates linear combinations of amino acid and secondary structure sequence similarity scores. Initially, only primary sequence information is used. Subsequently contributions from secondary structure are phased in and new homologous proteins are positively identified if their scores are consistent with the predetermined error rate. RESULTS: We used the SCOP40 database, where only PDB sequences that have 40% homology or less are included, to calibrate homology detection by the combined amino acid and secondary structure sequence alignments. Combining predicted secondary structure with sequence information results in a 8-15% increase in homology detection within SCOP40 relative to the pairwise alignments using only amino acid sequence data at an error rate of 0.01 errors per query; a 35% increase is observed when the actual secondary structure sequences are used. Incorporating predicted secondary structure information in the analysis of six small genomes yields an improvement in the homology detection of approximately 20% over SSEARCH pairwise alignments, but no improvement in the total number of homologs detected over PSI-BLAST, at an error rate of 0.01 errors per query. However, because the pairwise alignments based on combinations of amino acid and secondary structure similarity are different from those produced by PSI-BLAST and the error rates can be calibrated, it is possible to combine the results of both searches. An additional 25% relative improvement in the number of genes identified at an error rate of 0.01 is observed when the data is pooled in this way. Similarly for the SCOP40 dataset, PSI-BLAST detected 15% of all possible homologs, whereas the pooled results increased the total number of homologs detected to 19%. These results are compared with recent reports of homology detection using sequence profiling methods. AVAILABILITY: Secondary structure alignment homepage at http://lutece.rutgers.edu/ssas CONTACT: anders@rutchem.rutgers.edu; ronlevy@lutece.rutgers.edu SUPPLEMENTARY INFORMATION: Genome sequence/structure alignment results at http://lutece.rutgers.edu/ss_fold_predictions.

Algorithms↗

Quantifying the local reliability of a sequence alignment.

We present a method for attributing a measure of reliability to a residue pair in an optimal alignment of two protein sequences. Validation based on a database of structurally correct alignments [Pascarella and Argos (1992) Protein Engng, 5, 121-137] shows that correctly aligned parts of a sequence alignment systematically receive high scores in this measure. The higher the sequence similarity between two sequences, the larger is the fraction found of the correct parts of the alignment. We used these observations to design a program that draws a reliability curve along an optimal alignment reflecting the chances for each residue pair to be aligned correctly.

Algorithms↗

MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.

A multiple sequence alignment program, MAFFT, has been developed. The CPU time is drastically reduced as compared with existing methods. MAFFT includes two novel techniques. (i) Homo logous regions are rapidly identified by the fast Fourier transform (FFT), in which an amino acid sequence is converted to a sequence composed of volume and polarity values of each amino acid residue. (ii) We propose a simplified scoring system that performs well for reducing CPU time and increasing the accuracy of alignments even for sequences having large insertions or extensions as well as distantly related sequences of similar length. Two different heuristics, the progressive method (FFT-NS-2) and the iterative refinement method (FFT-NS-i), are implemented in MAFFT. The performances of FFT-NS-2 and FFT-NS-i were compared with other methods by computer simulations and benchmark tests; the CPU time of FFT-NS-2 is drastically reduced as compared with CLUSTALW with comparable accuracy. FFT-NS-i is over 100 times faster than T-COFFEE, when the number of input sequences exceeds 60, without sacrificing the accuracy.

Computer Simulation↗

NAST: a multiple sequence alignment server for comparative analysis of 16S rRNA genes.

Microbiologists conducting surveys of bacterial and archaeal diversity often require comparative alignments of thousands of 16S rRNA genes collected from a sample. The computational resources and bioinformatics expertise required to construct such an alignment has inhibited high-throughput analysis. It was hypothesized that an online tool could be developed to efficiently align thousands of 16S rRNA genes via the NAST (Nearest Alignment Space Termination) algorithm for creating multiple sequence alignments (MSA). The tool was implemented with a web-interface at http://greengenes.lbl.gov/NAST. Each user-submitted sequence is compared with Greengenes' 'Core Set', comprising approximately 10,000 aligned non-chimeric sequences representative of the currently recognized diversity among bacteria and archaea. User sequences are oriented and paired with their closest match in the Core Set to serve as a template for inserting gap characters. Non-16S data (sequence from vector or surrounding genomic regions) are conveniently removed in the returned alignment. From the resulting MSA, distance matrices can be calculated for diversity estimates and organisms can be classified by taxonomy. The ability to align and categorize large sequence sets using a simple interface has enabled researchers with various experience levels to obtain bacterial and archaeal community profiles.

Algorithms↗

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence↗

Low linkage disequilibrium indicative of recombination in foot-and-mouth disease virus gene sequence alignments.

We have applied tests for detecting recombination to genes of foot-and-mouth disease virus (FMDV). Our approach estimated summary statistics of linkage disequilibrium (LD), which are sensitive to recombination. Using the genealogical relationships, rate heterogeneity and mutation parameters estimated from individual sets of aligned gene sequences, we simulated matching RNA sequence datasets without recombination. These simulated datasets allowed for recurrent mutations at any site to mimic homoplasy in virus sequence data and allow construction of null distributions for LD parameters expected in the absence of recombination. We tested for recombination in two ways: by comparing LD in observed data with corresponding null distributions obtained from simulated data; and by testing for a negative relationship between observed LD between pairs of polymorphic nucleotide sites and inter-site distance. We applied these tests to six FMDV datasets from four serotypes and found some evidence for recombination in all of them.

Capsid Proteins↗

PRALINE: a multiple sequence alignment toolbox that integrates homology-extended and secondary structure information.

PRofile ALIgNEment (PRALINE) is a fully customizable multiple sequence alignment application. In addition to a number of available alignment strategies, PRALINE can integrate information from database homology searches to generate a homology-extended multiple alignment. PRALINE also provides a choice of seven different secondary structure prediction programs that can be used individually or in combination as a consensus for integrating structural information into the alignment process. The program can be used through two separate interfaces: one has been designed to cater to more advanced needs of researchers in the field, and the other for standard construction of high confidence alignments. The web-based output is designed to facilitate the comprehensive visualization of the generated alignments by means of five default colour schemes based on: residue type, position conservation, position reliability, residue hydrophobicity and secondary structure, depending on the options set. A user can also define a custom colour scheme by selecting which colour will represent one or more amino acids in the alignment. All generated alignments are also made available in the PDF format for easy figure generation for publications. The grouping of sequences, on which the alignment is based, can also be visualized as a dendrogram. PRALINE is available at http://ibivu.cs.vu.nl/programs/pralinewww/.

Computer Graphics↗

Detecting recombination in 4-taxa DNA sequence alignments with Bayesian hidden Markov models and Markov chain Monte Carlo.

This article presents a statistical method for detecting recombination in DNA sequence alignments, which is based on combining two probabilistic graphical models: (1) a taxon graph (phylogenetic tree) representing the relationship between the taxa, and (2) a site graph (hidden Markov model) representing interactions between different sites in the DNA sequence alignments. We adopt a Bayesian approach and sample the parameters of the model from the posterior distribution with Markov chain Monte Carlo, using a Metropolis-Hastings and Gibbs-within-Gibbs scheme. The proposed method is tested on various synthetic and real-world DNA sequence alignments, and we compare its performance with the established detection methods RECPARS, PLATO, and TOPAL, as well as with two alternative parameter estimation schemes.

Base Sequence↗