PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Entropy calculator: getting the best from your multiple protein alignments.

Amino acid sequence alignment is an extremely useful tool in protein family analysis. Most family characteristics, such as the localization of functional residues, structural constraints and evolutionary relationships may be retrieved through the observation of the conservation pattern highlighted by the alignments. A quantitative score for the conservation in the alignment allows different stages of an alignment to be compared and consequently the alignment information to be efficiently exploited. Many scoring methods have been proposed during the last three decades. Claude Shannon's theory of communication (1948) paved the way for a consistent scoring of protein alignments by considering the residue (or symbol) frequency. A number of modifications have been proposed since that time, but the core statistical approach is still considered one of the best. By combining many database managing tools for treatment of protein sequences, a ClustalW software integration, a flexible symbols treatment and gap normalization functions, Entropy Calculator software has been developed. This new tool provides a global and optimal approach to multiple sequence alignment scoring by offering an easy graphic interface and a series of modification options that help in interpreting alignments and allow conservation pattern inferences to be performed.

Amino Acid Sequence↗

[Effectiveness of a procedure for aligning sense sequences to make possible restoring the true alignment].

The method is presented which may estimate the reliability of any alignment procedure on the basis of comparing between genuine alignment obtained by generating of appropriate sequences with the result of alignment procedure work. The opportunities of the method are illustrated by testing the OPAL286 program. Dependence of reliability of the results i.e. reconstruction of genuine alignment from percent of mutations and deletion-insertions are pointed. The method developed provides correct optimal choice of parameters for alignment procedure in the sense of reliability and probability of correct reconstruction of the original alignment.

Mutagenesis, Insertional↗

Fast and sensitive multiple alignment of large genomic sequences.

BACKGROUND: Genomic sequence alignment is a powerful method for genome analysis and annotation, as alignments are routinely used to identify functional sites such as genes or regulatory elements. With a growing number of partially or completely sequenced genomes, multiple alignment is playing an increasingly important role in these studies. In recent years, various tools for pair-wise and multiple genomic alignment have been proposed. Some of them are extremely fast, but often efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast search program identifies a chain of strong local sequence similarities. In a second step, regions between these anchor points are aligned using a slower but more accurate method. RESULTS: Herein, we present CHAOS, a novel algorithm for rapid identification of chains of local pair-wise sequence similarities. Local alignments calculated by CHAOS are used as anchor points to improve the running time of DIALIGN, a slow but sensitive multiple-alignment tool. We show that this way, the running time of DIALIGN can be reduced by more than 95% for BAC-sized and longer sequences, without affecting the quality of the resulting alignments. We apply our approach to a set of five genomic sequences around the stem-cell-leukemia (SCL) gene and demonstrate that exons and small regulatory elements can be identified by our multiple-alignment procedure. CONCLUSION: We conclude that the novel CHAOS local alignment tool is an effective way to significantly speed up global alignment tools such as DIALIGN without reducing the alignment quality. We likewise demonstrate that the DIALIGN/CHAOS combination is able to accurately align short regulatory sequences in distant orthologues.

Algorithms↗

A sequence alignment-independent method for protein classification.

Annotation of the rapidly accumulating body of sequence data relies heavily on the detection of remote homologues and functional motifs in protein families. The most popular methods rely on sequence alignment. These include programs that use a scoring matrix to compare the probability of a potential alignment with random chance and programs that use curated multiple alignments to train profile hidden Markov models (HMMs). Related approaches depend on bootstrapping multiple alignments from a single sequence. However, alignment-based programs have limitations. They make the assumption that contiguity is conserved between homologous segments, which may not be true in genetic recombination or horizontal transfer. Alignments also become ambiguous when sequence similarity drops below 40%. This has kindled interest in classification methods that do not rely on alignment. An approach to classification without alignment based on the distribution of contiguous sequences of four amino acids (4-grams) was developed. Interest in 4-grams stemmed from the observation that almost all theoretically possible 4-grams (20(4)) occur in natural sequences and the majority of 4-grams are uniformly distributed. This implies that the probability of finding identical 4-grams by random chance in unrelated sequences is low. A Bayesian probabilistic model was developed to test this hypothesis. For each protein family in Pfam-A and PIR-PSD, a feature vector called a probe was constructed from the set of 4-grams that best characterised the family. In rigorous jackknife tests, unknown sequences from Pfam-A and PIR-PSD were compared with the probes for each family. A classification result was deemed a true positive if the probe match with the highest probability was in first place in a rank-ordered list. This was achieved in 70% of cases. Analysis of false positives suggested that the precision might approach 85% if selected families were clustered into subsets. Case studies indicated that the 4-grams in common between an unknown and the best matching probe correlated with functional motifs from PRINTS. The results showed that remote homologues and functional motifs could be identified from an analysis of 4-gram patterns.

Algorithms↗

Preferred positions of AA and TT dinucleotides in aligned nucleosomal DNA sequences.

Multiple alignment of 118 nucleosomal DNA sequences by maximizing simultaneously match of AA dinucleotides and match of TT dinucleotides results in a pattern of the dinucleotide distributions which is characteristic of the nucleosomal DNA sequences. The AA dinucleotides are found to be distributed symmetrically relative to the TT dinucleotide distribution, around the middle point of the nucleosomal DNA sequence. The distances between major peaks of the distributions are multiples of about 10.4 bases. The peaks of the TT distribution are shifted by 6 bases downstream from the peaks of the AA distribution.

Adenine↗

ViTO: tool for refinement of protein sequence-structure alignments.

UNLABELLED: ViTO is a graphical application, including an editor, of multiple sequence alignment and a three-dimensional (3D) structure viewer. It is possible to manipulate alignments containing hundreds of sequences and to display a dozen structures. ViTO can handle so-called 'multiparts' alignments to allow the visualization of complex structures (multi-chain proteins and/or small molecules and DNA) and the editing of the corresponding alignment. The 3D viewer and the alignment editor are connected together allowing rapid refinement of sequence-structure alignment by taking advantage of the immediate visualization of resulting insertions/deletions and strict conservations in their structural context. More generally, it allows the mapping of informations about the sequence conservation extracted from the alignment onto the 3D structures in a dynamic way. ViTO is also connected to two comparative modelling programs, SCWRL and MODELLER. These features make ViTO a powerful tool to characterize protein families and to optimize the alignments for comparative modelling. AVAILABILITY: http://bioserv.cbs.cnrs.fr/VITO/DOC/. SUPPLEMENTARY INFORMATION: http://bioserv.cbs.cnrs.fr/VITO/DOC/index.html.

Amino Acid Sequence↗

Structure alignment based on coding of local geometric measures.

BACKGROUND: A structure alignment method based on a local geometric property is presented and its performance is tested in pairwise and multiple structure alignments. In this approach, the writhing number, a quantity originating from integral formulas of Vassiliev knot invariants, is used as a local geometric measure. This measure is used in a sliding window to calculate the local writhe down the length of the protein chain. By encoding the distribution of writhing numbers across all the structures in the protein databank (PDB), protein geometries are represented in a 20-letter alphabet. This encoding transforms the structure alignment problem into a sequence alignment problem and allows the well-established algorithms of sequence alignment to be employed. Such geometric alignments offer distinct advantages over structural alignments in Cartesian coordinates as it better handles structural subtleties associated with slight twists and bends that distort one structure relative to another. RESULTS: The performance of programs for pairwise local alignment (TLOCAL) and multiple alignment (TCLUSTALW) are readily adapted from existing code for Smith-Waterman pairwise alignment and for multiple sequence alignment using CLUSTALW. The alignment algorithms employed a blocked scoring matrix (TBLOSUM) generated using the frequency of changes in the geometric alphabet of a block of protein structures. TLOCAL was tested on a set of 10 difficult proteins and found to give high quality alignments that compare favorably to those generated by existing pairwise alignment programs. A set of protein comparison involving hinged structures was also analyzed and TLOCAL was seen to compare favorably to other alignment methods. TCLUSTALW was tested on a family of protein kinases and reveal conserved regions similar to those previously identified by a hand alignment. CONCLUSION: These results show that the encoding of the writhing number as a geometric measure allow high quality structure alignments to be generated using standard algorithms of sequence alignment. This approach provides computationally efficient algorithms that allow fast database searching and multiple structure alignment. Because the geometric measure can employ different window sizes, the method allows the exploration of alignments on different, well-defined length scales.

Algorithms↗

Protein sequence-structure alignment based on site-alignment probabilities.

A protein sequence-structure alignment method for database searches is examined on how effectively this method together with a simple scoring function previously developed can identify compatibilities between sequences and structures of proteins. The scoring function consists of pairwise contact energies, repulsive packing potentials of residues for overly dense arrangement and short-range potentials for secondary structures. Pairwise contact interactions in a sequence-structure alignment are evaluated in a mean field approximation on the basis of probabilities of site pairs to be aligned. Gap penalties are assumed to be proportional to the number of contacts at each residue position, and as a result gaps will be more frequently placed on protein surfaces than in cores. In addition to minimum energy alignments, we use probability alignments made by successively aligning site pairs in order by pairwise alignment probabilities. Results show that the present energy function and alignment method can detect well both folds compatible with a given sequence and, inversely, sequences compatible with a given fold. Probability alignments consisting of most reliable site pairs only can yield small root mean square deviations, and including less reliable pairs increases the deviations. Remarkably, by this method some individual sequence-structure pairs are detected having only 5-20% sequence identity.

17-Hydroxysteroid Dehydrogenases↗

Phylogenetic relationships of entomopathogenic nematodes (Heterorhabditidae and Steinernematidae) inferred from partial 18S rRNA gene sequences.

Aligned 265-bp sequences of partial 18S rRNA gene were used to infer phylogenetic relationships among entomopathogenic nematodes by using maximum parsimony and likelihood methods. Phylogenetic analyses support Heterorhabditidae and Steinernematidae belonging to different monophylies. There was more sequence divergence in Steinernema species than in Heterorhabditis species. These results are congruent with the phylogenies based on morphological, life cycle, and distributional evidence. Examination of all trees within 1% of the length of the most parsimonious trees and bootstrap analyses support most relationships among Steinernema species but the relationships among Heterorhabditis species were not supported. We suggest that the partial 18S rRNA gene sequences may be too conserved for phylogenetic inference among Heterorhabditis species, but are well suited for phylogenetic inference within and among closely related families and genera of entomopathogenic nematodes and for inferring phylogenetic relationships among Steinernema species.

Animals↗

[The confirmation of putative natural hybrid species Meconopsis x cookei G. Taylor (Papaveraceae) based on nuclear ribosomal DNA ITS region sequence].

The Nuclear Ribosomal DNA internal transcribed spacers (ITS) region sequences from a putative natural hybrid species Meconopsis x cookei and its possible parents M. punicea and M. quintuplinervia were obtained by using direct sequencing method. The sequence length of ITS region (including ITS1, 5.8S and ITS2) is 667 bp for M. punicea, 668 bp for M. x cookei, and 668 bp for M. quintuplinervia. The sequences were aligned by the software Clustal X, and the bases per locus were compared by using software with manual method. The aligned sequence length is 688 bp, of which ITS1 is 254 bp, 5.8S is 162 bp, and ITS2 is 252 bp. 16 variable loci were detected from the aligned sequence, approximately 2.40% to the whole sequence length, of which ITS1 has nine variable loci (56.25%), ITS2 has six variable loci (37.50%), and 5.8S has one variable locus (6.25%). The results show that M. x cookei have two kinds of ITS sequences contributed from two species M. puniceaa and M. quintuplinervia, i.e. the variation of ITS gene among M. x cookei, M. punicea and M. quintuplinervia is congruence with the Mendel's genetics law. Therefore, the molecular evidences indicate that M. x cookei is a hybrid origin from M. punicea and M. quintuplinervia.

Base Sequence↗

An improved algorithm for statistical alignment of sequences related by a star tree.

The insertion-deletion model developed by Thorne, Kishino and Felsenstein (1991, J. Mol. Evol., 33, 114-124; the TKF91 model) provides a statistical framework of two sequences. The statistical alignment of a set of sequences related by a star tree is a generalization of this model. The known algorithm computes the probability of a set of such sequences in O(l2k) time, where l is the geometric mean of the sequence lengths and k is the number of sequences. An improved algorithm is presented whose running time is only O(2(2k)lk).

Algorithms↗

Using evolutionary trees in protein secondary structure prediction and other comparative sequence analyses.

Previously proposed methods for protein secondary structure prediction from multiple sequence alignments do not efficiently extract the evolutionary information that these alignments contain. The predictions of these methods are less accurate than they could be, because of their failure to consider explicitly the phylogenetic tree that relates aligned protein sequences. As an alternative, we present a hidden Markov model approach to secondary structure prediction that more fully uses the evolutionary information contained in protein sequence alignments. A representative example is presented, and three experiments are performed that illustrate how the appropriate representation of evolutionary relatedness can improve inferences. We explain why similar improvement can be expected in other secondary structure prediction methods and indeed any comparative sequence analysis method.

Amino Acid Sequence↗

Structure-derived substitution matrices for alignment of distantly related sequences.

Sequence alignment is a standard method to infer evolutionary, structural, and functional relationships among sequences. The quality of alignments depends on the substitution matrix used. Here we derive matrices based on superimpositions from protein pairs of similar structure, but of low or no sequence similarity. In a performance test the matrices are compared with 12 other previously published matrices. It is found that the structure-derived matrices are applicable for comparisons of distantly related sequences. We investigate the influence of evolutionary relationships of protein pairs on the alignment accuracy.

Models, Molecular↗

Analysis of sequence variation among legume lectins. A ring of hypervariable residues forms the perimeter of the carbohydrate-binding site.

Twelve plant lectins from the Papilionoideae subfamily were selected to represent a range of carbohydrate specificities, and their sequences were aligned. Two variability indices were applied to the aligned sequences and the results were analysed using the three-dimensional structures of concanavalin A and the pea lectin. The areas of greatest variability were located in the carbohydrate-binding site region, forming a perimeter around a well-conserved core. These residues are inferred to be specificity determining, in the manner of antibodies, and the most variable position corresponded to Tyr100 in concanavalin A, a known ligand contact residue. In addition to the five peptide loops known to form the binding site from crystallographic studies, a sixth segment with variable residues was located in the binding-site region, and this may contribute to oligosaccharide specificity. In their overall composition, the lectin sites resemble those of the sugar-transport proteins rather than antibodies. The prospects for modelling lectin binding sites by the methods used for antibodies were also assessed.

Amino Acid Sequence↗

ENVIRON: a software package to compare protein three-dimensional structures with homologous sequences using local structural motifs.

This work presents a method to compare local clusters of interacting residues as observed in a known three-dimensional protein structure with corresponding clusters inferred from homologous protein sequences, assuming conserved protein folding. For this purpose the local environment of a selected residue in a known protein structure is defined as the ensemble of amino acids in contact with it in the folded state. Using a multiple sequence alignment to identify corresponding residues in homologous proteins, a detailed comparison can be performed between the local environment of a selected amino acid in the template protein structure and the expected local environments at the sets of equivalent residues, derived from the aligned protein sequences. The comparison makes it possible to detect conserved local features such as hydrogen bonding or complementarity in residue substitution. A global measure of environmental similarity is also defined, to search for conserved amino acid clusters subject to functional or structural constraints. The proposed approach is useful for investigating protein function as well as for site-directed mutagenesis experiments, where appropriate amino acid substitutions can be suggested by observing naturally occurring protein variants.

Algorithms↗

Rapid motif-based prediction of circular permutations in multi-domain proteins.

MOTIVATION: Rearrangements of protein domains and motifs such as swaps and circular permutations (CPs) can produce erroneous results in searching sequence databases when using traditional methods based on linear sequence alignments. Circular permutations are also of biological relevance because they can help to better understand both protein evolution and functionality. RESULTS: We have developed an algorithm, RASPODOM, which is based on the classical recursive alignment scheme. Sequences are represented as strings of domains taken from precompiled resources of domain (motif) databases such as ProDom. The algorithm works several orders of magnitude faster than a reimplementation of the existing CP detection algorithm working on strings of amino acids, produces virtually no false positives and allows the discrimination of true CPs from 'intermediate' CPs (iCPs). Several true CPs which have not been reported in literature so far could be identified from Swiss-Prot/TrEMBL within minutes.

Algorithms↗

Comparative sequence analysis of a gene-rich cluster at human chromosome 12p13 and its syntenic region in mouse chromosome 6.

The Human Genome Project has created a formidable challenge: the extraction of biological information from extensive amounts of raw sequence. With the increasing availability of genomic sequence from other species, one approach to extracting coding and regulatory element information is through cross-species sequence comparison. To assess the strengths and weaknesses of this methodology for large-scale sequence analysis, 227 kb of mouse sequence syntenic to a gene-rich cluster on human chromosome 12p13 was obtained. Primarily through percent identity plots (PIPs) of SIM comparative sequence alignments, the sequence of coding regions, putative alternative exons, conserved noncoding regions, and correlation in repetitive element insertions were easily determined. The analysis demonstrated that the number, order, and orientation of all 17 genes are conserved between the two species, whereas two human pseudogenes are absent in mouse. In addition, apart from MIRs, no direct correlation of distribution or position of the majority of repetitive elements between the two species is seen. Finally, in examining the synonymous and nonsynonymous substitution rates in the conserved genes, a large variation in nonsynonymous rates is observed indicating that the genes in this region are diverging at different rates. This study indicates the utility and strength of large-scale cross-species sequence comparisons in the extraction of biological information from raw sequence, especially when combined with other computational tools such as GRAIL and BLAST.

Amino Acid Sequence↗

A fast algorithm for determining the best combination of local alignments to a query sequence.

BACKGROUND: Existing sequence alignment algorithms assume that similarities between DNA or amino acid sequences are linearly ordered. That is, stretches of similar nucleotides or amino acids are in the same order in both sequences. Recombination perturbs this order. An algorithm that can reconstruct sequence similarity despite rearrangement would be helpful for reconstructing the evolutionary history of recombined sequences. RESULTS: We propose a graph-based algorithm for combining multiple local alignments to a query sequence into the single combination of alignments that either covers the maximal portion of the query or results in the single highest alignment score to the query. This algorithm can help study the process of genome rearrangement, improve functional gene annotation, and reconstruct the evolutionary history of recombined proteins. The algorithm takes O(n2) time, where n is the number of local alignments considered. CONCLUSIONS: We discuss two example applications of the algorithm. The algorithm is able to provide useful reconstructions of the metazoan mitochondrial genome. It is also able to increase the percentage of a query sequence's amino acid residues for which similar stretches of amino acids can be found in sequence databases.

Algorithms↗