PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

The nucleoside-specific Tsx channel from the outer membrane of Salmonella typhimurium, Klebsiella pneumoniae and Enterobacter aerogenes: functional characterization and DNA sequence analysis of the tsx genes.

The Escherichia coli tsx gene encodes an integral outer-membrane protein (Tsx) that functions as a substrate-specific channel for deoxynucleosides and the antibiotic albicidin, and also serves as a receptor for bacteriophages and colicins. We cloned the structural genes of the Tsx proteins from Salmonella typhimurium, Klebsiella pneumoniae and Enterobacter aerogenes and expressed them in an E.coli tsx mutant. The heterologous Tsx proteins fully substituted the E.coli Tsx protein with respect to its function in deoxynucleoside and albicidin uptake, and as receptor for colicin K. The Tsx proteins from K. pneumoniae and Ent. aerogenes were also proficient as receptors for several Tsx-specific bacteriophages, whereas the corresponding protein from S. typhimurium did not confer sensitivity against these phages. The nucleotide sequence of the tsx genes from S. typhimurium, K. pneumoniae and Ent. aerogenes was established. Each of the Tsx proteins is initially synthesized with typical bacterial signal sequence peptides and the predicted mature forms of the Tsx proteins have a calculated M(r) of 30,567 (265 residues), 31,412 (272 residues) and 31,477 (272 residues), respectively. Multiple sequence alignments between the Tsx proteins showed a high degree of sequence identity and revealed the presence of four hypervariable regions, which are thought to constitute segments of the polypeptide chain exposed at the cell surface. Most notable was a deletion of 8 amino acids in one of these hypervariable domains in the S. typhimurium Tsx protein. When this deletion was introduced by site-directed mutagenesis into the corresponding region of the E.coli tsx gene, the mutant Tsx-515 protein lost its phage receptor function but still served as a colicin K receptor and as a substrate-specific channel, indicating that the region between residues 198 and 207 might be part of the bacteriophage receptor area. Multiple sequence alignments, structural predictions and the properties of previously characterized Tsx missense mutants were taken into account to develop a two-dimensional model for the topological organization of the Tsx protein within the outer membrane.

Amino Acid Sequence↗

TOPAL 2.0: improved detection of mosaic sequences within multiple alignments.

MOTIVATION: The Dss statistic was proposed by McGuire et al. (Mol. Biol. Evol., 14, 1125-1131, 1997) for scanning data sets for the presence of recombination, an important step in some phylogenetic analyses. The statistic, however, could not distinguish well between among-site rate variation and recombination, and had no statistical test for significant values. This paper addresses these shortfalls. RESULTS: A modification to the Dss statistic is proposed which accounts for rate variation to a large extent. A statistical test, based on parametric bootstrapping, is also suggested. AVAILABILITY: The TOPAL package (version 2) may be accessed from http:/ /www.bioss.sari.ac.uk/frank/Genetics and by anonymous ftp from typ://ftp.bioss.sari.ac.uk in the directory pub/phylogeny/topal. CONTACT: frank@bioss.sari.ac.uk

Algorithms↗

Recurring local sequence motifs in proteins.

We describe a completely automated approach to identifying local sequence motifs that transcend protein family boundaries. Cluster analysis is used to identify recurring patterns of variation at single positions and in short segments of contiguous positions in multiple sequence alignments for a non-redundant set of protein families. Parallel experiments on simulated data sets constructed with the overall residue frequencies of proteins but not the inter-residue correlations show that naturally occurring protein sequences are significantly more clustered than the corresponding random sequences for window lengths ranging from one to 13 contiguous positions. The patterns of variation at single positions are not in general surprising: chemically similar amino acids tend to be grouped together. More interesting patterns emerge as the window length increases. The patterns of variation for longer window lengths are in part recognizable patterns of hydrophobic and hydrophilic residues, and in part less obvious combinations. A particularly interesting class of patterns features highly conserved glycine residues. The patterns provide a means to abstract the information contained in multiple sequence alignments and may be useful for comparison of distantly related sequences or sequence families and for protein structure prediction.

Cluster Analysis↗

Theory and practice of parallel direct optimization.

Our ability to collect and distribute genomic and other biological data is growing at a staggering rate (Pagel, 1999). However, the synthesis of these data into knowledge of evolution is incomplete. Phylogenetic systematics provides a unifying intellectual approach to understanding evolution but presents formidable computational challenges. A fundamental goal of systematics, the generation of evolutionary trees, is typically approached as two distinct NP-complete problems: multiple sequence alignment and phylogenetic tree search. The number of cells in a multiple alignment matrix are exponentially related to sequence length. In addition, the number of evolutionary trees expands combinatorially with respect to the number of organisms or sequences to be examined. Biologically interesting datasets are currently comprised of hundreds of taxa and thousands of nucleotides and morphological characters. This standard will continue to grow with the advent of highly automated sequencing and development of character databases. Three areas of innovation are changing how evolutionary computation can be addressed: (1) novel concepts for determination of sequence homology, (2) heuristics and shortcuts in tree-search algorithms, and (3) parallel computing. In this paper and the online software documentation we describe the basic usage of parallel direct optimization as implemented in the software POY (ftp://ftp.amnh.org/pub/molecular/poy).

Animals↗

A hidden Markov model for progressive multiple alignment.

MOTIVATION: Progressive algorithms are widely used heuristics for the production of alignments among multiple nucleic-acid or protein sequences. Probabilistic approaches providing measures of global and/or local reliability of individual solutions would constitute valuable developments. RESULTS: We present here a new method for multiple sequence alignment that combines an HMM approach, a progressive alignment algorithm, and a probabilistic evolution model describing the character substitution process. Our method works by iterating pairwise alignments according to a guide tree and defining each ancestral sequence from the pairwise alignment of its child nodes, thus, progressively constructing a multiple alignment. Our method allows for the computation of each column minimum posterior probability and we show that this value correlates with the correctness of the result, hence, providing an efficient mean by which unreliably aligned columns can be filtered out from a multiple alignment.

Algorithms↗

Ab initio folding of proteins using restraints derived from evolutionary information.

We present our predictions in the ab initio structure prediction category of CASP3. Eleven targets were folded, using a method based on a Monte Carlo search driven by secondary and tertiary restraints derived from multiple sequence alignments. Our results can be qualitatively summarized as follows: The global fold can be considered "correct" for targets 65 and 74, "almost correct" for targets 64, 75, and 77, "half-correct" for target 79, and "wrong" for targets 52, 56, 59, and 63. Target 72 has not yet been solved experimentally. On average, for small helical and alpha/beta proteins (on the order of 110 residues or smaller), the method predicted low resolution structures with a reasonably good prediction of the global topology. Most encouraging is that in some situations, such as with target 75 and, particularly, target 77, the method can predict a substantial portion of a rare or even a novel fold. However, the current method still fails on some beta proteins, proteins over the 110-residue threshold, and sequences in which only a poor multiple sequence alignment can be built. On the other hand, for small proteins, the method gives results of quality at least similar to that of threading, with the advantage of not being restricted to known folds in the protein database. Overall, these results indicate that some progress has been made on the ab initio protein folding problem. Detailed information about our results can be obtained by connecting to http:/(/)www.bioinformatics.danforthcenter.org/+ ++CASP3.

Algorithms↗

A neural network method for prediction of beta-turn types in proteins using evolutionary information.

MOTIVATION: The prediction of beta-turns is an important element of protein secondary structure prediction. Recently, a highly accurate neural network based method Betatpred2 has been developed for predicting beta-turns in proteins using position-specific scoring matrices (PSSM) generated by PSI-BLAST and secondary structure information predicted by PSIPRED. However, the major limitation of Betatpred2 is that it predicts only beta-turn and non-beta-turn residues and does not provide any information of different beta-turn types. Thus, there is a need to predict beta-turn types using an approach based on multiple sequence alignment, which will be useful in overall tertiary structure prediction. RESULTS: In the present work, a method has been developed for the prediction of beta-turn types I, II, IV and VIII. For each turn type, two consecutive feed-forward back-propagation networks with a single hidden layer have been used where the first sequence-to-structure network has been trained on single sequences as well as on PSI-BLAST PSSM. The output from the first network along with PSIPRED predicted secondary structure has been used as input for the second-level structure-to-structure network. The networks have been trained and tested on a non-homologous dataset of 426 proteins chains by 7-fold cross-validation. It has been observed that the prediction performance for each turn type is improved significantly by using multiple sequence alignment. The performance has been further improved by using a second level structure-to-structure network and PSIPRED predicted secondary structure information. It has been observed that Type I and II beta-turns have better prediction performance than Type IV and VIII beta-turns. The final network yields an overall accuracy of 74.5, 93.5, 67.9 and 96.5% with MCC values of 0.29, 0.29, 0.23 and 0.02 for Type I, II, IV and VIII beta-turns, respectively, and is better than random prediction. AVAILABILITY: A web server for prediction of beta-turn types I, II, IV and VIII based on above approach is available at http://www.imtech.res.in/raghava/betaturns/ and http://bioinformatics.uams.edu/mirror/betaturns/ (mirror site).

Algorithms↗

Exploring the relationship between sequence similarity and accurate phylogenetic trees.

We have characterized the relationship between accurate phylogenetic reconstruction and sequence similarity, testing whether high levels of sequence similarity can consistently produce accurate evolutionary trees. We generated protein families with known phylogenies using a modified version of the PAML/EVOLVER program that produces insertions and deletions as well as substitutions. Protein families were evolved over a range of 100-400 point accepted mutations; at these distances 63% of the families shared significant sequence similarity. Protein families were evolved using balanced and unbalanced trees, with ancient or recent radiations. In families sharing statistically significant similarity, about 60% of multiple sequence alignments were 95% identical to true alignments. To compare recovered topologies with true topologies, we used a score that reflects the fraction of clades that were correctly clustered. As expected, the accuracy of the phylogenies was greatest in the least divergent families. About 88% of phylogenies clustered over 80% of clades in families that shared significant sequence similarity, using Bayesian, parsimony, distance, and maximum likelihood methods. However, for protein families with short ancient branches (ancient radiation), only 30% of the most divergent (but statistically significant) families produced accurate phylogenies, and only about 70% of the second most highly conserved families, with median expectation values better than 10(-60), produced accurate trees. These values represent upper bounds on expected tree accuracy for sequences with a simple divergence history; proteins from 700 Giardia families, with a similar range of sequence similarities but considerably more gaps, produced much less accurate trees. For our simulated insertions and deletions, correct multiple sequence alignments did not perform much better than those produced by T-COFFEE, and including sequences with expressed sequence tag-like sequencing errors did not significantly decrease phylogenetic accuracy. In general, although less-divergent sequence families produce more accurate trees, the likelihood of estimating an accurate tree is most dependent on whether radiation in the family was ancient or recent. Accuracy can be improved by combining genes from the same organism when creating species trees or by selecting protein families with the best bootstrap values in comprehensive studies.

Animals↗

ViTO: tool for refinement of protein sequence-structure alignments.

UNLABELLED: ViTO is a graphical application, including an editor, of multiple sequence alignment and a three-dimensional (3D) structure viewer. It is possible to manipulate alignments containing hundreds of sequences and to display a dozen structures. ViTO can handle so-called 'multiparts' alignments to allow the visualization of complex structures (multi-chain proteins and/or small molecules and DNA) and the editing of the corresponding alignment. The 3D viewer and the alignment editor are connected together allowing rapid refinement of sequence-structure alignment by taking advantage of the immediate visualization of resulting insertions/deletions and strict conservations in their structural context. More generally, it allows the mapping of informations about the sequence conservation extracted from the alignment onto the 3D structures in a dynamic way. ViTO is also connected to two comparative modelling programs, SCWRL and MODELLER. These features make ViTO a powerful tool to characterize protein families and to optimize the alignments for comparative modelling. AVAILABILITY: http://bioserv.cbs.cnrs.fr/VITO/DOC/. SUPPLEMENTARY INFORMATION: http://bioserv.cbs.cnrs.fr/VITO/DOC/index.html.

Amino Acid Sequence↗

Functional multimerization of human telomerase requires an RNA interaction domain in the N terminus of the catalytic subunit.

Functional human telomerase complexes are minimally composed of the human telomerase RNA (hTR) and a catalytic subunit (human telomerase reverse transcriptase [hTERT]) containing reverse transcriptase (RT)-like motifs. The N terminus of TERT proteins is unique to the telomerase family and has been implicated in catalysis, telomerase RNA binding, and telomerase multimerization, and conserved motifs have been identified by alignment of TERT sequences from multiple organisms. We studied hTERT proteins containing N-terminal deletions or substitutions to identify and characterize hTERT domains mediating telomerase catalytic activity, hTR binding, and hTERT multimerization. Using multiple sequence alignment, we identified two vertebrate-conserved TERT N-terminal regions containing vertebrate-specific residues that were required for human telomerase activity. We identified two RNA interaction domains, RID1 and RID2, the latter containing a vertebrate-specific RNA binding motif. Mutations in RID2 reduced the association of hTR with hTERT by 50 to 70%. Inactive mutants defective in RID2-mediated hTR binding failed to complement an inactive hTERT mutant containing an RT motif substitution to reconstitute activity. Our results suggest that functional hTERT complementation requires intact RID2 and RT domains on the same hTERT molecule and is dependent on hTR and the N terminus.

Amino Acid Sequence↗

Protein sequence threading: Averaging over structures.

Multiple sequence alignments are a routine tool in protein fold recognition, but multiple structure alignments are computationally less cooperative. This work describes a method for protein sequence threading and sequence-to-structure alignments that uses multiple aligned structures, the aim being to improve models from protein threading calculations. Sequences are aligned into a field due to corresponding sites in homologous proteins. On the basis of a test set of more than 570 protein pairs, the procedure does improve alignment quality, although no more than averaging over sequences. For the force field tested, the benefit of structure averaging is smaller than that of adding sequence similarity terms or a contribution from secondary structure predictions. Although there is a significant improvement in the quality of sequence-to-structure alignments, this does not directly translate to an immediate improvement in fold recognition capability.

Animals↗

Functional implications of structure-based sequence alignment of proteins in the extracellular pectate lyase superfamily.

Pectate lyases are plant virulence factors that degrade the pectate component of the plant cell wall. The enzymes share considerable sequence homology with plant pollen and style proteins, suggesting a shared structural topology and possibly functional relationships as well. The three-dimensional structures of two Erwinia chrysanthemi pectate lyases, C and E, have been superimposed and the structurally conserved amino acids have been identified. There are 232 amino acids that superimpose with a root-mean-square deviation of 3 A or less. These amino acids have been used to correct the primary sequence alignment derived from evolution-based techniques. Subsequently, multiple alignment techniques have allowed the realignment of other extracellular pectate lyases as well as all sequence homologs, including pectin lyases and the plant pollen and style proteins. The new multiple sequence alignment reveals amino acids likely to participate in the parallel beta helix motif, those involved in binding Ca2+, and those invariant amino acids with potential catalytic properties. The latter amino acids cluster in two well-separated regions on the pectate lyase structures, suggesting two distinct enzymatic functions for extracellular pectate lyases and their sequence homologs.

Amino Acid Sequence↗

QAlign: quality-based multiple alignments with dynamic phylogenetic analysis.

Integrating different alignment strategies, a layout editor and tools deriving phylogenetic trees in a 'multiple alignment environment' helps to investigate and enhance results of multiple sequence alignment by hand. QAlign combines algorithms for fast progressive and accurate simultaneous multiple alignment with a versatile editor and a dynamic phylogenetic analysis in a convenient graphical user interface.

Algorithms↗

Sequence alignments, variabilities, and vagaries.

It seems as if the algorithms and weighting matrices for multiple sequence alignments of the highly divergent members of the P450 gene superfamily have advanced to the point that unknown proteins can be aligned to structurally known members with reasonable accuracy. As stated earlier, the alignment tends to break down at gaps in the sequence alignments, but these regions can be improved manually. This type of alignment and analysis is especially useful for extracting and analyzing the various genome databases. Variations of the conservation analysis can be used to identify charged and uncharged residues that may be important in domain/domain interactions with redox partners or effector molecules (e.g., cytochrome b5). From these alignments and with comparative analysis within families and across P450 families, one can readily obtain an estimation of those residues that might be involved in substrate binding, in redox partner interaction, and in the catalytic mechanism.

Amino Acid Sequence↗

CINEMA-MX: a modular multiple alignment editor.

UNLABELLED: Analyzing and visualizing multiple sequence alignments is a common task in many areas of molecular biology and bioinformatics. Many tools exist for this purpose, but are not easily customizable for specific in-house uses. Here we report the development of an editor, CINEMA-MX, that addresses these issues. CINEMA-MX is highly modular and configurable, and we present examples to illustrate its extensibility. AVAILABILITY: The program and full source code, which are available from http://www.bioinf.man.ac.uk/dbbrowser/cinema-mx, are being released under a combination of the LGPL and GPL, for Unix or Windows platforms.

Computer Graphics↗