PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Analysis of the nucleotide and derived amino acid sequences of the SsoII restriction endonuclease and methyltransferase.

A 2648-bp fragment from the P4 plasmid of Shigella sonnei strain 47 coding for the SsoII restriction endonuclease (ENase) and methyltransferase (MTase) (recognition sequence 5'-CCNGG) was sequenced. Two divergently arranged open reading frames of 905 bp for the SsoII ENase (R.SsoII) and 1137 bp for the MTase (M.SsoII) were identified. The coding regions are separated by 110 bp. The calculated M(r) of R.SsoII (35937) and M.SsoII (42887) are in good agreement with values previously obtained by in vitro transcription-translation experiments, i.e., 35 and 43 kDa for the ENase and MTase, respectively. The M.SsoII amino acid (aa) sequence revealed a considerable similarity to m5C-MTases recognizing the related sequences--M.EcoRII, M.dcm, M.MspI, M.BsuFI, M.HpaII, and M.HhaI. Surprisingly, the greatest degree of homology has been observed between the aa sequences of M.SsoII and M.NlaX, with an unidentified recognition sequence. The multiple alignment of aa sequences helps to identify the blocks of conserved aa in variable regions of MTases. These conserved aa can play a key role in target recognition. Some aspects of evolution of m5C-MTases are discussed.

Amino Acid Sequence

Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.

A wealth of protein and DNA sequence data is being generated by genome projects and other sequencing efforts. A crucial barrier to deciphering these sequences and understanding the relations among them is the difficulty of detecting subtle local residue patterns common to multiple sequences. Such patterns frequently reflect similar molecular structures and biological properties. A mathematical definition of this "local multiple alignment" problem suitable for full computer automation has been used to develop a new and sensitive algorithm, based on the statistical method of iterative sampling. This algorithm finds an optimized local alignment model for N sequences in N-linear time, requiring only seconds on current workstations, and allows the simultaneous detection and optimization of multiple patterns and pattern repeats. The method is illustrated as applied to helix-turn-helix proteins, lipocalins, and prenyltransferases.

Algorithms

A method to recognize distant repeats in protein sequences.

An automated algorithm is presented that delineates protein sequence fragments which display similarity. The method incorporates a selection of a number of local nonoverlapping sequence alignments with the highest similarity scores and a graph-theoretical approach to elucidate the consistent start and end points of the fragments comprising one or more ensembles of related subsequences. The procedure allows the simultaneous identification of different types of repeats within one sequence. A multiple alignment of the resulting fragments is performed and a consensus sequence derived from the ensemble(s). Finally, a profile is constructed from the multiple alignment to detect possible and more distant members within the sequence. The method tolerates mutations in the repeats as well as insertions and deletions. The sequence spans between the various repeats or repeat clusters may be of different lengths. The technique has been applied to a number of proteins where the repeating fragments have been derived from information additional to the protein sequences.

Algorithms

Structure-based multiple alignment of extracellular pectate lyase sequences.

Pectate lyases are secreted virulence factors which degrade the pectate component of plant cell walls. The evolutionary-based multiple alignment of extracellular pectate lyases has been corrected using three-dimensional structural information derived from Erwinia chyrsanthemi pectate lyases C and E. The new multiple alignment reveals invariant amino acids likely to be involved in two different enzymatic functions.

Amino Acid Sequence

The orotidine-5'-monophosphate decarboxylase gene of Myxococcus xanthus. Comparison to the OMP decarboxylase gene family.

The nucleotide sequence of the Myxococcus xanthus orotidine-5'-monophosphate decarboxylase (OMP DCase) gene was determined. The derived protein sequence is not closely related to other prokaryotic OMP DCase sequences; nor is it closely related to any eukaryotic OMP DCase sequences. Progressive multiple alignment of the M. xanthus OMP DCase protein sequence with 19 other OMP DCase sequences revealed four conserved regions present in all 20 sequences. Ten entirely conserved residues were found in these four regions and one region contains a tight cluster of 5 conserved residues, certain of which may be catalytically active residues. A second open reading frame was found upstream of uraA and oriented in the same direction as uraA. A stretch of 21 consecutive pyrimidine (C or T) residues were found in the intercistronic region between the potential ribosome-binding site of uraA and the UGA stop codon of the upstream open reading frame. RNA directly upstream of the pyrimidine run, including the UGA stop codon of the upstream open reading frame, could be folded into a stable hairpin structure resembling Rho-independent terminators of Escherichia coli. Expression of the uraA gene may be regulated by an intercistronic transcription termination mechanism.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Phylogenetic relationships among megabats, microbats, and primates.

We present 744 nucleotide base positions from the mitochondrial 12S rRNA gene and 236 base positions from the mitochondrial cytochrome oxidase subunit I gene for a microbat, Brachyphylla cavernarum, and a megabat, Pteropus capestratus, in phylogenetic analyses with homologous DNA sequences from Homo sapiens, Mus musculus (house mouse), and Gallus gallus (chicken). We use information on evolutionary rate differences for different types of sequence change to establish phylogenetic character weights, and we consider alternative rRNA alignment strategies in finding that this mtDNA data set clearly supports bat monophyly. This result is found despite variations in outgroup used, gap coding scheme, and order of input for DNA sequences in multiple alignment bouts. These findings are congruent with morphological characters including details of wing structure as well as cladistic analyses of amino acid sequences for three globin genes and indicate that neurological similarities between megabats and primates are due to either retention of primitive characters or to convergent evolution rather than to inheritance from a common ancestor. This finding also indicates a single origin for flight among mammals.

Amino Acid Sequence

Flexible algorithm for direct multiple alignment of protein structures and sequences.

The recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach.

Algorithms

Calculating percent identity between protein or DNA sequences with a word processor.

Two macros, to calculate percentage identity between protein or DNA sequences using the Microsoft Word word processor, are described. The user prepares an alignment file of multiple sequences which is used by the macros to calculate number of matches, number of mismatches, total number of compared positions, and the percent identity. The macros are especially useful when alignment of multiple sequences is possible only by eye.

Algorithms

Multiple alignment and hierarchical clustering of conserved amino acid sequences in the replication-associated proteins of plant RNA viruses.

We have used multiple alignment computer programs to align and hierarchically cluster the conserved amino acid "signature" sequences found in the replication-associated proteins of all plant RNA viruses sequenced so far. These regions, called "polymerase", "nucleotide-binding" and "N-terminal" are well conserved even between viruses which are only distantly related, and are thus very well suited for this type of analysis. Our results show that the clusterings obtained using these very short amino acid sequences are very robust to computing parameters and are surprisingly well matched with the taxonomic grouping of RNA plant viruses. The possibility of using this system as a new taxonomic criterion is discussed.

Amino Acid Sequence

Evolutionary relationship between the TonB-dependent outer membrane transport proteins: nucleotide and amino acid sequences of the Escherichia coli colicin I receptor gene.

The nucleotide sequence of the Escherichia coli colicin I receptor gene (cir) has been determined. The predicted mature protein consists of 599 amino acids and has a molecular weight of 67,169. Several previously noted characteristics of other E. coli outer membrane protein sequences were also identified in the sequence of Cir. These include an overall acidic nature, the absence of long hydrophobic stretches of amino acids, and a lack of predicted alpha-helical secondary structure. Because two classes of outer membrane proteins (the TonB-dependent transport proteins and the porins) share some structural features, protein sequences from both of these groups were aligned pairwise and scored for sequence similarity. Statistical evidence suggested that the porins were not related to the proteins in the TonB-dependent group; however, there was a significant relationship between the proteins in the TonB-dependent group. On the basis of the multiple progressive sequence alignment and the similarity scores derived from it, a tree representing evolutionary distance between five TonB-dependent outer membrane transport proteins was generated.

Amino Acid Sequence

A multiple alignment of the capsid protein sequences of nepoviruses and comoviruses suggests a common structure.

The amino acid sequences of the regions encoding the structural proteins of eleven nepoviruses and five comoviruses, two genera of the family Comoviridae, have been aligned. The properties predicted by computer analysis (three-dimensional-3D-structure, hydrophobicity) are also correlated along this alignment, and aligned to the experimentally determined 3D structure of two comoviruses. It can thus be assumed that the 3D structure of the unique nepovirus coat protein matches that of the bipartite protomer found in the comovirus particles. In this model, the spatial locations of two amino-acid motifs characteristic of nepoviruses are in close vicinity, at the external surface of the virion. The coat proteins of nepoviruses and comoviruses may thus share a common evolutionary origin. A phylogenetic analysis was made using the multiple alignment, allowing a better understanding of the molecular relationships between these two groups of viruses.

Amino Acid Sequence

The inference of evolutionary trees from molecular data.

1. Procedures for multiple alignment of sequence data, subsequent phylogenetic inference, and testing of the trees derived are presented. 2. The assumptions underlying different approaches and the extent to which they are valid are discussed.

Amino Acid Sequence

Evidence for an evolutionary relationship among type-II restriction endonucleases.

Type-II restriction-modification (R-M) systems comprise two enzymes, a DNA methyltransferase (MTase) and a restriction endonuclease (ENase), each of which specifically interact with the same 4-8 bp sequence. All type-II MTases share several amino acid (aa) sequence motifs, which makes an evolutionary relatedness among these enzymes probable. The type-II ENases, in contrast, except for some homologous isoschizomers, do not share significant aa sequence similarity. Therefore, ENases in general have been considered unrelated. Here we show that in addition to the analysis of the genotype (aa sequence), a comparison of the phenotype (recognition sequence) of these enzymes can provide independent information regarding evolutionary relationships, and thereby, help to analyze the significance of weak aa sequence similarities. Multistep Monte-Carlo analyses were employed to demonstrate that the recognition sequences of those ENases, which were found to be related by a progressive multiple aa sequence alignment, are more similar to each other than would be expected by chance. This analysis supports the notion that not only type-II MTases, but also type-II ENases did not arise independently in evolution, but rather evolved from one or a few primordial DNA-modifying and DNA-cleaving enzymes, respectively.

Amino Acid Sequence

A novel RNA-binding motif in omnipotent suppressors of translation termination, ribosomal proteins and a ribosome modification enzyme?

Using computer methods for database search, multiple alignment, protein sequence motif analysis and secondary structure prediction, a putative new RNA-binding motif was identified. The novel motif is conserved in yeast omnipotent translation termination suppressor SUP1, the related DOM34 protein and its pseudogene homologue; three groups of eukaryotic and archaeal ribosomal proteins, namely L30e, L7Ae/S6e and S12e; an uncharacterized Bacillus subtilis protein related to the L7A/S6e group; and Escherichia coli ribosomal protein modification enzyme RimK. We hypothesize that a new type of RNA-binding domain may be utilized to deliver additional activities to the ribosome.

Amino Acid Sequence