PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A measure of the similarity of sets of sequences not requiring sequence alignment.

Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classifications. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism--for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.

Animals

Structure-based sequence alignment of three AdoMet-dependent DNA methyltransferases.

M.HhaI, M.TaqI and COMT are DNA methyltransferases (MTases) which catalyze the transfer of a methyl group from the cofactor AdoMet to C5 of cytosine, to N6 of adenine and to a hydroxyl group of catechol, respectively. The larger catalytic domains of the bilobal proteins, M.HhaI and M.TaqI, and the entire single domain of COMT have an alpha/beta structure containing a mixed central beta-sheet. These domains have very similar folding. By allowing appropriate 'insertions' or 'deletions' in the backbones of the three structures, it was possible to find more conserved motifs in M.TaqI and COMT. The similarity in protein folding and the equivalence of amino-acid sequences revealed by the structural alignment indicate that many AdoMet-dependent MTases may share a common catalytic domain structure.

Amino Acid Sequence

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence

Maximum entropy weighting of aligned sequences of proteins or DNA.

In a family of proteins or other biological sequences like DNA the various subfamilies are often very unevenly represented. For this reason a scheme for assigning weights to each sequence can greatly improve performance at tasks such as database searching with profiles or other consensus models based on multiple alignments. A new weighting scheme for this type of database search is proposed. In a statistical description of the searching problem it is derived from the maximum entropy principle. It can be proved that, in a certain sense, it corrects for uneven representation. It is shown that finding the maximum entropy weights is an easy optimization problem for which standard techniques are applicable.

Amino Acid Sequence

A flexible multiple sequence alignment program.

The 'regions' method for multisequence alignment used in the previously reported program MALIGN has been generalized to include recursive refinement so that unaligned portions between two regions at the current level of resolution can be handled with increased resolution. Additionally, there is incorporated a limiting of the number of regions to be used at any level of resolution from which to abstract an alignment. This provides a significant increase in speed over the unlimited version. The program GENALIGN uses this improved regions method to execute fast pairwise alignments in the framework of Taylor's multisequence alignment procedure using clustered pairwise alignments. Pairwise alignments by dynamic programming are also provided in the program.

Algorithms

Sequence alignment and evolutionary comparison of the L10 equivalent and L12 equivalent ribosomal proteins from archaebacteria, eubacteria, and eucaryotes.

The genes corresponding to the L10 and L12 equivalent ribosomal proteins (L10e and L12e) of Escherichia coli have been cloned and sequenced from two widely divergent species of archaebacteria, Halobacterium cutirubrum and Sulfolobus solfataricus. The deduced amino acid sequences of the L10e and L12e proteins have been compared to each other and to available eubacterial and eucaryotic sequences. We have identified the human P0 protein as the eucaryotic L10e. The L10e proteins from the three kingdoms were found to be colinear. The eubacterial L10e protein is much shorter than the archaebacterial-eucaryotic proteins because of two large deletions, one internal and one at the carboxy terminus. The archaebacterial and eucaryotic L12e proteins were also colinear; the eubacterial protein is homologous to the archaebacterial and eucaryotic L12e proteins, but has suffered rearrangement through what appear to be gene fusion events. Intraspecies comparisons between L10e and L12e sequences indicate the archaebacterial and eucaryotic L10e proteins contain a partial copy of the L12e protein fused to their carboxy terminus. In the eubacteria most of this fusion has been removed by the carboxy terminal deletion. Within the L12e-derived region, a 26-amino acid-long internal modular sequence reiterated thrice in the archaebacterial L10e, twice in the eucaryotic L10e, and once in the eubacterial L10e was discovered. This modular sequence also appears to be present as a single copy in all L12e proteins and may play a role in L12e dimerization, L10e-L12e complex formation, and the function of L10e-L12e complex in translation.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

The bacterial porin superfamily: sequence alignment and structure prediction.

The porins of Gram-negative bacteria are responsible for the 'molecular sieve' properties of the outer membrane. They form large water-filled channels which allow the diffusion of hydrophilic molecules into the periplasmic space. Owing to the strong hydrophilicity of their amino acid sequence and the nature of their secondary structure (beta strands), conventional hydropathy methods for predicting membrane topology are useless for this class of protein. The large number of available porin amino acid sequences was exploited to improve the accuracy of the prediction in combination with tools detecting amphipathicity of secondary structure. Using the constraints of beta-sheet structure these porins are predicted to contain 16 membrane-spanning strands, 14 of which are common to the two (enteric and the neisserial) porin subfamilies.

Amino Acid Sequence

Prediction of an inter-residue interaction in the chaperonin GroEL from multiple sequence alignment is confirmed by double-mutant cycle analysis.

A search for co-ordinated amino acid changes in the hsp60 family of chaperonins suggested that cysteine residues at positions 137 and 518 in the Escherichia coli chaperonin GroEL may interact with each other. In order to determine whether this interaction indeed exists we constructed a double-mutant cycle comprising wild-type GroEL, the single mutants Cys137-->Ser and Cys518-->Ser and the corresponding double mutant. The effects of the two mutations on the function of GroEL, in assisting the refolding of a non-folded protein substrate (rhodanese), are shown to be non-additive. It is also shown that ADP by itself specifically destabilizes the Cys518-->Ser mutant GroEL particle with this effect being suppressed in the double mutant. The observed pattern of co-ordinated mutations in the hsp60 family of chaperonins is thus shown to reflect a real interaction, though most likely indirect, between Cys137 and Cys518 in GroEL. Our study demonstrates that patterns of co-ordinated mutations combined with double-mutant cycle analysis can provide structural information on interactions in a protein without an available three-dimensional structure at atomic resolution.

Bacterial Proteins

ALIGNMENT SERVICE: creation and processing of alignments of sequences of unlimited length.

A package for the creation and processing of multiple sequence alignment is described. There is no limit on the lengths of the processed nucleotide or amino acid sequences, and the number of sequences in the alignment is also unlimited. The main groups of functions are: a semiautomatic alignment editor; a wide set of functions for technical processing of alignments; nucleotide alignment mapping and translation; and similarity search functions. A user-friendly interface and a set of generally used file actions provide a special operational subsystem for everyday tasks.

Amino Acid Sequence

An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.

The sensitivity of most protein sequence alignment methods depends strongly on the quality of the comparison matrices used. These matrices, which assign weights or similarity scores to every possible amino acid substitution pair, are utilized to differentiate amongst the various possible alignments of two or more sequences. There are many ways to generate these exchange weights and new matrices are constantly published. There has been no overall assessment of these various matrices when applied in different alignment techniques and over many protein folds and families, both close and distant and with the use of several gap penalty values. In this work, a set of amino acid sequences matched by superposition of known protein tertiary topologies is used to test the alignment accuracy of the different method/matrix/penalty combinations. The comparisons show relatively similar results for the top scoring matrices, a preference for the global alignment method of Needleman and Wunsch, and the importance of matrix modification and optimized gap penalties. The relationship between the percentage identity in a resulting alignment and the level of correctness to be expected are given for the top-performing matrix, resulting in a better definition of the so-called "twilight zone". Estimates are made for the probability that two sequences, aligned at a certain level of residue percentage identity, are in fact unrelated.

Amino Acid Sequence

Sequence divergence analysis for the prediction of seven-helix membrane protein structures: II. A 3-D model of human rhodopsin.

A three-dimensional (3-D) model of the transmembrane domain of human rhodopsin was predicted from the sequence divergence analysis of 42 sequences of rhodopsins and visual pigments without a template. The prediction steps include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, their secondary structure and packing shape in a helix bundle, prediction of side-chain conformations and structure refinement. The identification of the retinal binding site was assisted by its known covalent linkage with K296. The structural features of the predicted 3-D model are in good agreement with a low resolution electron density map of bovine rhodopsin and with residues in contact with retinal as determined experimentally.

Amino Acid Sequence

Genetic relatedness of human DNA polymerase beta and terminal deoxynucleotidyltransferase.

The Protein Identification Resource (PIR) protein sequence data bank was searched for sequence similarity between known proteins and human DNA polymerase beta (Pol beta) or human terminal deoxynucleotidyltransferase (TdT). Pol beta and TdT were found to exhibit amino acid sequence similarity only with each other and not with any other of the 4750 entries in release 12.0 of the PIR data bank. Optimal amino acid sequence alignment of the entire 39-kDa Pol beta polypeptide with the C-terminal two thirds of TdT revealed 24% identical aa residues and 21% conservative aa substitutions. The Monte Carlo score of 12.6 for the entire aligned sequences indicates highly significant aa sequence homology. The hydropathicity profiles of the aligned aa sequences were remarkably similar throughout, suggesting structural similarity of the polypeptides. The most significant regions of homology are aa residues 39-224 and 311-333 of Pol beta vs. aa residues 191-374 and 484-506 of TdT. In addition, weaker homology was seen between a large portion of the 'nonessential' N-terminal end of TdT (aa residues 33-130) and the first region of strong homology between the two proteins (aa residues 31-128 of Pol beta and aa residues 183-280 of TdT), suggestive of genetic duplication within the ancestral gene. On the basis of nucleotide differences between conserved regions of Pol beta and TdT genes (aligned according to optimally aligned aa sequences) it was estimated that Pol beta and TdT diverged on the order of 250 million years ago, corresponding roughly to a time before radiation of mammals and birds.

Amino Acid Sequence

Alignment of molecular sequences seen as random path analysis.

We propose a generating functional method--random path analysis (RPA)--that generalizes the classical dynamic programming (DP) method widely used in sequence alignments. For a given cost function, DP is a deterministic method that finds an optimal alignment by minimizing the total cost function for all possible alignments. By allowing uncertainty, RPA is a statistical method that weights fluctuating alignments by probabilities. Therefore, DP maybe thought of as the deterministic limit of RPA when the fluctuations approach zero. DP is the method of choice if one is only interested in optimal alignment. But we argue that, when information beyond the optimal alignment is desired, RPA gives a natural extension of DP for biological applications. As an algebraic approach, RPA is computationally intensive for long sequences, but it can provide better parametric control for developing analytical or perturbational results and it is more informative and biologically relevant. The idea of RPA opens up new opportunities for simulational approaches and more importantly it suggests a novel hardware implementation that has the potential of improving the way a sequence alignment is done. Here we focus on deriving a mathematically rigorous solution to RPA both in its combinatorial form and in its graphical representation; this puts DP in logical perspective under a more general conceptual framework.

Animals