PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

An analysis of the periodicity of conserved residues in sequence alignments of G-protein coupled receptors. Implications for the three-dimensional structure.

Twenty-three sequences from the family of G-protein coupled receptors have been aligned according to the 'historical alignment' procedure of Feng and Doolittle. Fourier transform analysis of this reveals that parts of five of the seven putative membrane-spanning regions exhibit a periodicity of conserved/nonconserved residues which is compatible with the periodicity of the alpha-helix. This would place the conserved residues on one side of the helix, which may face the inside of the proposed seven membered helical bundle.

Amino Acid Sequence

A measure of the similarity of sets of sequences not requiring sequence alignment.

Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classifications. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism--for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.

Animals

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence

A flexible multiple sequence alignment program.

The 'regions' method for multisequence alignment used in the previously reported program MALIGN has been generalized to include recursive refinement so that unaligned portions between two regions at the current level of resolution can be handled with increased resolution. Additionally, there is incorporated a limiting of the number of regions to be used at any level of resolution from which to abstract an alignment. This provides a significant increase in speed over the unlimited version. The program GENALIGN uses this improved regions method to execute fast pairwise alignments in the framework of Taylor's multisequence alignment procedure using clustered pairwise alignments. Pairwise alignments by dynamic programming are also provided in the program.

Algorithms

Sequence alignment and evolutionary comparison of the L10 equivalent and L12 equivalent ribosomal proteins from archaebacteria, eubacteria, and eucaryotes.

The genes corresponding to the L10 and L12 equivalent ribosomal proteins (L10e and L12e) of Escherichia coli have been cloned and sequenced from two widely divergent species of archaebacteria, Halobacterium cutirubrum and Sulfolobus solfataricus. The deduced amino acid sequences of the L10e and L12e proteins have been compared to each other and to available eubacterial and eucaryotic sequences. We have identified the human P0 protein as the eucaryotic L10e. The L10e proteins from the three kingdoms were found to be colinear. The eubacterial L10e protein is much shorter than the archaebacterial-eucaryotic proteins because of two large deletions, one internal and one at the carboxy terminus. The archaebacterial and eucaryotic L12e proteins were also colinear; the eubacterial protein is homologous to the archaebacterial and eucaryotic L12e proteins, but has suffered rearrangement through what appear to be gene fusion events. Intraspecies comparisons between L10e and L12e sequences indicate the archaebacterial and eucaryotic L10e proteins contain a partial copy of the L12e protein fused to their carboxy terminus. In the eubacteria most of this fusion has been removed by the carboxy terminal deletion. Within the L12e-derived region, a 26-amino acid-long internal modular sequence reiterated thrice in the archaebacterial L10e, twice in the eucaryotic L10e, and once in the eubacterial L10e was discovered. This modular sequence also appears to be present as a single copy in all L12e proteins and may play a role in L12e dimerization, L10e-L12e complex formation, and the function of L10e-L12e complex in translation.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

The bacterial porin superfamily: sequence alignment and structure prediction.

The porins of Gram-negative bacteria are responsible for the 'molecular sieve' properties of the outer membrane. They form large water-filled channels which allow the diffusion of hydrophilic molecules into the periplasmic space. Owing to the strong hydrophilicity of their amino acid sequence and the nature of their secondary structure (beta strands), conventional hydropathy methods for predicting membrane topology are useless for this class of protein. The large number of available porin amino acid sequences was exploited to improve the accuracy of the prediction in combination with tools detecting amphipathicity of secondary structure. Using the constraints of beta-sheet structure these porins are predicted to contain 16 membrane-spanning strands, 14 of which are common to the two (enteric and the neisserial) porin subfamilies.

Amino Acid Sequence

Genetic relatedness of human DNA polymerase beta and terminal deoxynucleotidyltransferase.

The Protein Identification Resource (PIR) protein sequence data bank was searched for sequence similarity between known proteins and human DNA polymerase beta (Pol beta) or human terminal deoxynucleotidyltransferase (TdT). Pol beta and TdT were found to exhibit amino acid sequence similarity only with each other and not with any other of the 4750 entries in release 12.0 of the PIR data bank. Optimal amino acid sequence alignment of the entire 39-kDa Pol beta polypeptide with the C-terminal two thirds of TdT revealed 24% identical aa residues and 21% conservative aa substitutions. The Monte Carlo score of 12.6 for the entire aligned sequences indicates highly significant aa sequence homology. The hydropathicity profiles of the aligned aa sequences were remarkably similar throughout, suggesting structural similarity of the polypeptides. The most significant regions of homology are aa residues 39-224 and 311-333 of Pol beta vs. aa residues 191-374 and 484-506 of TdT. In addition, weaker homology was seen between a large portion of the 'nonessential' N-terminal end of TdT (aa residues 33-130) and the first region of strong homology between the two proteins (aa residues 31-128 of Pol beta and aa residues 183-280 of TdT), suggestive of genetic duplication within the ancestral gene. On the basis of nucleotide differences between conserved regions of Pol beta and TdT genes (aligned according to optimally aligned aa sequences) it was estimated that Pol beta and TdT diverged on the order of 250 million years ago, corresponding roughly to a time before radiation of mammals and birds.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Simultaneous and multivariate alignment of protein sequences: correspondence between physicochemical profiles and structurally conserved regions (SCR).

A general protein sequence alignment methodology for detecting a priori unknown common structural and functional regions is described. The method proposed in this paper is based on two basic requirements for a meaningful alignment. First, each sequence or segment of a sequence is characterized by a multivariate physicochemical profile. Second, the alignment is performed by considering all the sequences simultaneously, and the algorithm detects those regions that form a set of similar profiles. In order to test the structural meaning of the alignment obtained from the sequences, quantitative comparisons are performed with structurally conserved regions (SCR) determined from the X-ray structures of three serine proteases. Results suggest that the limits of the SCR may be predicted from the similarities between the physicochemical profiles of the sequences. The procedures are not completely automated. The final step requires a visual screening of alternative pathways in order to determine an optimal alignment.

Algorithms

Modelling of binding sites of the nicotinic acetylcholine receptor and their relation to models of the whole receptor.

Models for the acetylcholine (ACh)-binding site of the nicotinic acetylcholine receptor (nAChR) are proposed. These models have been developed by using the concept of the ligand-gated ion-channel (LGIC) superfamily of receptors that have evolved from a common ancestor. An initial component of the binding site was identified as a highly conserved 15-residue stretch of primary structure in the N-terminal extracellular region of all known LGIC subunits, based on aligned sequence data of LGICs. This subregion, termed the Cys-loop, was modelled as an amphiphilic beta-hairpin and we propose that it forms a major determinant of the binding cleft for agonists. This initial, partial binding-site model has been extended to include residues biochemically identified as spatially adjacent to the binding cleft. A recently developed technique for rapidly scanning the known protein structural database for 'non-homologous similarity' using just sequence information identified the known structure of the enzyme pyrophosphatase (PPase) as a candidate scaffold for the N-terminal domain of the nAChR. This similarity was investigated further using sequence alignments. A framework model of the full N-terminal domain in which the position of the Cys-loop and other binding-site determinants, as well as the main immunogenic region (MIR), have been mapped on to the PPase structure.

Amino Acid Sequence

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms

Evaluation and improvements in the automatic alignment of protein sequences.

The accuracy of protein sequence alignment obtained by applying a commonly used global sequence comparison algorithm is assessed. Alignments based on the superposition of the three-dimensional structures are used as a standard for testing the automatic, sequence-based methods. Alignments obtained from the global comparison of five pairs of homologous protein sequences studied gave 54% agreement overall for residues in secondary structures. The inclusion of information about the secondary structure of one of the proteins in order to limit the number of gaps inserted in regions of secondary structure, improved this figure to 68%. A similarity score of greater than six standard deviation units suggests that an alignment which is greater than 75% correct within secondary structural regions can be obtained automatically for the pair of sequences.

Algorithms

Molecular cloning of a human thyrotropin receptor cDNA fragment. Use of highly degenerate, inosine containing primers derived from aligned amino acid sequences of a homologous family of glycoprotein hormone receptors.

Autoantibodies to the thyrotropin (TSH) hormone receptor (TSH-R) are present in the sera of patients with thyroid autoimmune disease which are pathogenetic leading to hyperthyroidism of Graves' disease. Considerable interest has been focused on the cloning of the human TSH-R, which has until very recently, proven exceedingly difficult due to the very low receptor level expression on thyroid cells. We have used polymerase chain reaction and highly degenerate, inosine containing oligonucleotides derived from sequence alignments of the transmembrane regions 2 and 7 of a number of G-binding protein receptors including the lutropin/choriogonadotropin (LH/CG) receptors to amplify various cDNAs from human thyroid cDNA. Sequencing analysis of 27 different clones revealed that they fall into eight different groups. The very recent publication of the complete nucleotide sequence of the human TSH-R revealed that one of the groups (GT1) containing seven clones which had been sequenced belong to the human TSH-receptor. The sequence of all 7 GT1 clones was identical and in complete concordance with transmembrane regions 2 and 7 of the published TSH-R sequence. Our results show that by designing oligonucleotides to common transmembrane regions of G-binding proteins where the primers are biased in their sequence to the LH/CG receptors it is possible to amplify the TSH-R receptor sequence.

Amino Acid Sequence

Preferred positions of AA and TT dinucleotides in aligned nucleosomal DNA sequences.

Multiple alignment of 118 nucleosomal DNA sequences by maximizing simultaneously match of AA dinucleotides and match of TT dinucleotides results in a pattern of the dinucleotide distributions which is characteristic of the nucleosomal DNA sequences. The AA dinucleotides are found to be distributed symmetrically relative to the TT dinucleotide distribution, around the middle point of the nucleosomal DNA sequence. The distances between major peaks of the distributions are multiples of about 10.4 bases. The peaks of the TT distribution are shifted by 6 bases downstream from the peaks of the AA distribution.

Adenine

Analysis of sequence variation among legume lectins. A ring of hypervariable residues forms the perimeter of the carbohydrate-binding site.

Twelve plant lectins from the Papilionoideae subfamily were selected to represent a range of carbohydrate specificities, and their sequences were aligned. Two variability indices were applied to the aligned sequences and the results were analysed using the three-dimensional structures of concanavalin A and the pea lectin. The areas of greatest variability were located in the carbohydrate-binding site region, forming a perimeter around a well-conserved core. These residues are inferred to be specificity determining, in the manner of antibodies, and the most variable position corresponded to Tyr100 in concanavalin A, a known ligand contact residue. In addition to the five peptide loops known to form the binding site from crystallographic studies, a sixth segment with variable residues was located in the binding-site region, and this may contribute to oligosaccharide specificity. In their overall composition, the lectin sites resemble those of the sugar-transport proteins rather than antibodies. The prospects for modelling lectin binding sites by the methods used for antibodies were also assessed.

Amino Acid Sequence