PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Phylogenies from amino acid sequences aligned with gaps: the problem of gap weighting.

The common but generally overlooked problem of how best to construct phylogenies from orthologous amino acid sequences, when their alignment requires the placement therein of gaps denoting insertions/deletions in the evolutionary history of their genes since their common ancestor, has been studied. Three diverse methods were examined: 1. each missing residue in a gap is weighted as equivalent to the average number of minimum nucleotide replacements in known conjugate amino acid pairs of those same two sequences, which weight necessarily differs for each pair of sequences; 2. each missing residue in a gap is weighted as equivalent to a fixed number of nucleotide replacements; and 3. each gap, regardless of length, is weighted as equivalent to a fixed number of nucleotide replacements. For the flavodoxins, each method yielded a different best tree and suggests that the choice of method may be crucial. For the plant ferredoxins, all methods give results inconsistent with botanical classification and suggests the sequences may not all be orthologous. For the bacterial ferredoxins, the method was less germane than the actual weight used, five different best trees being obtained depending upon the weight. The best tree for all ferredoxins (prokaryotic plus eukaryotic) combined proved to be greatly dependent upon the gap locations with several reasonable aligments yielding different best trees. They also suggest that functional equivalence may well prove to be a poor guide to which residues have a common ancestral codon. The rubredoxin sequences show that a partial internal gene duplication occurred in the Pseudomonas line, probably very soon after its divergence from the other genera. Together, the results clearly indicate that the phylogenetic answer one gets may greatly depend upon how one treats the gaps but they fail to indicate what treatment may be best.

Amino Acid Sequence

Sequence alignment of the G-protein coupled receptor superfamily.

The multitude of G-protein coupled receptor (GPR) superfamily cDNAs recently isolated has exceeded the number of receptor subtypes anticipated by pharmacological studies. Analysis of the sequence similarities and unique features of the members of this family is valuable for designing strategies to isolate related cDNAs, for developing hypotheses concerning substrate-ligand and receptor-effector interactions, and for understanding the evolution of these genes. We have compiled and aligned the 74 unique amino acid sequences published to date and review the present understanding of the structural motifs contributing to ligand binding and G-protein coupling.

Amino Acid Sequence

Sequence alignment and penalty choice. Review of concepts, case studies and implications.

Alignment algorithms to compare DNA or amino acid sequences are widely used tools in molecular biology. The algorithms depend on the setting of various parameters, most notably gap penalties. The effect that such parameters have on the resulting alignments is still poorly understood. This paper begins by reviewing two recent advances in algorithms and probability that enable us to take a new approach to this question. The first tool we introduce is a newly developed method to delineate efficiently all optimal alignments arising under all choices of parameters. The second tool comprises insights into the statistical behavior of optimal alignment scores. From this we gain a better understanding of the dependence of alignments on parameters in general. We propose novel criteria to detect biologically good alignments and highlight some specific features about the interaction between similarity matrices and gap penalties. To illustrate our analysis we present a detailed study of the comparison of two immunoglobulin sequences.

Algorithms

PROANAL version 2: multifunctional program for analysis of multiple protein sequence alignments and for studying the structure--activity relationships in protein families.

A new version of the program PROANAL is described. A multiple linear regression analysis of the protein structure--activity relationship allows one to investigate the combinations of protein sites and factors influencing the activity. The program also provides the possibility to seek out protein sites, conservative or variable in variations of physicochemical characteristics, and regions with high or low values of these characteristics. PROANAL2 may be useful in the simulation of protein-engineering experiments and in the search of a number of protein regions such as functional sites, secondary structures, solvent-exposed regions, T- and B-cell antigenic determinants, etc.

Algorithms

A new interactive protein sequence alignment program and comparison of its results with widely used algorithms.

A computer program that allows interactive sequence comparison is described. It graphically displays a search matrix using residue physiochemical characteristics and multilength segmental comparisons. The user selects through a mousing device and screen pointer the sequence spans to be matched. The results of this method are compared with those of ALIGN and BESTFIT.

Algorithms

An analysis of the periodicity of conserved residues in sequence alignments of G-protein coupled receptors. Implications for the three-dimensional structure.

Twenty-three sequences from the family of G-protein coupled receptors have been aligned according to the 'historical alignment' procedure of Feng and Doolittle. Fourier transform analysis of this reveals that parts of five of the seven putative membrane-spanning regions exhibit a periodicity of conserved/nonconserved residues which is compatible with the periodicity of the alpha-helix. This would place the conserved residues on one side of the helix, which may face the inside of the proposed seven membered helical bundle.

Amino Acid Sequence

A measure of the similarity of sets of sequences not requiring sequence alignment.

Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classifications. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism--for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.

Animals

Structure-based sequence alignment of three AdoMet-dependent DNA methyltransferases.

M.HhaI, M.TaqI and COMT are DNA methyltransferases (MTases) which catalyze the transfer of a methyl group from the cofactor AdoMet to C5 of cytosine, to N6 of adenine and to a hydroxyl group of catechol, respectively. The larger catalytic domains of the bilobal proteins, M.HhaI and M.TaqI, and the entire single domain of COMT have an alpha/beta structure containing a mixed central beta-sheet. These domains have very similar folding. By allowing appropriate 'insertions' or 'deletions' in the backbones of the three structures, it was possible to find more conserved motifs in M.TaqI and COMT. The similarity in protein folding and the equivalence of amino-acid sequences revealed by the structural alignment indicate that many AdoMet-dependent MTases may share a common catalytic domain structure.

Amino Acid Sequence

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence

Maximum entropy weighting of aligned sequences of proteins or DNA.

In a family of proteins or other biological sequences like DNA the various subfamilies are often very unevenly represented. For this reason a scheme for assigning weights to each sequence can greatly improve performance at tasks such as database searching with profiles or other consensus models based on multiple alignments. A new weighting scheme for this type of database search is proposed. In a statistical description of the searching problem it is derived from the maximum entropy principle. It can be proved that, in a certain sense, it corrects for uneven representation. It is shown that finding the maximum entropy weights is an easy optimization problem for which standard techniques are applicable.

Amino Acid Sequence

A flexible multiple sequence alignment program.

The 'regions' method for multisequence alignment used in the previously reported program MALIGN has been generalized to include recursive refinement so that unaligned portions between two regions at the current level of resolution can be handled with increased resolution. Additionally, there is incorporated a limiting of the number of regions to be used at any level of resolution from which to abstract an alignment. This provides a significant increase in speed over the unlimited version. The program GENALIGN uses this improved regions method to execute fast pairwise alignments in the framework of Taylor's multisequence alignment procedure using clustered pairwise alignments. Pairwise alignments by dynamic programming are also provided in the program.

Algorithms

Sequence alignment and evolutionary comparison of the L10 equivalent and L12 equivalent ribosomal proteins from archaebacteria, eubacteria, and eucaryotes.

The genes corresponding to the L10 and L12 equivalent ribosomal proteins (L10e and L12e) of Escherichia coli have been cloned and sequenced from two widely divergent species of archaebacteria, Halobacterium cutirubrum and Sulfolobus solfataricus. The deduced amino acid sequences of the L10e and L12e proteins have been compared to each other and to available eubacterial and eucaryotic sequences. We have identified the human P0 protein as the eucaryotic L10e. The L10e proteins from the three kingdoms were found to be colinear. The eubacterial L10e protein is much shorter than the archaebacterial-eucaryotic proteins because of two large deletions, one internal and one at the carboxy terminus. The archaebacterial and eucaryotic L12e proteins were also colinear; the eubacterial protein is homologous to the archaebacterial and eucaryotic L12e proteins, but has suffered rearrangement through what appear to be gene fusion events. Intraspecies comparisons between L10e and L12e sequences indicate the archaebacterial and eucaryotic L10e proteins contain a partial copy of the L12e protein fused to their carboxy terminus. In the eubacteria most of this fusion has been removed by the carboxy terminal deletion. Within the L12e-derived region, a 26-amino acid-long internal modular sequence reiterated thrice in the archaebacterial L10e, twice in the eucaryotic L10e, and once in the eubacterial L10e was discovered. This modular sequence also appears to be present as a single copy in all L12e proteins and may play a role in L12e dimerization, L10e-L12e complex formation, and the function of L10e-L12e complex in translation.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence