PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Optimal sequence alignments.

Current theory is adequate to the task of finding an optimal alignment between two character strings such as nucleic acids. Most algorithms currently in use must fail to find the homologous alignment between a set of codons for the chicken alpha- and beta-hemoglobin sequence when it is in fact discoverable by a more general treatment of gaps. Fundamental reasons for this are discussed.

Journal Article↗

Structure-based sequence alignment of three AdoMet-dependent DNA methyltransferases.

M.HhaI, M.TaqI and COMT are DNA methyltransferases (MTases) which catalyze the transfer of a methyl group from the cofactor AdoMet to C5 of cytosine, to N6 of adenine and to a hydroxyl group of catechol, respectively. The larger catalytic domains of the bilobal proteins, M.HhaI and M.TaqI, and the entire single domain of COMT have an alpha/beta structure containing a mixed central beta-sheet. These domains have very similar folding. By allowing appropriate 'insertions' or 'deletions' in the backbones of the three structures, it was possible to find more conserved motifs in M.TaqI and COMT. The similarity in protein folding and the equivalence of amino-acid sequences revealed by the structural alignment indicate that many AdoMet-dependent MTases may share a common catalytic domain structure.

Amino Acid Sequence↗

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence↗

Fold recognition and accurate sequence-structure alignment of sequences directing beta-sheet proteins.

The ability to predict structure from sequence is particularly important for toxins, virulence factors, allergens, cytokines, and other proteins of public health importance. Many such functions are represented in the parallel beta-helix and beta-trefoil families. A method using pairwise beta-strand interaction probabilities coupled with evolutionary information represented by sequence profiles is developed to tackle these problems for the beta-helix and beta-trefoil folds. The algorithm BetaWrapPro employs a "wrapping" component that may capture folding processes with an initiation stage followed by processive interaction of the sequence with the already-formed motifs. BetaWrapPro outperforms all previous motif recognition programs for these folds, recognizing the beta-helix with 100% sensitivity and 99.7% specificity and the beta-trefoil with 100% sensitivity and 92.5% specificity, in crossvalidation on a database of all nonredundant known positive and negative examples of these fold classes in the PDB. It additionally aligns 88% of residues for the beta-helices and 86% for the beta-trefoils accurately (within four residues of the exact position) to the structural template, which is then used with the side-chain packing program SCWRL to produce 3D structure predictions. One striking result has been the prediction of an unexpected parallel beta-helix structure for a pollen allergen, and its recent confirmation through solution of its structure. A Web server running BetaWrapPro is available and outputs putative PDB-style coordinates for sequences predicted to form the target folds.

Algorithms↗

Maximum entropy weighting of aligned sequences of proteins or DNA.

In a family of proteins or other biological sequences like DNA the various subfamilies are often very unevenly represented. For this reason a scheme for assigning weights to each sequence can greatly improve performance at tasks such as database searching with profiles or other consensus models based on multiple alignments. A new weighting scheme for this type of database search is proposed. In a statistical description of the searching problem it is derived from the maximum entropy principle. It can be proved that, in a certain sense, it corrects for uneven representation. It is shown that finding the maximum entropy weights is an easy optimization problem for which standard techniques are applicable.

Amino Acid Sequence↗

A flexible multiple sequence alignment program.

The 'regions' method for multisequence alignment used in the previously reported program MALIGN has been generalized to include recursive refinement so that unaligned portions between two regions at the current level of resolution can be handled with increased resolution. Additionally, there is incorporated a limiting of the number of regions to be used at any level of resolution from which to abstract an alignment. This provides a significant increase in speed over the unlimited version. The program GENALIGN uses this improved regions method to execute fast pairwise alignments in the framework of Taylor's multisequence alignment procedure using clustered pairwise alignments. Pairwise alignments by dynamic programming are also provided in the program.

Algorithms↗

Sequence alignment and evolutionary comparison of the L10 equivalent and L12 equivalent ribosomal proteins from archaebacteria, eubacteria, and eucaryotes.

The genes corresponding to the L10 and L12 equivalent ribosomal proteins (L10e and L12e) of Escherichia coli have been cloned and sequenced from two widely divergent species of archaebacteria, Halobacterium cutirubrum and Sulfolobus solfataricus. The deduced amino acid sequences of the L10e and L12e proteins have been compared to each other and to available eubacterial and eucaryotic sequences. We have identified the human P0 protein as the eucaryotic L10e. The L10e proteins from the three kingdoms were found to be colinear. The eubacterial L10e protein is much shorter than the archaebacterial-eucaryotic proteins because of two large deletions, one internal and one at the carboxy terminus. The archaebacterial and eucaryotic L12e proteins were also colinear; the eubacterial protein is homologous to the archaebacterial and eucaryotic L12e proteins, but has suffered rearrangement through what appear to be gene fusion events. Intraspecies comparisons between L10e and L12e sequences indicate the archaebacterial and eucaryotic L10e proteins contain a partial copy of the L12e protein fused to their carboxy terminus. In the eubacteria most of this fusion has been removed by the carboxy terminal deletion. Within the L12e-derived region, a 26-amino acid-long internal modular sequence reiterated thrice in the archaebacterial L10e, twice in the eucaryotic L10e, and once in the eubacterial L10e was discovered. This modular sequence also appears to be present as a single copy in all L12e proteins and may play a role in L12e dimerization, L10e-L12e complex formation, and the function of L10e-L12e complex in translation.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

A comparison of position-specific score matrices based on sequence and structure alignments.

Sequence comparison methods based on position-specific score matrices (PSSMs) have proven a useful tool for recognition of the divergent members of a protein family and for annotation of functional sites. Here we investigate one of the factors that affects overall performance of PSSMs in a PSI-BLAST search, the algorithm used to construct the seed alignment upon which the PSSM is based. We compare PSSMs based on alignments constructed by global sequence similarity (ClustalW and ClustalW-pairwise), local sequence similarity (BLAST), and local structure similarity (VAST). To assess performance with respect to identification of conserved functional or structural sites, we examine the accuracy of the three-dimensional molecular models predicted by PSSM-sequence alignments. Using the known structures of those sequences as the standard of truth, we find that model accuracy varies with the algorithm used for seed alignment construction in the pattern local-structure (VAST) > local-sequence (BLAST) > global-sequence (ClustalW). Using structural similarity of query and database proteins as the standard of truth, we find that PSSM recognition sensitivity depends primarily on the diversity of the sequences included in the alignment, with an optimum around 30-50% average pairwise identity. We discuss these observations, and suggest a strategy for constructing seed alignments that optimize PSSM-sequence alignment accuracy and recognition sensitivity.

Algorithms↗

The bacterial porin superfamily: sequence alignment and structure prediction.

The porins of Gram-negative bacteria are responsible for the 'molecular sieve' properties of the outer membrane. They form large water-filled channels which allow the diffusion of hydrophilic molecules into the periplasmic space. Owing to the strong hydrophilicity of their amino acid sequence and the nature of their secondary structure (beta strands), conventional hydropathy methods for predicting membrane topology are useless for this class of protein. The large number of available porin amino acid sequences was exploited to improve the accuracy of the prediction in combination with tools detecting amphipathicity of secondary structure. Using the constraints of beta-sheet structure these porins are predicted to contain 16 membrane-spanning strands, 14 of which are common to the two (enteric and the neisserial) porin subfamilies.

Amino Acid Sequence↗

Selective loss of either the epimerase or kinase activity of UDP-N-acetylglucosamine 2-epimerase/N-acetylmannosamine kinase due to site-directed mutagenesis based on sequence alignments.

N-Acetylneuraminic acid is the most common naturally occurring sialic acid, as well as being the biosynthetic precursor of this group of compounds. UDP-GlcNAc 2-epimerase/N-acetylmannosamine kinase has been shown to be the key enzyme of N-acetylneuraminic acid biosynthesis in rat liver, and it is a regulator of cell surface sialylation. The N-terminal region of this bifunctional enzyme displays sequence similarities with prokaryotic UDP-GlcNAc 2-epimerases, whereas the sequence of its C-terminal region is similar to sequences of members of the sugar kinase superfamily. High level overexpression of active enzyme was established by using the baculovirus/Sf9 system. For functional characterization, site-directed mutagenesis was performed on different conserved amino acid residues. The histidine mutants H45A, H110A, H132A, H155A, and H157A showed a drastic loss of epimerase activity with almost unchanged kinase activity. Conversely, the mutants D413N, D413K, and R420M in the putative kinase active site lost their kinase activity but retained their epimerase activity. To estimate the structural perturbation effect due to site-directed mutagenesis, the oligomeric state of all mutants was determined by gel filtration analysis. The mutants D413N, D413K, and R420M as well as H45A were shown to form a hexamer like the wild-type enzyme, indicating little influence of mutation on protein folding. Histidine mutants H155A and H157A formed mainly trimeric enzyme with small amounts of hexamer. Oligomerization of mutants H110A and H132A was also significantly different from that of the wild-type enzyme. Therefore the loss of epimerase activity in mutants H110A, H132A, H155A, and H157A can largely be attributed to incorrect protein folding. In contrast, the mutation site of mutant H45A seems to be involved directly in the epimerization process, and the amino acids Asp-413 and Arg-420 of UDP-GlcNAc 2-epimerase/N-acetylmannosamine kinase are essential for the phosphorylation process. The fact that either epimerase or kinase activity are lost selectively provides evidence for the existence of two active sites working quite independently.

Amino Acid Sequence↗

Improving protein secondary structure prediction with aligned homologous sequences.

Most recent protein secondary structure prediction methods use sequence alignments to improve the prediction quality. We investigate the relationship between the location of secondary structural elements, gaps, and variable residue positions in multiple sequence alignments. We further investigate how these relationships compare with those found in structurally aligned protein families. We show how such associations may be used to improve the quality of prediction of the secondary structure elements, using the Quadratic-Logistic method with profiles. Furthermore, we analyze the extent to which the number of homologous sequences influences the quality of prediction. The analysis of variable residue positions shows that surprisingly, helical regions exhibit greater variability than do coil regions, which are generally thought to be the most common secondary structure elements in loops. However, the correlation between variability and the presence of helices does not significantly improve prediction quality. Gaps are a distinct signal for coil regions. Increasing the coil propensity for those residues occurring in gap regions enhances the overall prediction quality. Prediction accuracy increases initially with the number of homologues, but changes negligibly as the number of homologues exceeds about 14. The alignment quality affects the prediction more than other factors, hence a careful selection and alignment of even a small number of homologues can lead to significant improvements in prediction accuracy.

Amino Acid Sequence↗

Novel use of a genetic algorithm for protein structure prediction: searching template and sequence alignment space.

A novel genetic algorithm was applied to all CASP5 targets. The algorithm simultaneously searches template and alignment space. Results show that the current implementation of the method is perhaps most useful in recognizing and refining remote homology targets. This new method is briefly described and results are analyzed. Strengths and weaknesses of the current implementation of the algorithm are discussed.

Algorithms↗

Toward an accurate statistics of gapped alignments.

Sequence alignment has been an invaluable tool for finding homologous sequences. The significance of the homology found is often quantified statistically by p-values. Theory for computing p-values exists for gapless alignments [Karlin, S., Altschul, S.F., 1990. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes. Proc. Natl. Acad. Sci. USA 87, 2264-2268; Karlin, S., Dembo A., 1992. Limit distributions of maximal segmental score among Markov-dependent partial sums. Adv. Appl. Probab. 24, 13-140], but a full generalization to alignments with gaps is not yet complete. We present a unified statistical analysis of two common sequence comparison algorithms: maximum-score (Smith-Waterman) alignments and their generalized probabilistic counterparts, including maximum-likelihood alignments and hidden Markov models. The most important statistical characteristic of these algorithms is the distribution function of the maximum score S(max), resp. the maximum free energy F(max), for mutually uncorrelated random sequences. This distribution is known empirically to be of the Gumbel form with an exponential tail P(S(max)>x) approximately exp(-lambdax) for maximum-score alignment and P(F(max)>x) approximately exp(-lambdax) for some classes of probabilistic alignment. We derive an exact expression for lambda for particular probabilistic alignments. This result is then used to obtain accurate lambda values for generic probabilistic and maximum-score alignments. Although the result demonstrated uses a simple match-mismatch scoring system, it is expected to be a good starting point for more general scoring functions.

Algorithms↗

Catalytic properties, molecular composition and sequence alignments of pyruvate: ferredoxin oxidoreductase from the methanogenic archaeon Methanosarcina barkeri (strain Fusaro).

Methanosarcina barkeri (strain Fusaro) was grown on pyruvate as methanogenic substrate [Bock, A. K., Prieger-Kraft, A. & Schönheit, P. (1994) Arch. Microbiol. 161, 33-46]. The first enzyme of pyruvate catabolism, pyruvate oxidoreductase, which catalyzes oxidation of pyruvate to acetyl-CoA was purified about 90-fold to apparent electrophoretic homogeneity. The purified enzyme catalyzed the CoA-dependent oxidation of pyruvate with ferredoxin as an electron acceptor which defines the enzyme as a pyruvate: ferredoxin oxidoreductase. The deazaflavin, coenzyme F420, which has been proposed to be the physiological electron acceptor of pyruvate oxidoreductase in methanogens, was not reduced by the purified enzyme. In addition to ferredoxin and viologen dyes, flavin nucleotides served as electron acceptors. Pyruvate: ferredoxin oxidoreductase also catalyzed the oxidation of 2-oxobutyrate but not the oxidation of 2-oxoglutarate, indolepyruvate, phenylpyruvate, glyoxylate, 3-hydroxypyruvate and oxaloacetate. The apparent Km values of pyruvate:ferredoxin oxidoreductase were 70 microM for pyruvate, 6 microM for CoA and 30 microM for clostridial ferredoxin. The apparent Vmax with ferredoxin was about 30 U/mg (at 37 degrees C) with a pH optimum of approximately 7. The temperature optimum was approximately 60 degrees C and the Arrhenius activation energy was 40 kJ/mol (between 30 degrees C and 60 degrees C). The enzyme was extremely oxygen sensitive, losing 90% of its activity upon exposure to air for 1 h at 0 degrees C. Sodium nitrite inhibited the enzyme with a Ki of about 10 mM. The native enzyme had an apparent molecular mass of approximately 130 kDa and was composed of four different subunits with apparent molecular masses of 48, 30, 25, and 15 kDa which indicates that the enzyme has an alpha beta gamma delta structure. The enzyme contained 1 mol/mol thiamine diphosphate, and about 12 mol/mol each of non-heme iron and acid-labile sulfur. FAD, FMN and lipoic acid were not found. The N-terminal amino acid sequences of the four subunits were determined. The sequence of the alpha-subunit was similar to the N-terminal amino acid sequence of the alpha-subunit of the heterotetrameric pyruvate:ferredoxin oxidoreductases of the hyperthermophiles Archaeoglobus fulgidus, Pyrococcus furiosus and Thermotoga maritima and of the mesophile Helicobacter pylori, and to the N-terminal amino acid sequence of the homodimeric pyruvate:ferredoxin oxidoreductase from proteobacteria and from cyanobacteria. No sequence similarities were found, however, between the alpha-subunit of the M. barkeri enzyme and the heterodimeric pyruvate:ferredoxin oxidoreductase of the archaeon Halobacterium halobium.

Amino Acid Sequence↗