PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

GeneSilico protein structure prediction meta-server.

Rigorous assessments of protein structure prediction have demonstrated that fold recognition methods can identify remote similarities between proteins when standard sequence search methods fail. It has been shown that the accuracy of predictions is improved when refined multiple sequence alignments are used instead of single sequences and if different methods are combined to generate a consensus model. There are several meta-servers available that integrate protein structure predictions performed by various methods, but they do not allow for submission of user-defined multiple sequence alignments and they seldom offer confidentiality of the results. We developed a novel WWW gateway for protein structure prediction, which combines the useful features of other meta-servers available, but with much greater flexibility of the input. The user may submit an amino acid sequence or a multiple sequence alignment to a set of methods for primary, secondary and tertiary structure prediction. Fold-recognition results (target-template alignments) are converted into full-atom 3D models and the quality of these models is uniformly assessed. A consensus between different FR methods is also inferred. The results are conveniently presented on-line on a single web page over a secure, password-protected connection. The GeneSilico protein structure prediction meta-server is freely available for academic users at http://genesilico.pl/meta.

Internet↗

EnteriX 2003: Visualization tools for genome alignments of Enterobacteriaceae.

We describe EnteriX, a suite of three web-based visualization tools for graphically portraying alignment information from comparisons among several fixed and user-supplied sequences from related enterobacterial species, anchored on a reference genome (http://bio.cse.psu.edu/). The first visualization, Enteric, displays stacked pairwise alignments between a reference genome and each of the related bacteria, represented schematically as PIPs (Percent Identity Plots). Encoded in the views are large-scale genomic rearrangement events and functional landmarks. The second visualization, Menteric, computes and displays 1 Kb views of nucleotide-level multiple alignments of the sequences, together with annotations of genes, regulatory sites and conserved regions. The third, a Java-based tool named Maj, displays alignment information in two formats, corresponding roughly to the Enteric and Menteric views, and adds zoom-in capabilities. The uses of such tools are diverse, from examining the multiple sequence alignment to infer conserved sites with potential regulatory roles, to scrutinizing the commonalities and differences between the genomes for pathogenicity or phylogenetic studies. The EnteriX suite currently includes >15 enterobacterial genomes, generates views centered on four different anchor genomes and provides support for including user sequences in the alignments.

Computer Graphics↗

Effects of nucleotide sequence alignment on phylogeny estimation: a case study of 18S rDNAs of apicomplexa.

The reconstruction of phylogenetic history is predicated on being able to accurately establish hypotheses of character homology, which involves sequence alignment for studies based on molecular sequence data. In an empirical study investigating nucleotide sequence alignment, we inferred phylogenetic trees for 43 species of the Apicomplexa and 3 of Dinozoa based on complete small-subunit rDNA sequences, using six different multiple-alignment procedures: manual alignment based on the secondary structure of the 18S rRNA molecule, and automated similarity-based alignment algorithms using the PileUp, ClustalW, TreeAlign, MALIGN, and SAM computer programs. Trees were constructed using neighboring-joining, weighted-parsimony, and maximum-likelihood methods. All of the multiple sequence alignment procedures yielded the same basic structure for the estimate of the phylogenetic relationship among the taxa, which presumably represents the underlying phylogenetic signal. However, the placement of many of the taxa was sensitive to the alignment procedure used; and the different alignments produced trees that were on average more dissimilar from each other than did the different tree-building methods used. The multiple alignments from the different procedures varied greatly in length, but aligned sequence length was not a good predictor of the similarity of the resulting phylogenetic trees. We also systematically varied the gap weights (the relative cost of inserting a new gap into a sequence or extending an already-existing gap) for the ClustalW program, and this produced alignments that were at least as different from each other as those produced by the different alignment algorithms. Furthermore, there was no combination of gap weights that produced the same tree as that from the structure alignment, in spite of the fact that many of the alignments were similar in length to the structure alignment. We also investigated the phylogenetic information content of the helical and nonhelical regions of the rDNA, and conclude that the helical regions are the most informative. We therefore conclude that many of the literature disagreements concerning the phylogeny of the Apicomplexa are probably based on differences in sequence alignment strategies rather than differences in data or tree-building methods.

Algorithms↗

Classification and phylogenetic analysis of the cAMP-dependent protein kinase regulatory subunit family.

The members of the PKA regulatory subunit family (PKA-R family) were analyzed by multiple sequence alignment and clustering based on phylogenetic tree construction. According to the phylogenetic trees generated from multiple sequence alignment of the complete sequences, the PKA-R family was divided into four subfamilies (types I to IV). Members of each subfamily were exclusively from animals (types I and II), fungi (type III), and alveolates (type IV). Application of the same methodology to the cAMP-binding domains, and subsequently to the region delimited by beta-strands 6 and 7 of the crystal structures of bovine RIalpha and rat RIIbeta (the phosphate-binding cassette; PBC), proved that this highly conserved region was enough to classify unequivocally the members of the PKA-R family. A single signature sequence, F-G-E-[LIV]-A-L-[LIMV]-x(3)-[PV]-R-[ANQV]-A, corresponding to the PBC was identified which is characteristic of the PKA-R family and is sufficient to distinguish it from other members of the cyclic nucleotide-binding protein superfamily. Specific determinants for the A and B domains of each R-subunit type were also identified. Conserved residues defining the signature motif are important for interaction with cAMP or for positioning the residues that directly interact with cAMP. Conversely, residues that define subfamilies or domain types are not conserved and are mostly located on the loop that connects alpha-helix B' and beta strand 7.

Amino Acid Sequence↗

Hidden Markov models that use predicted secondary structures for fold recognition.

There are many proteins that share the same fold but have no clear sequence similarity. To predict the structure of these proteins, so called "protein fold recognition methods" have been developed. During the last few years, improvements of protein fold recognition methods have been achieved through the use of predicted secondary structures (Rice and Eisenberg, J Mol Biol 1997;267:1026-1038), as well as by using multiple sequence alignments in the form of hidden Markov models (HMM) (Karplus et al., Proteins Suppl 1997;1:134-139). To test the performance of different fold recognition methods, we have developed a rigorous benchmark where representatives for all proteins of known structure are matched against each other. Using this benchmark, we have compared the performance of automatically-created hidden Markov models with standard-sequence-search methods. Further, we combine the use of predicted secondary structures and multiple sequence alignments into a combined method that performs better than methods that do not use this combination of information. Using only single sequences, the correct fold of a protein was detected for 10% of the test cases in our benchmark. Including multiple sequence information increased this number to 16%, and when predicted secondary structure information was included as well, the fold was correctly identified in 20% of the cases. Moreover, if the correct secondary structure was used, 27% of the proteins could be correctly matched to a fold. For comparison, blast2, fasta, and ssearch identifies the fold correctly in 13-17% of the cases. Thus, standard pairwise sequence search methods perform almost as well as hidden Markov models in our benchmark. This is probably because the automatically-created multiple sequence alignments used in this study do not contain enough diversity and because the current generation of hidden Markov models do not perform very well when built from a few sequences.

Amino Acid Sequence↗

Comparative analysis of seven multiple protein sequence alignment servers: clues to enhance reliability of predictions.

MOTIVATION: The prediction reliability of seven multiple alignment servers currently available on the Internet (ClustalW, MAP, PIMA, Block Maker, MSA, MEME and Match-Box) has been evaluated in terms of power (sensitivity) and confidence (selectivity). Therefore, the alignments obtained have been respectively compared to refined structural alignments for 20 families of related proteins with low levels of identity. RESULTS: Results clearly show that any powerful method remains reliable when the rate of identity falls. For some methods, power and confidence decrease linearly with the rate of identity, while other methods emphasize reliability at the cost of a lower power. Increasing the number of related sequences included in the alignment may either improve or decrease the quality of the predictions substantially. For some methods, the gain in power or in confidence is quite systematic; for others, the effect of the addition of homologous sequences is highly unpredictable. Extracting the consensus between two different methods may increase the overall confidence of the predictions tremendously. Our conclusions induce users of sequence alignment methods on the Internet to select the most suitable technique according to their requirements in terms of selectivity and sensitivity. AVAILABILITY: The aligned sequences of the 20 alignments of structure can be obtained automatically by sending the message 'send: cabios_tests.txt' by e-mail to 'matchbox@biq.fundp.ac.be'. CONTACT: eric.depiereux@fundp.ac.be

Computer Communication Networks↗

HHsenser: exhaustive transitive profile search using HMM-HMM comparison.

HHsenser is the first server to offer exhaustive intermediate profile searches, which it combines with pairwise comparison of hidden Markov models. Starting from a single protein sequence or a multiple alignment, it can iteratively explore whole superfamilies, producing few or no false positives. The output is a multiple alignment of all detected homologs. HHsenser's sensitivity should make it a useful tool for evolutionary studies. It may also aid applications that rely on diverse multiple sequence alignments as input, such as homology-based structure and function prediction, or the determination of functional residues by conservation scoring and functional subtyping.HHsenser can be accessed at http://hhsenser.tuebingen.mpg.de/. It has also been integrated into our structure and function prediction server HHpred (http://hhpred.tuebingen.mpg.de/) to improve predictions for near-singleton sequences.

Internet↗

Divide-and-conquer multiple alignment with segment-based constraints.

A large number of methods for multiple sequence alignment are currently available. Recent benchmarking tests demonstrated that strengths and drawbacks of these methods differ substantially. Global strategies can be outperformed by approaches based on local similarities and vice versa, depending on the characteristics of the input sequences. In recent years, mixed approaches that include both global and local features have shown promising results. Herein, we introduce a new algorithm for multiple sequence alignment that integrates the global divide-and-conquer approach with the local segment-based approach, thereby combining the strengths of those two strategies.

Algorithms↗

VIR: a computational tool for analysis of immunoglobulin sequences.

In this paper a microcomputer software named VIR (Variable domains of the Immune Receptors) is reported. This package can be used in sequence studies of immunoglobulin variable domains. The main features of the VIR software in the sequences management are: (1) ease of information recovery/extraction from amino acid sequences; and (2) its capability to obtain multiple sequence alignments with predefined characteristics (i.e. specie and/or specificity). As an analytical tool, the VIR package employs such multiple sequence alignments to compute: (1) tables showing amino acid frequencies; (2) three variability indexes; (3) identity matrices; (4) random samples; and (5) sequences with possible canonical structures. Thus the software reported here is proposed as a useful tool to carry out detailed studies of immunoglobulin variable domains.

Amino Acid Sequence↗

The PredictProtein server.

PredictProtein (PP, http://cubic.bioc.columbia.edu/pp/) is an internet service for sequence analysis and the prediction of aspects of protein structure and function. Users submit protein sequence or alignments; the server returns a multiple sequence alignment, PROSITE sequence motifs, low-complexity regions (SEG), ProDom domain assignments, nuclear localisation signals, regions lacking regular structure and predictions of secondary structure, solvent accessibility, globular regions, transmembrane helices, coiled-coil regions, structural switch regions and disulfide-bonds. Upon request, fold recognition by prediction-based threading is available. For all services, users can submit their query either by electronic mail or interactively from World Wide Web.

Databases, Protein↗

Method for low resolution prediction of small protein tertiary structure.

A new method for the de novo prediction of protein structures at low resolution has been developed. Starting from a multiple sequence alignment, protein secondary structure is predicted, and only those topological elements with high reliability are selected. Then, the multiple sequence alignment and the secondary structure prediction are combined to predict side chain contacts. Such contact map prediction is carried out in two stages. First, an analysis of correlated mutations is carried out to identify pairs of topological elements of secondary structure which are in contact. Then, inverse folding is used to select compatible fragments in contact, thereby enriching the number and identity of predicted side chain contacts. The final outcome of the procedure is a set of noisy secondary and tertiary restraints. These are used as a restrained potential in a Monte Carlo simulation of simplified protein models driven by statistical potentials. Low energy structures are then searched for by using simulated annealing techniques. Implementation of the restraints is carried out so as to take into account of their low resolution. Using this procedure, it has been possible to predict de novo the structure of three very different protein topologies: an alpha/beta protein, the bovine pancreatic trypsin inhibitor (6pti), an alpha-helical protein, calbindin (3icb), and an all beta- protein, the SH3 domain of spectrin (1shg). In all cases, low resolution folds have been obtained with a root mean square deviation (RMSD) of 4.5-5.5 A with respect to the native structure. Some misfolded topologies appear in the simulations, but it is possible to select the native one on energetic grounds. Thus, it is demonstrated that the methodology is general for all protein motifs. Work is in progress in order to test the methodology on a larger set of protein structures.

Amino Acid Sequence↗

Folding and assembly of hepatitis B virus core protein: a new model proposal.

Hepatitis B core antigen has been intensively studied. Recently, cryoelectron microscopy studies have determined the structure of human and duck hepatitis B virus nucleocapsids at low resolution. Both viruses assemble into core particles of two sizes with icosahedral dimer-clustered T = 3 and T = 4 symmetries. Both capsids present tightly clustered dimers composed of a shell and a protruding domain. The present work introduces a model for HBc folding, dimer formation, and assembly. The model is based in multiple alignments of HBc sequences from 20 mammalian and avian isolates and secondary structure predictions. The 54% alpha-helical conformation predicted is in good agreement with CD results reporting 53-71% content of alpha-helices. Despite the sequence divergence of mammalian and avian proteins, the secondary structure prediction of both shows a high degree of coincidence, according to the multiple sequence alignment. The proposed fold of HBc monomers is built from five alpha-helices. In dimers, pairs of two of those helices conform the protruding domain. The model also suggests the convergence of the region preceding the protamine domain around the sixfold symmetry axes. The model gives answers to most of the standing questions concerning the nucleocapsid assembly and antigenic behavior of HBc protein.

Amino Acid Sequence↗

Parameterization studies for the SAM and HMMER methods of hidden Markov model generation.

Multiple sequence alignment of distantly related viral proteins remains a challenge to all currently available alignment methods. The hidden Markov model approach offers a new, flexible method for the generation of multiple sequence alignments. The results of studies attempting to infer appropriate parameter constraints for the generation of de novo HMMs for globin, kinase, aspartic acid protease, and ribonuclease H sequences by both the SAM and HMMER methods are described.

Aspartic Acid Endopeptidases↗

Combining evolutionary information and neural networks to predict protein secondary structure.

Using evolutionary information contained in multiple sequence alignments as input to neural networks, secondary structure can be predicted at significantly increased accuracy. Here, we extend our previous three-level system of neural networks by using additional input information derived from multiple alignments. Using a position-specific conservation weight as part of the input increases performance. Using the number of insertions and deletions reduces the tendency for overprediction and increases overall accuracy. Addition of the global amino acid content yields a further improvement, mainly in predicting structural class. The final network system has sustained overall accuracy of 71.6% in a multiple cross-validation test on 126 unique protein chains. A test on a new set of 124 recently solved protein structures that have no significant sequence similarity to the learning set confirms the high level of accuracy. The average cross-validated accuracy for all 250 sequence-unique chains is above 72%. Using various data sets, the method is compared to alternative prediction methods, some of which also use multiple alignments: the performance advantage of the network system is at least 6 percentage points in three-state accuracy. In addition, the network estimates secondary structure content from multiple sequence alignments about as well as circular dichroism spectroscopy on a single protein and classifies 75% of the 250 proteins correctly into one of four protein structural classes. Of particular practical importance is the definition of a position-specific reliability index. For 40% of all residues the method has a sustained three-state accuracy of 88%, as high as the overall average for homology modelling. A further strength of the method is greatly increased accuracy in predicting the placement of secondary structure segments.

Amino Acid Sequence↗

AdoMet radical proteins--from structure to evolution--alignment of divergent protein sequences reveals strong secondary structure element conservation.

Eighteen subclasses of S-adenosyl-l-methionine (AdoMet) radical proteins have been aligned in the first bioinformatics study of the AdoMet radical superfamily to utilize crystallographic information. The recently resolved X-ray structure of biotin synthase (BioB) was used to guide the multiple sequence alignment, and the recently resolved X-ray structure of coproporphyrinogen III oxidase (HemN) was used as the control. Despite the low 9% sequence identity between BioB and HemN, the multiple sequence alignment correctly predicted all but one of the core helices in HemN, and correctly predicted the residues in the enzyme active site. This alignment further suggests that the AdoMet radical proteins may have evolved from half-barrel structures (alphabeta)4 to three-quarter-barrel structures (alphabeta)6 to full-barrel structures (alphabeta)8. It predicts that anaerobic ribonucleotide reductase (RNR) activase, an ancient enzyme that, it has been suggested, serves as a link between the RNA and DNA worlds, will have a half-barrel structure, whereas the three-quarter barrel, exemplified by HemN, will be the most common architecture for AdoMet radical enzymes, and fewer members of the superfamily will join BioB in using a complete (alphabeta)8 TIM-barrel fold to perform radical chemistry. These differences in barrel architecture also explain how AdoMet radical enzymes can act on substrates that range in size from 10 atoms to 608 residue proteins.

Amino Acid Sequence↗

PROANAL version 2: multifunctional program for analysis of multiple protein sequence alignments and for studying the structure--activity relationships in protein families.

A new version of the program PROANAL is described. A multiple linear regression analysis of the protein structure--activity relationship allows one to investigate the combinations of protein sites and factors influencing the activity. The program also provides the possibility to seek out protein sites, conservative or variable in variations of physicochemical characteristics, and regions with high or low values of these characteristics. PROANAL2 may be useful in the simulation of protein-engineering experiments and in the search of a number of protein regions such as functional sites, secondary structures, solvent-exposed regions, T- and B-cell antigenic determinants, etc.

Algorithms↗

An algorithm for progressive multiple alignment of sequences with insertions.

Dynamic programming algorithms guarantee to find the optimal alignment between two sequences. For more than a few sequences, exact algorithms become computationally impractical, and progressive algorithms iterating pairwise alignments are widely used. These heuristic methods have a serious drawback because pairwise algorithms do not differentiate insertions from deletions and end up penalizing single insertion events multiple times. Such an unrealistically high penalty for insertions typically results in overmatching of sequences and an underestimation of the number of insertion events. We describe a modification of the traditional alignment algorithm that can distinguish insertion from deletion and avoid repeated penalization of insertions and illustrate this method with a pair hidden Markov model that uses an evolutionary scoring function. In comparison with a traditional progressive alignment method, our algorithm infers a greater number of insertion events and creates gaps that are phylogenetically consistent but spatially less concentrated. Our results suggest that some insertion/deletion "hot spots" may actually be artifacts of traditional alignment algorithms.

Algorithms↗

Context-dependent optimal substitution matrices.

Substitution matrices are a key tool in important applications such as identifying sequence homologies, creating sequence alignments and more recently using evolutionary patterns for the prediction of protein structure. We have derived a novel approach to the derivation of these matrices that utilizes not only multiple sequence alignments, but also the associated evolutionary trees. The key to our method is the use of a Bayesian formalism to calculate the probability that a given substitution matrix fits the tree structures and multiple sequence alignment data. Using this procedure, we can determine optimal substitution matrices for various local environments, depending on parameters such as secondary structure and surface accessibility.

Evolution, Molecular↗