PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Multiple alignment and sorting of peptides derived from phage-displayed random peptide libraries with polyclonal sera allows discrimination of relevant phagotopes.

Biopanning of phage-displayed random peptide libraries is a powerful technique for identifying peptides that mimic epitopes (mimotopes) for monoclonal antibodies (mAbs). However, peptides derived using polyclonal antisera may represent epitopes for a diverse range of antibodies. Hence following screening of phage libraries with polyclonal antisera, including autoimmune disease sera, a procedure is required to distinguish relevant from irrelevant phagotopes. We therefore applied the multiple sequence alignment algorithm PILEUP together with a matrix for scoring amino acid substitutions based on physicochemical properties to generate guide trees depicting relatedness of selected peptides. A random heptapeptide library was biopanned nine times using no selecting antibodies, immunoglobulin G (IgG) from sera of subjects with autoimmune diseases (primary biliary cirrhosis (PBC) and type 1 diabetes) and three murine ascites fluids that contained mAbs to overlapping epitope(s) on the Ross River Virus envelope protein 2. Peptides randomly sampled from the library were distributed throughout the guide tree of the total set of peptides whilst many of the peptides derived in the absence of selecting antibody aligned to a single cluster. Moreover peptides selected by different sources of IgG aligned to separate clusters, each with a different amino acid motif. These alignments were validated by testing all of the 53 phagotopes derived using IgG from PBC sera for reactivity by capture ELISA with antibodies affinity purified on the E2 subunit of the pyruvate dehydrogenase complex (PDC-E2), the major autoantigen in PBC: only those phagotopes that aligned to PBC-associated clusters were reactive. Hence the multiple sequence alignment procedure discriminates relevant from irrelevant phagotopes and thus a major difficulty with biopanning phage-displayed random peptide libraries with polyclonal antibodies is surmounted.

Algorithms↗

On the significance of sequence alignments when using multiple scoring matrices.

MOTIVATION: Pairwise local sequence alignment is commonly used to search data bases for sequences related to some query sequence. Alignments are obtained using a scoring matrix that takes into account the different frequencies of occurrence of the various types of amino acid substitutions. Software like BLAST provides the user with a set of scoring matrices available to choose from, and in the literature it is sometimes recommended to try several scoring matrices on the sequences of interest. The significance of an alignment is usually assessed by looking at E-values and p-values. While sequence lengths and data base sizes enter the standard calculations of significance, it is much less common to take the use of several scoring matrices on the same sequences into account. Altschul proposed corrections of the p-value that account for the simultaneous use of an infinite number of PAM matrices. Here we consider the more realistic situation where the user may choose from a finite set of popular PAM and BLOSUM matrices, in particular the ones available in BLAST. It turns out that the significance of a result can be considerably overestimated, if a set of substitution matrices is used in an alignment problem and the most significant alignment is then quoted. RESULTS: Based on extensive simulations, we study the multiple testing problem that occurs when several scoring matrices for local sequence alignment are used. We consider a simple Bonferroni correction of the p-values and investigate its accuracy. Finally, we propose a more accurate correction based on extreme value distributions fitted to the maximum of the normalized scores obtained from different scoring matrices. For various sets of matrices we provide correction factors which can be easily applied to adjust p- and E-values reported by software packages.

Data Interpretation, Statistical↗

Applicability of the multiple alignment algorithm for detection of weak patterns: periodically distributed DNA pattern as a study case.

MOTIVATION: A nucleosome DNA positioning pattern is known to be one of the weakest (highly degenerated) patterns. The alignment procedure that has been developed recently for the extraction of such a pattern is based on a statistical matching of the sequences, and its success depends on the pattern/background ratio in the individual sequences and in the generated pattern. The heuristic nature of the method and distinctive properties of the pattern bring up the question of efficiency and sensitivity in the procedure. This paper presents a method of verification for this multiple sequence alignment algorithm. RESULTS: To verify the applicability of the multiple alignment approach, we constructed a set of sequences carrying the hidden pattern. The pattern was presented by weak ('signal') oscillations of occurrences of AA and TT dinucleotides along otherwise random sequences. Only a few dinucleotides of any given 145 base long sequence would correspond to the signal, appearing in about the same phase within the simulated periodic pattern. The novelty of our simulation approach is that we simulated a database as a whole, as opposed to simulating each sequence separately. The correlation between the hidden pattern and a sequence from the database is negligible on average, but our statistical multicycle alignment procedure produced the pattern with attributes very close to the simulated ones. The accuracy of the procedure was tested and calibrated. The presence in a typical sequence of as little as three dinucleotides corresponding to the signal is sufficient to generate (detect) the pattern hidden in a collection of 204 sequences.

Algorithms↗

Structure and sequence variation of the trypanosome spliced leader transcript.

We have assessed the potential of using the spliced leader (SL) or mini-exon gene as a marker for molecular phylogenetic analysis of genus Trypanosoma. A total of 27 trypanosome sequences were compared, 18 of these being newly reported. In contrast to genus Leishmania, we found the non-transcribed spacer region of the SL locus in trypanosomes to be far too variable for informative comparison of all but the most closely related species. At the other extreme, the short (39 nt) SL exon was usually completely conserved and hence uninformative. The SL RNA showed variation in both length (97-152 nt) and sequence among different trypanosome species, with most variation occurring in stem-loop II. Consequently, this region could not be aligned with confidence in multiple sequence alignment, severely reducing the number of phylogenetically informative nucleotide positions. In computer simulation, most of the SL RNAs readily folded into the 3 stem-loop secondary structure predicted previously, but again stem-loop II was highly variable. No obvious correlation could be seen between the length of this stem-loop and trypanosome biology. We conclude that the SL repeat is not an informative phylogenetic marker for long range evolutionary studies of genus Trypanosoma.

Animals↗

Structure-guided recombination creates an artificial family of cytochromes P450.

Creating artificial protein families affords new opportunities to explore the determinants of structure and biological function free from many of the constraints of natural selection. We have created an artificial family comprising 3,000 P450 heme proteins that correctly fold and incorporate a heme cofactor by recombining three cytochromes P450 at seven crossover locations chosen to minimize structural disruption. Members of this protein family differ from any known sequence at an average of 72 and by as many as 109 amino acids. Most (>73%) of the properly folded chimeric P450 heme proteins are catalytically active peroxygenases; some are more thermostable than the parent proteins. A multiple sequence alignment of 955 chimeras, including both folded and not, is a valuable resource for sequence-structure-function studies. Logistic regression analysis of the multiple sequence alignment identifies key structural contributions to cytochrome P450 heme incorporation and peroxygenase activity and suggests possible structural differences between parents CYP102A1 and CYP102A2.

Amino Acid Sequence↗

Molecular modeling of family GH16 glycoside hydrolases: potential roles for xyloglucan transglucosylases/hydrolases in cell wall modification in the poaceae.

Family GH16 glycoside hydrolases can be assigned to five subgroups according to their substrate specificities, including xyloglucan transglucosylases/hydrolases (XTHs), (1,3)-beta-galactanases, (1,4)-beta-galactanases/kappa-carrageenases, "nonspecific" (1,3/1,3;1,4)-beta-D-glucan endohydrolases, and (1,3;1,4)-beta-D-glucan endohydrolases. A structured family GH16 glycoside hydrolase database has been constructed (http://www.ghdb.uni-stuttgart.de) and provides multiple sequence alignments with functionally annotated amino acid residues and phylogenetic trees. The database has been used for homology modeling of seven glycoside hydrolases from the GH16 family with various substrate specificities, based on structural coordinates for (1,3;1,4)-beta-D-glucan endohydrolases and a kappa-carrageenase. In combination with multiple sequence alignments, the models predict the three-dimensional (3D) dispositions of amino acid residues in the substrate-binding and catalytic sites of XTHs and (1,3/1,3;1,4)-beta-d-glucan endohydrolases; there is no structural information available in the databases for the latter group of enzymes. Models of the XTHs, compared with the recently determined structure of a Populus tremulos x tremuloides XTH, reveal similarities with the active sites of family GH11 (1,4)-beta-D-xylan endohydrolases. From a biological viewpoint, the classification, molecular modeling and a new 3D structure of the P. tremulos x tremuloides XTH establish structural and evolutionary connections between XTHs, (1,3;1,4)-beta-D-glucan endohydrolases and xylan endohydrolases. These findings raise the possibility that XTHs from higher plants could be active not only on cell wall xyloglucans, but also on (1,3;1,4)-beta-D-glucans and arabinoxylans, which are major components of walls in grasses. A role for XTHs in (1,3;1,4)-beta-D-glucan and arabinoxylan modification would be consistent with the apparent overrepresentation of XTH sequences in cereal expressed sequence tags databases.

Amino Acid Sequence↗

Seventy-five percent accuracy in protein secondary structure prediction.

In this study we present an accurate secondary structure prediction procedure by using an query and related sequences. The most novel aspect of our approach is its reliance on local pairwise alignment of the sequence to be predicted with each related sequence rather than utilization of a multiple alignment. The residue-by-residue accuracy of the method is 75% in three structural states after jack-knife tests. The gain in prediction accuracy compared with the existing techniques, which are at best 72%, is achieved by secondary structure propensities based on both local and long-range effects, utilization of similar sequence information in the form of carefully selected pairwise alignment fragments, and reliance on a large collection of known protein primary structures. The method is especially appropriate for large-scale sequence analysis of efforts such as genome characterization, where precise and significant multiple sequence alignments are not available or achievable.

Algorithms↗

Phylogenetic analysis of family 6 glycoside hydrolases.

Multiple sequence alignment separates members of glycoside hydrolase Family 6 into eight subfamilies: one of mainly actinobacterial endoglucanases (EGs), one of ascomycotal EGs, one of chytridiomycotal EGs and cellobiohydrolases (CBHs), one of actinobacterial and proteobacterial CBHs, one of chytridiomycotal CBHs, two of ascomycotal CBHs, and one of basidiomycotal CBHs. Each also has some proteins of unknown function. Multiple sequence alignment also extends to all of Family 6 the observation that lengths of loops that form the active-site tunnel in CBHs vary among subfamilies, and along with loop conformations, determine enzyme function.

Amino Acid Sequence↗

A simulated annealing algorithm for finding consensus sequences.

MOTIVATION: A consensus sequence for a family of related sequences is, as the name suggests, a sequence that captures the features common to most members of the family. Consensus sequences are important in various DNA sequencing applications and are a convenient way to characterize a family of molecules. RESULTS: This paper describes a new algorithm for finding a consensus sequence, using the popular optimization method known as simulated annealing. Unlike the conventional approach of finding a consensus sequence by first forming a multiple sequence alignment, this algorithm searches for a sequence that minimises the sum of pairwise distances to each of the input sequences. The resulting consensus sequence can then be used to induce a multiple sequence alignment. The time required by the algorithm scales linearly with the number of input sequences and quadratically with the length of the consensus sequence. We present results demonstrating the high quality of the consensus sequences and alignments produced by the new algorithm. For comparison, we also present similar results obtained using ClustalW. The new algorithm outperforms ClustalW in many cases.

Algorithms↗

Improving contact predictions by the combination of correlated mutations and other sources of sequence information.

We have previously developed a method for predicting interresidue contacts using information about correlated mutations in multiple sequence alignments. The predictions generated with this method were clearly better than random but not enough for their use in de novo protein folding experiments. We assess the possibility of improving contact predictions combining information from the following variables: correlated mutations, sequence conservation, sequence separation along the chain, alignment stability, family size, residue-specific contact occupancy and formation of contact networks. The application of a protocol for combining these independent variables leads to contact predictions that are on average two times better than those obtained initially with correlated mutations. Correlated mutations can be effectively combined with other types of information derived from multiple sequence alignments. Among the different variables tried, sequence conservation and contact density are particularly relevant for the combination with correlated mutations.

Amino Acid Sequence↗

Incorporating background frequency improves entropy-based residue conservation measures.

BACKGROUND: Several entropy-based methods have been developed for scoring sequence conservation in protein multiple sequence alignments. High scoring amino acid positions may correlate with structurally or functionally important residues. However, amino acid background frequencies are usually not taken into account in these entropy-based scoring schemes. RESULTS: We demonstrate that using a relative entropy measure that incorporates amino acid background frequency results in improved performance in identifying functional sites from protein multiple sequence alignments. CONCLUSION: Our results suggest that the application of appropriate background frequency information may lead to more biologically relevant results in many areas of bioinformatics.

Amino Acid Sequence↗

Sequence diversification of the FK506-binding proteins in several different genomes.

Sequences of FK506-binding proteins (FKBPs) from four genomes of the following organisms were compared: the prokaryote Escherichia coli, the lower eukaryote Saccharomyces cerevisiae, the plant Arabidopsis thaliana, the nematode Caenorhabditis elegans and a composite of 14 unique FKBPs from two mammalian organisms Homo sapiens (man) and Mus musculus (domestic mouse). A singular FK506-like binding domain (FKBD) has about 12 kDa and occurs in the form of archetypal FKBP-12 and as a part of different proteins ranging in size from 13 to 135 kDa. Some organisms may contain a variable number of proteins which consist from two to four consecutively fused FKBDs. In the 12-kDa subgroup of archetypal FKBPs sequence identity (ID) varies from 100 to 83% (mammalian FKBPs-12), 75-50% in mammalian vs. invertebrate FKBPs-12, and fall to about 30% for pairwise sequence comparisons of mammalian and bacterial FKBPs-12 which suggests that their sequences are divergent. Multiple sequence alignment of FKBPs from the four genomes and a set of unique mammalian FKBPs does not contain any explicit consensus sequence but certain sequence positions have conserved physico-chemical characteristics. Variations of hydrophobicity and bulkiness in the multiple sequence alignment are nonsymmetrical because the physico-chemical properties of the aligned sequences changed during evolution. These variations at the sequence positions which are crucial for binding the immunosuppressive macrolide FK506 and peptidyl-prolyl cis/trans isomerase (PPIase) activity are small.

Amino Acid Sequence↗

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms↗

Combined multiple sequence reduced protein model approach to predict the tertiary structure of small proteins.

By incorporating predicted secondary and tertiary restraints into ab initio folding simulations, low resolution tertiary structures of a test set of 20 nonhomologous proteins have been predicted. These proteins, which represent all secondary structural classes, contain from 37 to 100 residues. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Predicted tertiary restraints are obtained from multiple sequence alignments via a two-step process: First, "seed" side chain contacts are identified from a correlated mutation analysis, and then, the seed contacts are "expanded" by an inverse folding algorithm. These predicted restraints are then incorporated into a lattice based, reduced protein model. Depending upon fold complexity, the resulting nativelike topologies exhibit a coordinate root-mean-square deviation, cRMSD, from native between 3.1 and 6.7 A. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

PdbAlign, PdbDist and DistAlign: tools to aid in relating sequence variability to structure.

Many sequence analysis problems involve consideration of a multiple sequence alignment where the 3-dimensional structure of one (or more) of the aligned sequences is known. In such cases, it is useful to map the sequence variability onto the atomic co-ordinates of known structure. If the structure also includes a bound ligand (or the location of the active site is known), each column position in the multiple sequence alignment may be annotated with its 'distance' from the binding site. These annotations, together with a measure of sequence variability, provide additional insights into drug specificity, for example among viral mutants. This paper describes several useful programs that automate this analysis.

Amino Acid Sequence↗

Four-helix bundle: a ubiquitous sensory module in prokaryotic signal transduction.

MOTIVATION: Transmembrane chemoreceptors in Escherichia coli utilize ligand-binding domains for detecting various external signals. The structure of this domain in the E.coli aspartate receptor, Tar, is known and its signal transduction mechanism is under investigation. Current domain models for this important sensory module are inaccurate and, therefore, cannot reveal the distribution of this domain within the current genomic landscape. RESULTS: We carried out sensitive and exhaustive PSI-BLAST searches initiated with the sequence corresponding to a known structure of the four-helix, ligand-binding domain of the aspartate chemoreceptor. From the resulting sequences, we built a multiple sequence alignment for this domain family, which confirmed that the current TarH model is erroneous and fails to detect most of the domain homologs. In the process, we developed a technique that visualizes the secondary structure prediction of each protein sequence in order to improve the multiple sequence alignment. We found that the four-helix up-and-down bundle represents a large domain family and includes representatives of all major classes of prokaryotic signal transduction, namely histidine kinases, di-guanylate cyclases and chemotaxis receptors.

Amino Acid Sequence↗

Can three-dimensional contacts in protein structures be predicted by analysis of correlated mutations?

A method has been developed to detect pairs of positions with correlated mutations in protein multiple sequence alignments. The method is based on reconstruction of the phylogenetic tree for a set of sequences and statistical analysis of the distribution of mutations in the branches of the tree. The database of homology-derived protein structures (HSSP) is used as the source of multiple sequence alignments for proteins of known three-dimensional structure. We analyse pairs of positions with correlated mutations in 67 protein families and show quantitatively that the presence of such positions is a typical feature of protein families. A significant but weak tendency is observed for correlated residue pairs to be close in the three-dimensional structure. With further improvements, methods of this type may be useful for the prediction of residue--residue contacts and subsequent prediction of protein structure using distance geometry algorithms. In conclusion, we suggest a new experimental approach to protein structure determination in which selection of functional mutants after random mutagenesis and analysis of correlated mutations provide sufficient proximity constraints for calculation of the protein fold.

Amino Acid Sequence↗

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157↗