PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Can three-dimensional contacts in protein structures be predicted by analysis of correlated mutations?

A method has been developed to detect pairs of positions with correlated mutations in protein multiple sequence alignments. The method is based on reconstruction of the phylogenetic tree for a set of sequences and statistical analysis of the distribution of mutations in the branches of the tree. The database of homology-derived protein structures (HSSP) is used as the source of multiple sequence alignments for proteins of known three-dimensional structure. We analyse pairs of positions with correlated mutations in 67 protein families and show quantitatively that the presence of such positions is a typical feature of protein families. A significant but weak tendency is observed for correlated residue pairs to be close in the three-dimensional structure. With further improvements, methods of this type may be useful for the prediction of residue--residue contacts and subsequent prediction of protein structure using distance geometry algorithms. In conclusion, we suggest a new experimental approach to protein structure determination in which selection of functional mutants after random mutagenesis and analysis of correlated mutations provide sufficient proximity constraints for calculation of the protein fold.

Amino Acid Sequence↗

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157↗

Rfam: an RNA family database.

Rfam is a collection of multiple sequence alignments and covariance models representing non-coding RNA families. Rfam is available on the web in the UK at http://www.sanger.ac.uk/Software/Rfam/ and in the US at http://rfam.wustl.edu/. These websites allow the user to search a query sequence against a library of covariance models, and view multiple sequence alignments and family annotation. The database can also be downloaded in flatfile form and searched locally using the INFERNAL package (http://infernal.wustl.edu/). The first release of Rfam (1.0) contains 25 families, which annotate over 50 000 non-coding RNA genes in the taxonomic divisions of the EMBL nucleotide database.

Animals↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.

Animals↗

Sequence comparison by sequence harmony identifies subtype-specific functional sites.

Multiple sequence alignments are often used to reveal functionally important residues within a protein family. They can be particularly useful for the identification of key residues that determine functional differences between protein subfamilies. We present a new entropy-based method, Sequence Harmony (SH) that accurately detects subfamily-specific positions from a multiple sequence alignment. The SH algorithm implements a novel formula, able to score compositional differences between subfamilies, without imposing conservation, in a simple manner on an intuitive scale. We compare our method with the most important published methods, i.e. AMAS, TreeDet and SDP-pred, using three well-studied protein families: the receptor-binding domain (MH2) of the Smad family of transcription factors, the Ras-superfamily of small GTPases and the MIP-family of integral membrane transporters. We demonstrate that SH accurately selects known functional sites with higher coverage than the other methods for these test-cases. This shows that compositional differences between protein subfamilies provide sufficient basis for identification of functional sites. In addition, SH selects a number of sites of unknown function that could be interesting candidates for further experimental investigation.

Algorithms↗

Amino acid sequence analysis of bovine rotavirus B223 reveals a unique outer capsid protein VP4 and confirms a third bovine VP4 type.

The nucleotide and deduced amino acid sequence of the gene 4 of bovine rotavirus strain B223 is described. The open reading frame is predicted to encode a VP4 of 772 amino acids, shorter than described for any other rotavirus strain sequenced to date. B223 VP4 shows 70 to 73% similarity to other rotavirus VP4 proteins, demonstrating the presence of a unique VP4 type, and confirming a third VP4 allele in the bovine rotavirus population. Multiple sequence alignment with several other rotavirus strains created gaps in the sequence to account for a shorter VP4. The alignment shows a two contiguous amino acid deletions within the trypsin cleavage region of B223 VP4. Comparisons of two regions flanking the trypsin cleavage site, (aa 224 to 235, and aa 257 to 271) which show high homologies between strains, demonstrate that the region 5' to the trypsin cut site has a low homology (66%) to other rotavirus strains, although the region 3' to the trypsin cleavage site shows high homologies (86 to 93%) with other rotavirus strains. The lack of a conserved proline residue within the 5' flanking region suggests a possible altered local conformation of this site in B223 VP4. A second gap inserted into the VP4 of B223 on multiple sequence alignment is a three contiguous amino acid deletion at position 613-615 in the VP5* subunit. Previously defined biologic properties of this strain in relation to the determination of the amino acid composition of VP4 are discussed.

Alleles↗

Homology modeling and active-site residues probing of the thermophilic Alicyclobacillus acidocaldarius esterase 2.

The moderate thermophilic eubacterium Alicyclobacillus (formerly Bacillus) acidocaldarius expresses a thermostable carboxylesterase (esterase 2) belonging to the hormone-sensitive lipase (HSL)-like group of the esterase/lipase family. Based on secondary structures predictions and a secondary structure-driven multiple sequence alignment with remote homologous protein of known three-dimensional (3D) structure, we previously hypothesized for this enzyme the alpha/beta-hydrolase fold typical of several lipases and esterases and identified Ser155, Asp252, and His282 as the putative members of the catalytic triad. In this paper we report the construction of a 3D model for this enzyme based on the structure of mouse acetylcholinesterase complexed with fasciculin. The model reveals the topological organization of the fold corroborating our predictions. As regarding the active-site residues, Ser155, Asp252, and His282 are located close to each other at hydrogen bond distances. Their catalytic role was here probed by biochemical and mutagenic studies. Moreover, on the basis of the secondary structure-driven multiple sequence alignment and the 3D structural model, a residue supposed important for catalysis, Gly84, was mutated to Ser. The activity of the mutated enzyme was drastically reduced. We propose that Gly84 is part of a putative "oxyanion hole" involved in the stabilization of the transition state similar to the C group of the esterase/lipase family.

Amino Acid Sequence↗

ESPript/ENDscript: Extracting and rendering sequence and 3D information from atomic structures of proteins.

The fortran program ESPript was created in 1993, to display on a PostScript figure multiple sequence alignments adorned with secondary structure elements. A web server was made available in 1999 and ESPript has been linked to three major web tools: ProDom which identifies protein domains, PredictProtein which predicts secondary structure elements and NPS@ which runs sequence alignment programs. A web server named ENDscript was created in 2002 to facilitate the generation of ESPript figures containing a large amount of information. ENDscript uses programs such as BLAST, Clustal and PHYLODENDRON to work on protein sequences and such as DSSP, CNS and MOLSCRIPT to work on protein coordinates. It enables the creation, from a single Protein Data Bank identifier, of a multiple sequence alignment figure adorned with secondary structure elements of each sequence of known 3D structure. Similar 3D structures are superimposed in turn with the program PROFIT and a final figure is drawn with BOBSCRIPT, which shows sequence and structure conservation along the Calpha trace of the query. ESPript and ENDscript are available at http://genopole.toulouse.inra.fr/ESPript.

DNA-Binding Proteins↗

CoSMoS: Conserved Sequence Motif Search in the proteome.

BACKGROUND: With the ever-increasing number of gene sequences in the public databases, generating and analyzing multiple sequence alignments becomes increasingly time consuming. Nevertheless it is a task performed on a regular basis by researchers in many labs. RESULTS: We have now created a database called CoSMoS to find the occurrences and at the same time evaluate the significance of sequence motifs and amino acids encoded in the whole genome of the model organism Escherichia coli K12. We provide a precomputed set of multiple sequence alignments for each individual E. coli protein with all of its homologues in the RefSeq database. The alignments themselves, information about the occurrence of sequence motifs together with information on the conservation of each of the more than 1.3 million amino acids encoded in the E. coli genome can be accessed via the web interface of CoSMoS. CONCLUSION: CoSMoS is a valuable tool to identify highly conserved sequence motifs, to find regions suitable for mutational studies in functional analyses and to predict important structural features in E. coli proteins.

Amino Acid Motifs↗

Evolution of duplications in the transferrin family of proteins.

The transferrin family is a group of proteins, defined by conserved amino acid motifs and putative function, found in both vertebrates and invertebrates. Included in this group are molecules known to bind iron, including serum transferrin, ovotransferrin, lactotransferrin, and melanotransferrin (MTF). Additional members of this family include inhibitor of carbonic anhydrase (ICA; mammals), major yolk protein (sea urchins), saxiphilin (frog), pacifastin (crayfish), and TTF-1 (algae). Most family members contain two lobes (N and C) of around 340 amino acids, the result of an ancient duplication event. In this article, we review the known functions of these proteins and speculate as to when the different homologs arose. From multiple-sequence alignments and neighbor-joining trees using 71 transferrin family sequences from 51 different species, including several novel sequences found in the Takifugu and Ciona genome databases, we conclude that melanotransferrins are much older (>670 MY) and more pervasive than previously thought, and the serum transferrin/melanotransferrin split may have occurred not long after lobe duplication. All subsequent duplication events diverged from the serum transferrin gene. The creation of such a large multiple-sequence alignment provides important information and could, in the future, highlight the role of specific residues in protein function.

Amino Acid Sequence↗

Efficient estimation of emission probabilities in profile hidden Markov models.

MOTIVATION: Profile hidden Markov models provide a sensitive method for performing sequence database search and aligning multiple sequences. One of the drawbacks of the hidden Markov model is that the conserved amino acids are not emphasized, but signal and noise are treated equally. For this reason, the number of estimated emission parameters is often enormous. Focusing the analysis on conserved residues only should increase the accuracy of sequence database search. RESULTS: We address this issue with a new method for efficient emission probability (EEP) estimation, in which amino acids are divided into effective and ineffective residues at each conserved alignment position. A practical study with 20 protein families demonstrated that the EEP method is capable of detecting family members from other proteins with sensitivity of 98% and specificity of 99% on the average, even if the number of free emission parameters was decreased to 15% of the original. In the database search for TIM barrel sequences, EEP recognizes the family members nearly as accurately as HMMER or Blast, but the number of false positive sequences was significantly less than that obtained with the other methods. AVAILABILITY: The algorithms written in C language are available on request from the authors.

Algorithms↗

Molecular dynamics and circular dichroism studies of human and rat C-peptides.

Proinsulin C-peptide has been recently described as an endogenous peptide hormone, responsible for important physiological functions others than its role in proinsulin processing. Accumulating evidences that C-peptide exerts beneficial effects in the treatment of long term complications of patients with type 1 diabetes mellitus indicate that this molecule may be administered together with insulin in future therapies. Despite its clear pharmacological interest, the secondary and three-dimensional (3D) structures of human C-peptide are still points of controversy. In the present work we report molecular dynamics (MD) simulations of human, rat I and rat II C-peptides. A common experimental strategy applied to all peptides consisted of homology building followed by multinanosecond MD simulations in vacuum and water. Circular dichroism (CD) experiments of each peptide in the absence and presence of 2,2,2-trifluoroethanol (TFE) were performed to support validation of the theoretical models. A multiple sequence alignment of 23 known mammalian C-peptides was constructed to identify significant conserved sites that would be important for the maintenance of secondary and tertiary structures. The analysis of the molecular dynamics trajectories for the human, rat I and rat II molecules have shown quite different general behavior, being the human C-peptide more flexible than the two others. Human and rat C-peptides exhibit very stable turn-like structures at the middle and C-terminal regions, which have been described as potential active sites of C-peptides. Human C-peptide also presented a short alpha-helix throughout the MD, which was not found in the rat molecules. CD data is in very good agreement with the MD results and both methods were able to identify a greater structural stability and potential in rat C-peptides when compared to the human C-peptide. The simulation results are discussed and validated in the light of multiple sequence alignment, recent experimental data from the literature and our own CD experiments.

Amino Acid Sequence↗

Prediction of cis/trans isomerization in proteins using PSI-BLAST profiles and secondary structure information.

BACKGROUND: The majority of peptide bonds in proteins are found to occur in the trans conformation. However, for proline residues, a considerable fraction of Prolyl peptide bonds adopt the cis form. Proline cis/trans isomerization is known to play a critical role in protein folding, splicing, cell signaling and transmembrane active transport. Accurate prediction of proline cis/trans isomerization in proteins would have many important applications towards the understanding of protein structure and function. RESULTS: In this paper, we propose a new approach to predict the proline cis/trans isomerization in proteins using support vector machine (SVM). The preliminary results indicated that using Radial Basis Function (RBF) kernels could lead to better prediction performance than that of polynomial and linear kernel functions. We used single sequence information of different local window sizes, amino acid compositions of different local sequences, multiple sequence alignment obtained from PSI-BLAST and the secondary structure information predicted by PSIPRED. We explored these different sequence encoding schemes in order to investigate their effects on the prediction performance. The training and testing of this approach was performed on a newly enlarged dataset of 2424 non-homologous proteins determined by X-Ray diffraction method using 5-fold cross-validation. Selecting the window size 11 provided the best performance for determining the proline cis/trans isomerization based on the single amino acid sequence. It was found that using multiple sequence alignments in the form of PSI-BLAST profiles could significantly improve the prediction performance, the prediction accuracy increased from 62.8% with single sequence to 69.8% and Matthews Correlation Coefficient (MCC) improved from 0.26 with single local sequence to 0.40. Furthermore, if coupled with the predicted secondary structure information by PSIPRED, our method yielded a prediction accuracy of 71.5% and MCC of 0.43, 9% and 0.17 higher than the accuracy achieved based on the singe sequence information, respectively. CONCLUSION: A new method has been developed to predict the proline cis/trans isomerization in proteins based on support vector machine, which used the single amino acid sequence with different local window sizes, the amino acid compositions of local sequence flanking centered proline residues, the position-specific scoring matrices (PSSMs) extracted by PSI-BLAST and the predicted secondary structures generated by PSIPRED. The successful application of SVM approach in this study reinforced that SVM is a powerful tool in predicting proline cis/trans isomerization in proteins and biological sequence analysis.

Databases, Protein↗

PANTHER: a browsable database of gene products organized by biological function, using curated protein family and subfamily classification.

The PANTHER database was designed for high-throughput analysis of protein sequences. One of the key features is a simplified ontology of protein function, which allows browsing of the database by biological functions. Biologist curators have associated the ontology terms with groups of protein sequences rather than individual sequences. Statistical models (Hidden Markov Models, or HMMs) are built from each of these groups. The advantage of this approach is that new sequences can be automatically classified as they become available. To ensure accurate functional classification, HMMs are constructed not only for families, but also for functionally distinct subfamilies. Multiple sequence alignments and phylogenetic trees, including curator-assigned information, are available for each family. The current version of the PANTHER database includes training sequences from all organisms in the GenBank non-redundant protein database, and the HMMs have been used to classify gene products across the entire genomes of human, and Drosophila melanogaster. The ontology terms and protein families and subfamilies, as well as Drosophila gene c;assifications, can be browsed and searched for free. Due to outstanding contractual obligations, access to human gene classifications and to protein family trees and multiple sequence alignments will temporarily require a nominal registration fee. PANTHER is publicly available on the web at http://panther.celera.com.

Animals↗

The relationship between the L1 and L2 domains of the insulin and epidermal growth factor receptors and leucine-rich repeat modules.

BACKGROUND: Leucine-rich repeats are one of the more common modules found in proteins. The leucine-rich repeat consensus motif is LxxLxLxxNxLxxLxxLxxLxx- where the first 11-12 residues are highly conserved and the remainder of the repeat can vary in size Leucine-rich repeat proteins have been subdivided into seven subfamilies, none of which include members of the epidermal growth factor receptor or insulin receptor families despite the similarity between the 3D structure of the L domains of the type I insulin-like growth factor receptor and some leucine-rich repeat proteins. RESULTS: Here we have used profile searches and multiple sequence alignments to identify the repeat motif Ixx-LxIxx-Nx-Lxx-Lxx-Lxx-Lxx- in the L1 and L2 domains of the insulin receptor and epidermal growth factor receptors. These analyses were aided by reference to the known three dimensional structures of the insulin-like growth factor type I receptor L domains and two members of the leucine rich repeat family, porcine ribonuclease inhibitor and internalin 1B. Pectate lyase, another beta helix protein, can also be seen to contain the sequence motif and much of the structural features characteristic of leucine-rich repeat proteins, despite the existence of major insertions in some of its repeats. CONCLUSION: Multiple sequence alignments and comparisons of the 3D structures has shown that right-handed beta helix proteins such as pectate lyase and the L domains of members of the insulin receptor and epidermal growth factor receptor families, are members of the leucine-rich repeat superfamily.

Amino Acid Sequence↗

The nucleoside-specific Tsx channel from the outer membrane of Salmonella typhimurium, Klebsiella pneumoniae and Enterobacter aerogenes: functional characterization and DNA sequence analysis of the tsx genes.

The Escherichia coli tsx gene encodes an integral outer-membrane protein (Tsx) that functions as a substrate-specific channel for deoxynucleosides and the antibiotic albicidin, and also serves as a receptor for bacteriophages and colicins. We cloned the structural genes of the Tsx proteins from Salmonella typhimurium, Klebsiella pneumoniae and Enterobacter aerogenes and expressed them in an E.coli tsx mutant. The heterologous Tsx proteins fully substituted the E.coli Tsx protein with respect to its function in deoxynucleoside and albicidin uptake, and as receptor for colicin K. The Tsx proteins from K. pneumoniae and Ent. aerogenes were also proficient as receptors for several Tsx-specific bacteriophages, whereas the corresponding protein from S. typhimurium did not confer sensitivity against these phages. The nucleotide sequence of the tsx genes from S. typhimurium, K. pneumoniae and Ent. aerogenes was established. Each of the Tsx proteins is initially synthesized with typical bacterial signal sequence peptides and the predicted mature forms of the Tsx proteins have a calculated M(r) of 30,567 (265 residues), 31,412 (272 residues) and 31,477 (272 residues), respectively. Multiple sequence alignments between the Tsx proteins showed a high degree of sequence identity and revealed the presence of four hypervariable regions, which are thought to constitute segments of the polypeptide chain exposed at the cell surface. Most notable was a deletion of 8 amino acids in one of these hypervariable domains in the S. typhimurium Tsx protein. When this deletion was introduced by site-directed mutagenesis into the corresponding region of the E.coli tsx gene, the mutant Tsx-515 protein lost its phage receptor function but still served as a colicin K receptor and as a substrate-specific channel, indicating that the region between residues 198 and 207 might be part of the bacteriophage receptor area. Multiple sequence alignments, structural predictions and the properties of previously characterized Tsx missense mutants were taken into account to develop a two-dimensional model for the topological organization of the Tsx protein within the outer membrane.

Amino Acid Sequence↗

Recurring local sequence motifs in proteins.

We describe a completely automated approach to identifying local sequence motifs that transcend protein family boundaries. Cluster analysis is used to identify recurring patterns of variation at single positions and in short segments of contiguous positions in multiple sequence alignments for a non-redundant set of protein families. Parallel experiments on simulated data sets constructed with the overall residue frequencies of proteins but not the inter-residue correlations show that naturally occurring protein sequences are significantly more clustered than the corresponding random sequences for window lengths ranging from one to 13 contiguous positions. The patterns of variation at single positions are not in general surprising: chemically similar amino acids tend to be grouped together. More interesting patterns emerge as the window length increases. The patterns of variation for longer window lengths are in part recognizable patterns of hydrophobic and hydrophilic residues, and in part less obvious combinations. A particularly interesting class of patterns features highly conserved glycine residues. The patterns provide a means to abstract the information contained in multiple sequence alignments and may be useful for comparison of distantly related sequences or sequence families and for protein structure prediction.

Cluster Analysis↗

Theory and practice of parallel direct optimization.

Our ability to collect and distribute genomic and other biological data is growing at a staggering rate (Pagel, 1999). However, the synthesis of these data into knowledge of evolution is incomplete. Phylogenetic systematics provides a unifying intellectual approach to understanding evolution but presents formidable computational challenges. A fundamental goal of systematics, the generation of evolutionary trees, is typically approached as two distinct NP-complete problems: multiple sequence alignment and phylogenetic tree search. The number of cells in a multiple alignment matrix are exponentially related to sequence length. In addition, the number of evolutionary trees expands combinatorially with respect to the number of organisms or sequences to be examined. Biologically interesting datasets are currently comprised of hundreds of taxa and thousands of nucleotides and morphological characters. This standard will continue to grow with the advent of highly automated sequencing and development of character databases. Three areas of innovation are changing how evolutionary computation can be addressed: (1) novel concepts for determination of sequence homology, (2) heuristics and shortcuts in tree-search algorithms, and (3) parallel computing. In this paper and the online software documentation we describe the basic usage of parallel direct optimization as implemented in the software POY (ftp://ftp.amnh.org/pub/molecular/poy).

Animals↗