PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Improving the sensitivity of the sequence profile method.

The sequence profile method (Gribskov M, McLachlan AD, Eisenberg D, 1987, Proc Natl Acad Sci USA 84:4355-4358) is a powerful tool to detect distant relationships between amino acid sequences. A profile is a table of position-specific scores and gap penalties, providing a generalized description of a protein motif, which can be used for sequence alignments and database searches instead of an individual sequence. A sequence profile is derived from a multiple sequence alignment. We have found 2 ways to improve the sensitivity of sequence profiles: (1) Sequence weights: Usage of individual weights for each sequence avoids bias toward closely related sequences. These weights are automatically assigned based on the distance of the sequences using a published procedure (Sibbald PR, Argos P, 1990, J Mol Biol 216:813-818). (2) Amino acid substitution table: In addition to the alignment, the construction of a profile also needs an amino acid substitution table. We have found that in some cases a new table, the BLOSUM45 table (Henikoff S, Henikoff JG, 1992, Proc Natl Acad Sci USA 89:10915-10919), is more sensitive than the original Dayhoff table or the modified Dayhoff table used in the current implementation. Profiles derived by the improved method are more sensitive and selective in a number of cases where previous methods have failed to completely separate true members from false positives.

Amino Acid Sequence↗

Introducing variable gap penalties to sequence alignment in linear space.

The problem of finding an optimal sequence alignment has been solved by Hirschberg (1975) in quadratic time and linear space. Myers and Miller (1988) presented an implementation of this algorithm for aligning biological sequences, incorporating affine gap penalties. The algorithm, has been essential in allowing progressive multiple sequence alignments to be performed on microcomputers with limited memory capacity. This paper presents a further development of the Myers and Miller algorithm. Here, we maximize similarity scores and, more significantly, introduce position-specific gap penalties. Thus, residue-dependent information such as structure preferences and existing gaps in a partial alignment can be applied to the solution of the alignment problem.

Algorithms↗

CDD: a database of conserved domain alignments with links to domain three-dimensional structure.

The Conserved Domain Database (CDD) is a compilation of multiple sequence alignments representing protein domains conserved in molecular evolution. It has been populated with alignment data from the public collections Pfam and SMART, as well as with contributions from colleagues at NCBI. The current version of CDD (v.1.54) contains 3693 such models. CDD alignments are linked to protein sequence and structure data in Entrez. The molecular structure viewer Cn3D serves as a tool to interactively visualize alignments and three-dimensional structure, and to link three-dimensional residue coordinates to descriptions of evolutionary conservation. CDD can be accessed on the World Wide Web at http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml. Protein query sequences may be compared against databases of position-specific score matrices derived from alignments in CDD, using a service named CD-Search, which can be found at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi. CD-Search runs reverse-position-specific BLAST (RPS-BLAST), a variant of the widely used PSI-BLAST algorithm. CD-Search is run by default for protein-protein queries submitted to NCBI's BLAST service at http://www.ncbi.nlm.nih.gov/BLAST.

Animals↗

Connexin 35: a gap-junctional protein expressed preferentially in the skate retina.

We have used low stringency hybridization to clone a novel connexin from a skate retinal cDNA library. A rat connexin 32 clone was used to isolate a single partial clone that was subsequently used to isolate seven more overlapping clones of the same cDNA. Two clones containing the entire open reading frame have a consensus sequence of 1456 bp and predict a protein of 302 amino acids length and molecular mass of 35,044 daltons, referred to as connexin 35 or Cx35. Southern blot analysis suggests that the cloned sequence lies in a single gene with one intron. Polymerase chain reaction amplification from genomic DNA and partial sequencing of this intron showed that it was approximately 950 bp in length, and located within the coding region 71 bp after the translation start site. Hydropathy analysis of the predicted protein and alignments with previously cloned connexins indicate that Cx35 has a long cytoplasmic loop and a relatively short carboxyl terminal tail. Multiple sequence alignments show that Cx35 has similarities to both alpha and beta groups of connexins and suggests that its origins may be near the divergence point for the two groups. Consensus sequences consistent with sites for phosphorylation by protein kinase C and by cAMP - or cGMP -dependent protein kinase were identified. Two transcripts were detected in Northern blot analysis: a 1.95-kb primary transcript and a 4.6-kb minor transcript. In RNA samples from 10 tissues, transcripts were detected only in the retina.

Amino Acid Sequence↗

Prediction of transmembrane alpha-helices in prokaryotic membrane proteins: the dense alignment surface method.

A new, simple method for predicting transmembrane segments in integral membrane proteins has been developed. It is based on low-stringency dot-plots of the query sequence against a collection of non-homologous membrane proteins using a previously derived scoring matrix [Cserzö et al., 1994, J. Mol. Biol., 243, 388-396]. This so-called dense alignment surface (DAS) method is shown to perform on par with earlier methods that require extra information in the form of multiple sequence alignments or the distribution of positively charged residues outside the transmembrane segments, and thus improves prediction abilities when only single-sequence information is available or for classes of membrane proteins that do not follow the 'positive inside' rule.

Cell Membrane↗

Application of genetic semihomology algorithm to theoretical studies on various protein families.

Several protein families of different nature were studied for genetic relationship, correct alignment at non-homologous fragments, optimal sequence consensus construction, and confirmation of their actual relevance. A comparison of the genetic semihomology approach with statistical approaches indicates a high accuracy and cognition significance of the former. This is particularly pronounced in the study of related proteins that show a low degree of homology. The sequence multiple alignments were verified and corrected with respect to the questionable, non-homologous fragments. The verified alignments were the basis for consensus sequence formation. The frequency of six-codon amino acids occurrence versus position variability was studied and their possible role in amino acid mutational exchange at variable positions is discussed.

Algorithms↗

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans↗

A workbench for large-scale sequence homology analysis.

When routinely analysing very long stretches of DNA sequences produced by genome sequencing projects, detailed analysis of database search results becomes exceedingly time consuming. To reduce the tedious browsing of large quantities of protein similarities, two programs, MSPcrunch and Blixem, were developed, which assist in processing the results from the database search programs in the BLAST suite. MSPcrunch removes biased composition and redundant matches while keeping weak matches that are consistent with a larger gapped alignment. This makes BLAST searching in practice more sensitive and reduces the risk of overlooking distant similarities. Blixem is a multiple sequence alignment viewer for X-windows which makes it significantly easier to scan and evaluate the matches ratified by MSPcrunch. In Blixem, matches to the translated DNA query sequence are simultaneously aligned in three frames. Also, the distribution of matches over the whole DNA query is displayed. Examples of usage are drawn from 36 C. elegans cosmid clones totalling 1.2 megabases, to which these tools were applied.

Algorithms↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

Sensitive pattern discovery with 'fuzzy' alignments of distantly related proteins.

MOTIVATION: Evolutionary comparison leads to efficient functional characterisation of hypothetical proteins. Here, our goal is to map specific sequence patterns to putative functional classes. The evolutionary signal stands out most clearly in a maximally diverse set of homologues. This diversity, however, leads to a number of technical difficulties. The targeted patterns-as gleaned from structure comparisons-are too sparse for statistically significant signals of sequence similarity and accurate multiple sequence alignment. RESULTS: We address this problem by a fuzzy alignment model, which probabilistically assigns residues to structurally equivalent positions (attributes) of the proteins. We then apply multivariate analysis to the 'attributes x proteins' matrix. The dimensionality of the space is reduced using non-negative matrix factorization. The method is general, fully automatic and works without assumptions about pattern density, minimum support, explicit multiple alignments, phylogenetic trees, etc. We demonstrate the discovery of biologically meaningful patterns in an extremely diverse superfamily related to urease.

Algorithms↗

Multiple sequence information for threading algorithms.

Threading algorithms attempt to solve the inverse protein folding problem: given a group of structures and a sequence, identify the structure that is most compatible with this sequence. A recent study of this class of algorithms by S. J. Wodak and colleagues suggests that while threading algorithms are capable of recognizing many folding motifs, their performance in truly blind predictions is disappointing, and the underlying alignments upon which the selections are based are frequently errant. To help overcome this problem we have developed a Test of Optimal Mutagenesis algorithm (TOM) that exploits information inherent in the variation between several homologues in a multiple sequence alignment. This information is used to help select the correct structural motif for the sequence from a database of known structures. A total of 305 high-resolution structures were selected to represent the set of known folds; 56 proteins were chosen that had at least one close structural match in this set. To test TOM, we attempted to determine which of the 305 folds was a match to each of the 56 protein sequences. TOM correctly predicts a close structural match for 45% of these proteins. THREADER, an algorithm chosen as a literature standard, correctly matched 20% of the test set. By comparing the performance of TOM, THREADER, and TOM NOVAR (a version of TOM without variability information), we conclude that the tendency of an amino acid to be buried or exposed is the dominant determinant of the success of threading algorithms. In addition, the structural alignments produced by TOM suggest that the exact alignment of just 30 to 50% of the residues in a sequence with the correct fold is necessary to select it as the highest scoring match in a set of folds.

Algorithms↗

A simple method for aligning many protein sequences.

A simple extension of the Needleman and Wunsch algorithm for aligning pairs of protein sequences allows it to be used for the efficient generation of very large multiple-sequence alignments whose members are similar. This technique could have applications in a broad range of high-volume genomics projects.

Algorithms↗

Membrane-associated proteins in eicosanoid and glutathione metabolism (MAPEG). A widespread protein superfamily.

The members of the MAPEG superfamily have been aligned and found to be distantly related, with a common pattern of hydropathy. Figure 2A shows the multiple sequence alignments of the human members and Figure 2B the corresponding superimposed hydropathy profiles. The alignment in Figure 2A demonstrates a total of six strictly conserved residues. The Arg-51 in LTC4 synthase has been suggested to function as proton donor for the opening of the LTA4 epoxide. This arginine is found in all but the FLAP sequences in accordance with the observation that FLAP has no known enzyme activity. Also the Tyr-93 in LTC4 synthase has been suggested to function as a base for the formation of the thiolate anion of glutathione. This tyrosine is not conserved in MGST1 or MGST1-L1. Table 1 summarizes some other properties of the individual human proteins. They are all of the same size, ranging from 147 to 161 amino acids. Only FLAP differs in that its isoelectric point is more neutral than that of the other, more basic proteins. The genes encoding these proteins all reside on different chromosomes (when known) (Table 1). In addition to the human proteins, MAPEG members have been identified in plants, fungi, and bacteria. It is clearly a challenge to elucidate their role in these different phyla in relation to their defined physiological functions in humans.

Amino Acid Sequence↗

Detecting recombination with MCMC.

MOTIVATION: We present a statistical method for detecting recombination, whose objective is to accurately locate the recombinant breakpoints in DNA sequence alignments of small numbers of taxa (4 or 5). Our approach explicitly models the sequence of phylogenetic tree topologies along a multiple sequence alignment. Inference under this model is done in a Bayesian way, using Markov chain Monte Carlo (MCMC). The algorithm returns the site-dependent posterior probability of each tree topology, which is used for detecting recombinant regions and locating their breakpoints. RESULTS: The method was tested on a synthetic and three real DNA sequence alignments, where it was found to outperform the established detection methods PLATO, RECPARS, and TOPAL.

Algorithms↗

Suboptimal sequence alignment in molecular biology. Alignment with error analysis.

A molecular sequence alignment algorithm based on dynamic programming has been extended to allow the computation of all pairs of residues that can be part of optimal and suboptimal sequence alignments. The uncertainties inherent in sequence alignment can be displayed using a new form of dot plot. The method allows the qualitative assessment of whether or not two sequences are related, and can reveal what parts of the alignment are better determined than others. It also permits the computation of representative optimal and suboptimal alignments. The relation between alignment reliability and alignment parameters is discussed. Other applications are to cyclical permutations of sequences and the detection of self-similarity. An application to multiple sequence alignment is noted.

Algorithms↗

Conservation analysis and structure prediction of the protein serine/threonine phosphatases. Sequence similarity with diadenosine tetraphosphatase from Escherichia coli suggests homology to the protein phosphatases.

A multiple sequence alignment of 44 serine/threonine-specific protein phosphatases has been performed. This reveals the position of a common conserved catalytic core, the location of invariant residues, insertions and deletions. The multiple alignment has been used to guide and improve a consensus secondary-structure prediction for the common catalytic core. The location of insertions and deletions has aided in defining the positions of surface loops and turns. The prediction suggests that the core protein phosphatase structure comprises two domains: the first has a single, beta sheet flanked by alpha helices, while the second is predominantly alpha helical. Knowledge of the core secondary structures provides a guide for the design of site-directed-mutagenesis experiments that will not disrupt the native phosphatase fold. A sequence similarity between eukaryotic serine/threonine protein phosphatases and the Escherichia coli diadenosine tetraphosphatase has been identified. This extends over the N-terminal 100 residues of bacteriophage phosphatases and E. coli diadenosine tetraphosphatase. Residues which are invariant amongst these classes are likely to be important in catalysis and protein folding. These include Arg92, Asn138, Asp59, Asp88, Gly58, Gly62, Gly87, Gly93, Gly137, His61, His139 and Val90 and fall into three clusters with the consensus sequences GD(IVTL)HG, GD(LYF)V(DA)RG and GNH, where brackets surround alternative amino acids. The first two consensus sequences are predicted to fall in the beta-alpha and beta-beta loops of a beta-alpha-beta-beta secondary-structure motif. This places the predicted phosphate-binding site at the N-terminus of the alpha helix, where phosphate binding may be stabilised by the alpha-helix dipole.

Acid Anhydride Hydrolases↗

New features of the Blocks Database servers.

Blocks are ungapped multiple sequence alignments representing conserved protein regions, and the Blocks Database consists of blocks from documented protein families. World Wide Web (http://www. blocks.fhcrc.org) and Email (blocks@blocks.fhcrc.org) servers provide tools for homology searching and for analyzing protein family relationships. New enhancements include a multiple alignment processor that extends the use of these tools to imported multiple alignments of families not present in the database and a PCR primer designer that implements a new strategy for gene isolation.

DNA Primers↗