PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

MUSTANG: a multiple structural alignment algorithm.

Multiple structural alignment is a fundamental problem in structural genomics. In this article, we define a reliable and robust algorithm, MUSTANG (MUltiple STructural AligNment AlGorithm), for the alignment of multiple protein structures. Given a set of protein structures, the program constructs a multiple alignment using the spatial information of the C(alpha) atoms in the set. Broadly based on the progressive pairwise heuristic, this algorithm gains accuracy through novel and effective refinement phases. MUSTANG reports the multiple sequence alignment and the corresponding superposition of structures. Alignments generated by MUSTANG are compared with several handcurated alignments in the literature as well as with the benchmark alignments of 1033 alignment families from the HOMSTRAD database. The performance of MUSTANG was compared with DALI at a pairwise level, and with other multiple structural alignment tools such as POSA, CE-MC, MALECON, and MultiProt. MUSTANG performs comparably to popular pairwise and multiple structural alignment tools for closely related proteins, and performs more reliably than other multiple structural alignment methods on hard data sets containing distantly related proteins or proteins that show conformational changes.

Algorithms↗

Human-specific insertions and deletions inferred from mammalian genome sequences.

It has been suggested that insertions and deletions (indels) have contributed to the sequence divergence between the human and chimpanzee genomes more than do nucleotide changes (3% vs. 1.2%). However, although there have been studies of large indels between the two genomes, no systematic analysis of small indels (i.e., indels </= 100 bp) has been published. In this study, we first estimated that the false-positive rate of small indels inferred from human-chimpanzee pairwise sequence alignments is quite high, suggesting that the chimpanzee genome draft is not sufficiently accurate for our purpose. We have therefore inferred only human-specific indels using multiple sequence alignments of mammalian genomes. We identified >840,000 "small" indels, which affect >7000 UCSC-annotated human genes (>11,000 transcripts). These indels, however, amount to only approximately 0.21% sequence change in the human lineage for the regions compared, whereas in pseudogenes indels contribute to a sequence divergence of 1.40%, suggesting that most of the indels that occurred in genic regions have been eliminated. Functional analysis reveals that the genes whose coding exons have been affected by human-specific indels are enriched in transcription and translation regulatory activities but are underrepresented in catalytic and transporter activities, cellular and physiological processes, and extracellular region/matrix. This functional bias suggests that human-specific indels might have contributed to human unique traits by causing changes at the RNA and protein level.

Animals↗

Improving the sensitivity of the sequence profile method.

The sequence profile method (Gribskov M, McLachlan AD, Eisenberg D, 1987, Proc Natl Acad Sci USA 84:4355-4358) is a powerful tool to detect distant relationships between amino acid sequences. A profile is a table of position-specific scores and gap penalties, providing a generalized description of a protein motif, which can be used for sequence alignments and database searches instead of an individual sequence. A sequence profile is derived from a multiple sequence alignment. We have found 2 ways to improve the sensitivity of sequence profiles: (1) Sequence weights: Usage of individual weights for each sequence avoids bias toward closely related sequences. These weights are automatically assigned based on the distance of the sequences using a published procedure (Sibbald PR, Argos P, 1990, J Mol Biol 216:813-818). (2) Amino acid substitution table: In addition to the alignment, the construction of a profile also needs an amino acid substitution table. We have found that in some cases a new table, the BLOSUM45 table (Henikoff S, Henikoff JG, 1992, Proc Natl Acad Sci USA 89:10915-10919), is more sensitive than the original Dayhoff table or the modified Dayhoff table used in the current implementation. Profiles derived by the improved method are more sensitive and selective in a number of cases where previous methods have failed to completely separate true members from false positives.

Amino Acid Sequence↗

Introducing variable gap penalties to sequence alignment in linear space.

The problem of finding an optimal sequence alignment has been solved by Hirschberg (1975) in quadratic time and linear space. Myers and Miller (1988) presented an implementation of this algorithm for aligning biological sequences, incorporating affine gap penalties. The algorithm, has been essential in allowing progressive multiple sequence alignments to be performed on microcomputers with limited memory capacity. This paper presents a further development of the Myers and Miller algorithm. Here, we maximize similarity scores and, more significantly, introduce position-specific gap penalties. Thus, residue-dependent information such as structure preferences and existing gaps in a partial alignment can be applied to the solution of the alignment problem.

Algorithms↗

The PRALINE online server: optimising progressive multiple alignment on the web.

We introduce the online server for PRALINE (http://ibium.cs.vu.nl/programs/pralinewww/), an iterative versatile progressive multiple sequence alignment (MSA) tool. PRALINE provides various MSA optimisation strategies including weighted global and local profile pre-processing, secondary structure-guided alignment and a reliability measure for aligned individual residue positions. The latter can also be used to optimise the alignment when the profile pre-processing strategies are iterated. In addition, we have modelled the server output to enable comprehensive visualisation of the generated alignment and easy figure generation for publications. The alignment is represented in five default colour schemes based on: residue type, position conservation, position reliability, residue hydrophobicity and secondary structure; depending on the options set. We have also implemented a custom colour scheme that allows the user to select which colour will represent one or more amino acids in the alignment. The grouping of sequences, on which the alignment is based, can also be visualised as a dendrogram. The PRALINE algorithm is designed to work more as a toolkit for MSA rather than a one step process.

Internet↗

CDD: a database of conserved domain alignments with links to domain three-dimensional structure.

The Conserved Domain Database (CDD) is a compilation of multiple sequence alignments representing protein domains conserved in molecular evolution. It has been populated with alignment data from the public collections Pfam and SMART, as well as with contributions from colleagues at NCBI. The current version of CDD (v.1.54) contains 3693 such models. CDD alignments are linked to protein sequence and structure data in Entrez. The molecular structure viewer Cn3D serves as a tool to interactively visualize alignments and three-dimensional structure, and to link three-dimensional residue coordinates to descriptions of evolutionary conservation. CDD can be accessed on the World Wide Web at http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml. Protein query sequences may be compared against databases of position-specific score matrices derived from alignments in CDD, using a service named CD-Search, which can be found at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi. CD-Search runs reverse-position-specific BLAST (RPS-BLAST), a variant of the widely used PSI-BLAST algorithm. CD-Search is run by default for protein-protein queries submitted to NCBI's BLAST service at http://www.ncbi.nlm.nih.gov/BLAST.

Animals↗

Connexin 35: a gap-junctional protein expressed preferentially in the skate retina.

We have used low stringency hybridization to clone a novel connexin from a skate retinal cDNA library. A rat connexin 32 clone was used to isolate a single partial clone that was subsequently used to isolate seven more overlapping clones of the same cDNA. Two clones containing the entire open reading frame have a consensus sequence of 1456 bp and predict a protein of 302 amino acids length and molecular mass of 35,044 daltons, referred to as connexin 35 or Cx35. Southern blot analysis suggests that the cloned sequence lies in a single gene with one intron. Polymerase chain reaction amplification from genomic DNA and partial sequencing of this intron showed that it was approximately 950 bp in length, and located within the coding region 71 bp after the translation start site. Hydropathy analysis of the predicted protein and alignments with previously cloned connexins indicate that Cx35 has a long cytoplasmic loop and a relatively short carboxyl terminal tail. Multiple sequence alignments show that Cx35 has similarities to both alpha and beta groups of connexins and suggests that its origins may be near the divergence point for the two groups. Consensus sequences consistent with sites for phosphorylation by protein kinase C and by cAMP - or cGMP -dependent protein kinase were identified. Two transcripts were detected in Northern blot analysis: a 1.95-kb primary transcript and a 4.6-kb minor transcript. In RNA samples from 10 tissues, transcripts were detected only in the retina.

Amino Acid Sequence↗

Prediction of transmembrane alpha-helices in prokaryotic membrane proteins: the dense alignment surface method.

A new, simple method for predicting transmembrane segments in integral membrane proteins has been developed. It is based on low-stringency dot-plots of the query sequence against a collection of non-homologous membrane proteins using a previously derived scoring matrix [Cserzö et al., 1994, J. Mol. Biol., 243, 388-396]. This so-called dense alignment surface (DAS) method is shown to perform on par with earlier methods that require extra information in the form of multiple sequence alignments or the distribution of positively charged residues outside the transmembrane segments, and thus improves prediction abilities when only single-sequence information is available or for classes of membrane proteins that do not follow the 'positive inside' rule.

Cell Membrane↗

Homology assessment and molecular sequence alignment.

Hypotheses of homology are the basis of phylogenetic analysis. All character data are considered to be equivalent regardless of the source of those characters. Putative homology statements are designated based on observations of similarity. Pairwise sequence alignment using the Needleman-Wunsch algorithm is the basis for similarity maximization between molecular sequences. Multiple sequence alignment uses this algorithm in a topologically hierarchical framework. The resulting hypotheses of homology are tested in conjunction with character congruence through parsimony. This review introduces some underlying principles of phylogenetic analysis as they pertain homology testing and DNA sequence alignment.

Algorithms↗

Application of genetic semihomology algorithm to theoretical studies on various protein families.

Several protein families of different nature were studied for genetic relationship, correct alignment at non-homologous fragments, optimal sequence consensus construction, and confirmation of their actual relevance. A comparison of the genetic semihomology approach with statistical approaches indicates a high accuracy and cognition significance of the former. This is particularly pronounced in the study of related proteins that show a low degree of homology. The sequence multiple alignments were verified and corrected with respect to the questionable, non-homologous fragments. The verified alignments were the basis for consensus sequence formation. The frequency of six-codon amino acids occurrence versus position variability was studied and their possible role in amino acid mutational exchange at variable positions is discussed.

Algorithms↗

A secondary structural model of the 28S rRNA expansion segments D2 and D3 for Chalcidoid wasps (Hymenoptera: Chalcidoidea).

We analyze the secondary structure of two expansion segments (D2, D3) of the 28S ribosomal (rRNA)-encoding gene region from 527 chalcidoid wasp taxa (Hymenoptera: Chalcidoidea) representing 18 of the 19 extant families. The sequences are compared in a multiple sequence alignment, with secondary structure inferred primarily from the evidence of compensatory base changes in conserved helices of the rRNA molecules. This covariation analysis yielded 36 helices that are composed of base pairs exhibiting positional covariation. Several additional regions are also involved in hydrogen bonding, and they form highly variable base-pairing patterns across the alignment. These are identified as regions of expansion and contraction or regions of slipped-strand compensation. Additionally, 31 single-stranded locales are characterized as regions of ambiguous alignment based on the difficulty in assigning positional homology in the presence of multiple adjacent indels. Based on comparative analysis of these sequences, the largest genetic study on any hymenopteran group to date, we report an annotated secondary structural model for the D2, D3 expansion segments that will prove useful in assigning positional nucleotide homology for phylogeny reconstruction in these and closely related apocritan taxa.

Animals↗

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans↗

FusionDB: a database for in-depth analysis of prokaryotic gene fusion events.

FusionDB (http://igs-server.cnrs-mrs.fr/FusionDB/) constitutes a resource dedicated to in-depth analysis of bacterial and archaeal gene fusion events. Such events can provide the 'Rosetta stone' in the search for potential protein-protein interactions, as well as metabolic and regulatory networks. However, the false positive rate of this approach may be quite high, prompting a detailed scrutiny of putative gene fusion events. FusionDB readily provides much of the information required for that task. Moreover, FusionDB extends the notion of gene fusion from that of a single gene to that of a family of genes by assembling pairs of genes from different genomes that belong to the same Cluster of Orthogonal Groups (COG). Multiple sequence alignments and phylogenetic tree reconstruction for the N- and C-terminal parts of these 'COG fusion' events are provided to distinguish single and multiple fusion events from cases of gene fission, pseudogenes and other false positives. Finally, gene fusion events with matches to known structures of heterodimers in the Protein Data Bank (PDB) are identified and may be visualized. FusionDB is fully searchable with access to sequence and alignment data at all levels. A number of different scores are provided to easily differentiate 'real' from 'questionable' cases, especially when larger database searches are performed. FusionDB is cross-linked with the 'Phylogenomic Display of Bacterial Genes' (PhydBac) online web server. Together, these servers provide the complete set of information required for in-depth analysis of non-homology-based gene function attribution.

Artificial Gene Fusion↗

A workbench for large-scale sequence homology analysis.

When routinely analysing very long stretches of DNA sequences produced by genome sequencing projects, detailed analysis of database search results becomes exceedingly time consuming. To reduce the tedious browsing of large quantities of protein similarities, two programs, MSPcrunch and Blixem, were developed, which assist in processing the results from the database search programs in the BLAST suite. MSPcrunch removes biased composition and redundant matches while keeping weak matches that are consistent with a larger gapped alignment. This makes BLAST searching in practice more sensitive and reduces the risk of overlooking distant similarities. Blixem is a multiple sequence alignment viewer for X-windows which makes it significantly easier to scan and evaluate the matches ratified by MSPcrunch. In Blixem, matches to the translated DNA query sequence are simultaneously aligned in three frames. Also, the distribution of matches over the whole DNA query is displayed. Examples of usage are drawn from 36 C. elegans cosmid clones totalling 1.2 megabases, to which these tools were applied.

Algorithms↗

PCOAT: positional correlation analysis using multiple methods.

UNLABELLED: PCOAT (Positional COrrelation Analysis Tool) is a program to perform positional correlation analysis for protein multiple sequence alignment in order to identify structurally or functionally important interactions between positions in a protein family. We implement different statistical methods to detect highly correlated position pairs, amino acid pairs, individual positions and networks of correlated positions, and utilize multiple sequence weighting and sampling methods to eliminate background correlations caused by phylogeny and stochastic events. Our program runs relatively fast and is suitable for analyzing alignments containing large number of sequences. AVAILABILITY: ftp://iole.swmed.edu/pub/PCOAT/. SUPPLEMENTARY INFORMATION: The PCOAT ftp site contains a detailed description of the program, and the results of PCOAT analysis on C2H2 alignment and ACT domain alignment.

Algorithms↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

Sensitive pattern discovery with 'fuzzy' alignments of distantly related proteins.

MOTIVATION: Evolutionary comparison leads to efficient functional characterisation of hypothetical proteins. Here, our goal is to map specific sequence patterns to putative functional classes. The evolutionary signal stands out most clearly in a maximally diverse set of homologues. This diversity, however, leads to a number of technical difficulties. The targeted patterns-as gleaned from structure comparisons-are too sparse for statistically significant signals of sequence similarity and accurate multiple sequence alignment. RESULTS: We address this problem by a fuzzy alignment model, which probabilistically assigns residues to structurally equivalent positions (attributes) of the proteins. We then apply multivariate analysis to the 'attributes x proteins' matrix. The dimensionality of the space is reduced using non-negative matrix factorization. The method is general, fully automatic and works without assumptions about pattern density, minimum support, explicit multiple alignments, phylogenetic trees, etc. We demonstrate the discovery of biologically meaningful patterns in an extremely diverse superfamily related to urease.

Algorithms↗

Multiple sequence information for threading algorithms.

Threading algorithms attempt to solve the inverse protein folding problem: given a group of structures and a sequence, identify the structure that is most compatible with this sequence. A recent study of this class of algorithms by S. J. Wodak and colleagues suggests that while threading algorithms are capable of recognizing many folding motifs, their performance in truly blind predictions is disappointing, and the underlying alignments upon which the selections are based are frequently errant. To help overcome this problem we have developed a Test of Optimal Mutagenesis algorithm (TOM) that exploits information inherent in the variation between several homologues in a multiple sequence alignment. This information is used to help select the correct structural motif for the sequence from a database of known structures. A total of 305 high-resolution structures were selected to represent the set of known folds; 56 proteins were chosen that had at least one close structural match in this set. To test TOM, we attempted to determine which of the 305 folds was a match to each of the 56 protein sequences. TOM correctly predicts a close structural match for 45% of these proteins. THREADER, an algorithm chosen as a literature standard, correctly matched 20% of the test set. By comparing the performance of TOM, THREADER, and TOM NOVAR (a version of TOM without variability information), we conclude that the tendency of an amino acid to be buried or exposed is the dominant determinant of the success of threading algorithms. In addition, the structural alignments produced by TOM suggest that the exact alignment of just 30 to 50% of the residues in a sequence with the correct fold is necessary to select it as the highest scoring match in a set of folds.

Algorithms↗