PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Spreadsheet macros for coloring sequence alignments.

This article describes a set of Microsoft Excel macros designed to color amino acid and nucleotide sequence alignments for review and preparation of visual aids. The colored alignments can then be modified to emphasize features of interest. Procedures for importing and coloring sequences are described. The macro file adds a new menu to the menu bar containing sequence-related commands to enable users unfamiliar with Excel to use the macros more readily. The macros were designed for use with Macintosh computers but will also run with the DOS version of Excel.

Amino Acid Sequence↗

Sequence alignments in the neighborhood of the optimum with general application to dynamic programming.

When applying dynamic programming techniques to obtain optimal sequence alignments, a set of weights must be assigned to mismatches, insertion/deletions, etc. These weights are not predetermined, although efforts are being made to deduce biologically meaningful values from data. In addition, there are sometimes unknown constraints on the sequences that cause the "true" alignment to disagree with the optimum (computer) solution. To assist in overcoming these difficulties, an algorithm has been developed to produce all alignments within a specified distance of the optimum. The distance can be chosen after the optimum is computed, and the algorithm can be repeated at will. Earlier algorithms to solve this problem were very complex and not practical for any case involving sequences with significant time or storage requirements. The algorithm presented here overcomes these difficulties and has application to general, discrete dynamic programming problems.

Journal Article↗

Tertiary structure prediction of the KIX domain of CBP using Monte Carlo simulations driven by restraints derived from multiple sequence alignments.

Using a recently developed protein folding algorithm, a prediction of the tertiary structure of the KIX domain of the CREB binding protein is described. The method incorporates predicted secondary and tertiary restraints derived from multiple sequence alignments in a reduced protein model whose conformational space is explored by Monte Carlo dynamics. Secondary structure restraints are provided by the PHD secondary structure prediction algorithm that was modified for the presence of predicted U-turns, i.e., regions where the chain reverses global direction. Tertiary restraints are obtained via a two-step process: First, seed side-chain contacts are identified from a correlated mutation analysis, and then, a threading-based algorithm expands the number of these seed contacts. Blind predictions indicate that the KIX domain is a putative three-helix bundle, although the chirality of the bundle could not be uniquely determined. The expected root-mean-square deviation for the correct chirality of the KIX domain is between 5.0 and 6.2 A. This is to be compared with the estimate of 12.9 A that would be expected by a random prediction, using the model of F. Cohen and M. Sternberg (J. Mol. Biol. 138:321-333, 1980).

Algorithms↗

Mix'n'Match: an improved multiple sequence alignment procedure for distantly related proteins using secondary structure predictions, designed to be independent of the choice of gap penalty and scoring matrix.

A new multiple sequence alignment procedure is presented. Several different multiple alignments are made using differing criteria. Having divided the sequences into strongly conserved regions (SCRs) and loosely conserved regions (LCRs), the 'best' alignment for each LCR is chosen, independently of the other LCRs, from a selection of possibilities in the multiple alignments. To help make this choice for each LCR, the secondary structure is predicted and shown alongside each different possible alignment. One advantage of this method over automatic, non-interactive methods, is that the final alignment is not dependent on the choice of a single set of scoring parameters. Another is that, by allowing interactive choice and by taking account of secondary structural information, the final alignment is based more on biological rather than mathematical factors. This method can produce better alignments than any of the initial automatic multiple alignment methods used.

Amino Acid Sequence↗

Comparative analysis of seven multiple protein sequence alignment servers: clues to enhance reliability of predictions.

MOTIVATION: The prediction reliability of seven multiple alignment servers currently available on the Internet (ClustalW, MAP, PIMA, Block Maker, MSA, MEME and Match-Box) has been evaluated in terms of power (sensitivity) and confidence (selectivity). Therefore, the alignments obtained have been respectively compared to refined structural alignments for 20 families of related proteins with low levels of identity. RESULTS: Results clearly show that any powerful method remains reliable when the rate of identity falls. For some methods, power and confidence decrease linearly with the rate of identity, while other methods emphasize reliability at the cost of a lower power. Increasing the number of related sequences included in the alignment may either improve or decrease the quality of the predictions substantially. For some methods, the gain in power or in confidence is quite systematic; for others, the effect of the addition of homologous sequences is highly unpredictable. Extracting the consensus between two different methods may increase the overall confidence of the predictions tremendously. Our conclusions induce users of sequence alignment methods on the Internet to select the most suitable technique according to their requirements in terms of selectivity and sensitivity. AVAILABILITY: The aligned sequences of the 20 alignments of structure can be obtained automatically by sending the message 'send: cabios_tests.txt' by e-mail to 'matchbox@biq.fundp.ac.be'. CONTACT: eric.depiereux@fundp.ac.be

Computer Communication Networks↗

A kinase sequence database: sequence alignments and family assignment.

UNLABELLED: The Kinase Sequence Database (KSD) located at http://kinase.ucsf.edu/ksd contains information on 290 protein kinase families derived by profile-based clustering of the non-redundant list of sequences obtained from a GenBank-wide search. Included in the database are a total of 5,041 protein kinases from over 100 organisms. Clustering into families is based on the extent of homology within the kinase catalytic domain (250-300 residues in length). Alignments of the families are viewed by interactive Excel-based sequence spreadsheets. In addition, KSD features evolutionary trees derived for each family and detailed information on each sequence as well as links to the corresponding GenBank entries. Sequence manipulation tools, such as evolutionary tree generation, novel sequence assignment, and statistical analysis, are also provided. AVAILABILITY: The kinase sequence database is a web-based service accessible at http://kinase.ucsf.edu/ksd CONTACT: buzko@cmp.ucsf.edu; shokat@cmp.ucsf.edu/ksd

Cluster Analysis↗

Consistency of optimal sequence alignments.

Pairwise optimal alignments between three or more sequences are not necessarily consistent as a whole, but consistent and inconsistent residues are usually distributed in clusters. An efficient method has been developed for locating consistent regions when each pairwise alignment is given in the form of a "skeletal representation" (Bull. math. Biol. 52, 359-373). This method is further extended so that the combination of pairwise alignments that gives the greatest consistency is found when possibly many alignments are equally optimal for each pairwise comparison. A method for acceleration of simultaneous multiple sequence alignment is proposed in which consistent regions serve as "anchor points" limiting application of direct multi-way alignment to the rest of "inconsistent" regions.

Algorithms↗

A phylogenetic study of the Anopheles punctulatus group of malaria vectors comparing rDNA sequence alignments derived from the mitochondrial and nuclear small ribosomal subunits.

A phylogenetic study of the members of the Anopheles punctulatus group was performed using structural and similarity-based DNA sequence alignments of the small ribosomal subunit (SSU) from both the nuclear and the mitochondrial genomes. The mitochondrial SSU gene (12S, approximately 650 bp) proved to be highly restricted by its secondary structure and displayed little informative sequence variation. Consequently, it was considered unsuitable for a phylogenetic study of these closely related mosquito species. A structural alignment of the nuclear ribosomal DNA SSU (18S, approximately 2000 bp) proved to be more informative than similarity-based alignments. Analyses showed the A. punctulatus group to be monophyletic with two major clades; a Farauti clade containing members displaying an all-black-scaled proboscis (A. farauti 1-3 and 5-7) and the Punctulatus clade containing members displaying extensive white scaling on the apical half of the proboscis (A. farauti 4, A. punctulatus, and An. sp. near punctulatus). Anopheles koliensis was positioned basal to the Farauti clade.

Animals↗

DNA binding properties in vivo and target recognition domain sequence alignment analyses of wild-type and mutant RsrI [N6-adenine] DNA methyltransferases.

A genetic selection method, the P22 challenge-phage assay, was used to characterize DNA binding in vivo by the prokaryotic beta class [N:6-adenine] DNA methyltransferase M.RSR:I. M.RSR:I mutants with altered binding affinities in vivo were isolated. Unlike the wild-type enzyme, a catalytically compromised mutant, M.RSR:I (L72P), demonstrated site-specific DNA binding in vivo. The L72P mutation is located near the highly conserved catalytic motif IV, DPPY (residues 65-68). A double mutant, M.RSR:I (L72P/D173A), showed less binding in vivo than did M.RSR:I (L72P). Thus, introduction of the D173A mutation deleteriously affected DNA binding. D173 is located in the putative target recognition domain (TRD) of the enzyme. Sequence alignment analyses of several beta class MTases revealed a TRD sequence element that contains the D173 residue. Phylogenetic analysis suggested that divergence in the amino acid sequences of these methyltransferases correlated with differences in their DNA target recognition sequences. Furthermore, MTases of other classes (alpha and gamma) having the same DNA recognition sequence as the beta class MTases share related regions of amino acid sequences in their TRDs.

Adenine↗

On the accuracy of homology modeling and sequence alignment methods applied to membrane proteins.

In this study, we investigate the extent to which techniques for homology modeling that were developed for water-soluble proteins are appropriate for membrane proteins as well. To this end we present an assessment of current strategies for homology modeling of membrane proteins and introduce a benchmark data set of homologous membrane protein structures, called HOMEP. First, we use HOMEP to reveal the relationship between sequence identity and structural similarity in membrane proteins. This analysis indicates that homology modeling is at least as applicable to membrane proteins as it is to water-soluble proteins and that acceptable models (with C alpha-RMSD values to the native of 2 A or less in the transmembrane regions) may be obtained for template sequence identities of 30% or higher if an accurate alignment of the sequences is used. Second, we show that secondary-structure prediction algorithms that were developed for water-soluble proteins perform approximately as well for membrane proteins. Third, we provide a comparison of a set of commonly used sequence alignment algorithms as applied to membrane proteins. We find that high-accuracy alignments of membrane protein sequences can be obtained using state-of-the-art profile-to-profile methods that were developed for water-soluble proteins. Improvements are observed when weights derived from the secondary structure of the query and the template are used in the scoring of the alignment, a result which relies on the accuracy of the secondary-structure prediction of the query sequence. The most accurate alignments were obtained using template profiles constructed with the aid of structural alignments. In contrast, a simple sequence-to-sequence alignment algorithm, using a membrane protein-specific substitution matrix, shows no improvement in alignment accuracy. We suggest that profile-to-profile alignment methods should be adopted to maximize the accuracy of homology models of membrane proteins.

Algorithms↗

Genomic multiple sequence alignments: refinement using a genetic algorithm.

BACKGROUND: Genomic sequence data cannot be fully appreciated in isolation. Comparative genomics--the practice of comparing genomic sequences from different species--plays an increasingly important role in understanding the genotypic differences between species that result in phenotypic differences as well as in revealing patterns of evolutionary relationships. One of the major challenges in comparative genomics is producing a high-quality alignment between two or more related genomic sequences. In recent years, a number of tools have been developed for aligning large genomic sequences. Most utilize heuristic strategies to identify a series of strong sequence similarities, which are then used as anchors to align the regions between the anchor points. The resulting alignment is globally correct, but in many cases is suboptimal locally. We describe a new program, GenAlignRefine, which improves the overall quality of global multiple alignments by using a genetic algorithm to improve local regions of alignment. Regions of low quality are identified, realigned using the program T-Coffee, and then refined using a genetic algorithm. Because a better COFFEE (Consistency based Objective Function For alignmEnt Evaluation) score generally reflects greater alignment quality, the algorithm searches for an alignment that yields a better COFFEE score. To improve the intrinsic slowness of the genetic algorithm, GenAlignRefine was implemented as a parallel, cluster-based program. RESULTS: We tested the GenAlignRefine algorithm by running it on a Linux cluster to refine sequences from a simulation, as well as refine a multiple alignment of 15 Orthopoxvirus genomic sequences approximately 260,000 nucleotides in length that initially had been aligned by Multi-LAGAN. It took approximately 150 minutes for a 40-processor Linux cluster to optimize some 200 fuzzy (poorly aligned) regions of the orthopoxvirus alignment. Overall sequence identity increased only slightly; but significantly, this occurred at the same time that the overall alignment length decreased--through the removal of gaps--by approximately 200 gapped regions representing roughly 1,300 gaps. CONCLUSION: We have implemented a genetic algorithm in parallel mode to optimize multiple genomic sequence alignments initially generated by various alignment tools. Benchmarking experiments showed that the refinement algorithm improved genomic sequence alignments within a reasonable period of time.

Algorithms↗

Exact asymptotic results for the Bernoulli matching model of sequence alignment.

Finding analytically the statistics of the longest common subsequence (LCS) of a pair of random sequences drawn from c alphabets is a challenging problem in computational evolutionary biology. We present exact asymptotic results for the distribution of the LCS in a simpler, yet nontrivial, variant of the original model called the Bernoulli matching (BM) model. We show that in the BM model, for all c , the distribution of the asymptotic length of the LCS, suitably scaled, is identical to the Tracy-Widom distribution of the largest eigenvalue of a random matrix whose entries are drawn from a Gaussian unitary ensemble.

Binomial Distribution↗

Identifying sequence-structure pairs undetected by sequence alignments.

We examine how effectively simple potential functions previously developed can identify compatibilities between sequences and structures of proteins for database searches. The potential function consists of pairwise contact energies, repulsive packing potentials of residues for overly dense arrangement and short-range potentials for secondary structures, all of which were estimated from statistical preferences observed in known protein structures. Each potential energy term was modified to represent compatibilities between sequences and structures for globular proteins. Pairwise contact interactions in a sequence-structure alignment are evaluated in a mean field approximation on the basis of probabilities of site pairs to be aligned. Gap penalties are assumed to be proportional to the number of contacts at each residue position, and as a result gaps will be more frequently placed on protein surfaces than in cores. In addition to minimum energy alignments, we use probability alignments made by successively aligning site pairs in order by pairwise alignment probabilities. The results show that the present energy function and alignment method can detect well both folds compatible with a given sequence and, inversely, sequences compatible with a given fold, and yield mostly similar alignments for these two types of sequence and structure pairs. Probability alignments consisting of most reliable site pairs only can yield extremely small root mean square deviations, and including less reliable pairs increases the deviations. Also, it is observed that secondary structure potentials are usefully complementary to yield improved alignments with this method. Remarkably, by this method some individual sequence-structure pairs are detected having only 5-20% sequence identity.

Algorithms↗

Identification of potential ferric binding residues in the iron-binding protein of pathogenic Neisseria meningitidis through structure-based multiple sequence alignments.

The ferric iron-binding proteins of pathogenic Neisseria display structural and metal-binding properties characteristic of the transferrin family. In the absence of structural data for the ferric iron-binding proteins, spacial folding templates have been derived for the meningococcal protein from complete and partial structure-based multiple sequence alignments with structurally related proteins. The templates have been used to identify a number of potential iron-binding residues. These include four residues that are identical with the iron coordinating ligands of transferrin, but only two reside within equivalent structural elements.

Amino Acid Sequence↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗

Prediction of membrane protein topology utilizing multiple sequence alignments.

A technique for prediction of protein membrane topology (intra- and extracellular sidedness) has been developed. Membrane-spanning segments are first predicted using an algorithm based upon multiply aligned amino acid sequences. The compositional differences in the protein segments exposed at each side of the membrane are then investigated. The ratios are calculated for Asn, Asp, Gly, Phe, Pro, Trp, Tyr, and Val, mostly found on the extracellular side, and for Ala, Arg, Cys, and Lys, mostly occurring on the intracellular side. The consensus over these 12 residue distributions is used for sidedness prediction. The method was developed with a set of 42 protein families for which all but one were correctly predicted with the new algorithm. This represents an improvement over previous techniques. The new method, applied to a set of 12 membrane protein families different from the test set and with recently determined topologies, performed well, with 11 of 12 sidedness assignments agreeing with experimental results. The method has also been applied to several membrane protein families for which the topology has yet to be determined. An electronic prediction service is available at the E-mail address tmap@embl-heidelberg.de and on WWW via http://www.embl-heidelberg.de.

Amino Acid Sequence↗