PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence↗

A transition probability model for amino acid substitutions from blocks.

Substitution matrices have been useful for sequence alignment and protein sequence comparisons. The BLOSUM series of matrices, which had been derived from a database of alignments of protein blocks, improved the accuracy of alignments previously obtained from the PAM-type matrices estimated from only closely related sequences. Although BLOSUM matrices are scoring matrices now widely used for protein sequence alignments, they do not describe an evolutionary model. BLOSUM matrices do not permit the estimation of the actual number of amino acid substitutions between sequences by correcting for multiple hits. The method presented here uses the Blocks database of protein alignments, along with the additivity of evolutionary distances, to approximate the amino acid substitution probabilities as a function of actual evolutionary distance. The PMB (Probability Matrix from Blocks) defines a new evolutionary model for protein evolution that can be used for evolutionary analyses of protein sequences. Our model is directly derived from, and thus compatible with, the BLOSUM matrices. The model has the additional advantage of being easily implemented.

Amino Acid Substitution↗

Critical aspartic acid residues in pseudouridine synthases.

The pseudouridine synthases catalyze the isomerization of uridine to pseudouridine at particular positions in certain RNA molecules. Genomic data base searches and sequence alignments using the first four identified pseudouridine synthases led Koonin (Koonin, E. V. (1996) Nucleic Acids Res. 24, 2411-2415) and, independently, Santi and co-workers (Gustafsson, C., Reid, R., Greene, P. J., and Santi, D. V. (1996) Nucleic Acids Res. 24, 3756-3762) to group this class of enzyme into four families, which display no statistically significant global sequence similarity to each other. Upon further scrutiny (Huang, H. L., Pookanjanatavip, M., Gu, X. G., and Santi, D. V. (1998) Biochemistry 37, 344-351), the Santi group discovered that a single aspartic acid residue is the only amino acid present in all of the aligned sequences; they then demonstrated that this aspartic acid residue is catalytically essential in one pseudouridine synthase. To test the functional significance of the sequence alignments in light of the global dissimilarity between the pseudouridine synthase families, we changed the aspartic acid residue in representatives of two additional families to both alanine and cysteine: the mutant enzymes are catalytically inactive but retain the ability to bind tRNA substrate. We have also verified that the mutant enzymes do not release uracil from the substrate at a rate significant relative to turnover by the wild-type pseudouridine synthases. Our results clearly show that the aligned aspartic acid residue is critical for the catalytic activity of pseudouridine synthases from two additional families of these enzymes, supporting the predictive power of the sequence alignments and suggesting that the sequence motif containing the aligned aspartic acid residue might be a prerequisite for pseudouridine synthase function.

Amino Acid Sequence↗

Improvement of the GenTHREADER method for genomic fold recognition.

MOTIVATION: In order to enhance genome annotation, the fully automatic fold recognition method GenTHREADER has been improved and benchmarked. The previous version of GenTHREADER consisted of a simple neural network which was trained to combine sequence alignment score, length information and energy potentials derived from threading into a single score representing the relationship between two proteins, as designated by CATH. The improved version incorporates PSI-BLAST searches, which have been jumpstarted with structural alignment profiles from FSSP, and now also makes use of PSIPRED predicted secondary structure and bi-directional scoring in order to calculate the final alignment score. Pairwise potentials and solvation potentials are calculated from the given sequence alignment which are then used as inputs to a multi-layer, feed-forward neural network, along with the alignment score, alignment length and sequence length. The neural network has also been expanded to accommodate the secondary structure element alignment (SSEA) score as an extra input and it is now trained to learn the FSSP Z-score as a measurement of similarity between two proteins. RESULTS: The improvements made to GenTHREADER increase the number of remote homologues that can be detected with a low error rate, implying higher reliability of score, whilst also increasing the quality of the models produced. We find that up to five times as many true positives can be detected with low error rate per query. Total MaxSub score is doubled at low false positive rates using the improved method. AVAILABILITY: http://www.psipred.net.

Amino Acid Sequence↗

Retrieval and on-the-fly alignment of sequence fragments from the HIV database.

MOTIVATION: The amount of HIV-1 sequence data generated (presently around 42000 sequences, of which more than 22000 are from the V3 region of the viral envelope) presents a challenge for anyone working on the analysis of these data. A major problem is obtaining the region of interest from the stored sequences, which often contain but are not limited to that region. In addition, multiple alignment programs generally cannot deal with the large numbers of sequences that are available for many HIV-1 regions. We set out to provide our users with a tool that will retrieve and create an initial alignment of the HIV sequences that are available for a given genomic region. RESULTS: The MPAlign (Multiple Pairwise Alignment) web interface is a collection of Perl scripts that retrieves sequences from the Los Alamos HIV sequence database based on a number of search parameters. All sequences were pairwise-aligned to a model sequence using the Hidden Markov Model-based program HMMER. The HMMER model is general enough to accommodate virtually all HIV-1 sequences stored in the database. To create a multiple sequence alignment, gaps were inserted into the sequences during retrieval, so that they are aligned to one another. Retrieving and aligning the almost 560 gp120 sequences (approximately>1500 nt) stored in the database is at least 1500 times faster than a similar Clustal alignment.

Algorithms↗

Sequence-based alignment of sorghum chromosome 3 and rice chromosome 1 reveals extensive conservation of gene order and one major chromosomal rearrangement.

The completed rice genome sequence will accelerate progress on the identification and functional classification of biologically important genes and serve as an invaluable resource for the comparative analysis of grass genomes. In this study, methods were developed for sequence-based alignment of sorghum and rice chromosomes and for refining the sorghum genetic/physical map based on the rice genome sequence. A framework of 135 BAC contigs spanning approximately 33 Mbp was anchored to sorghum chromosome 3. A limited number of sequences were collected from 118 of the BACs and subjected to BLASTX analysis to identify putative genes and BLASTN analysis to identify sequence matches to the rice genome. Extensive conservation of gene content and order between sorghum chromosome 3 and the homeologous rice chromosome 1 was observed. One large-scale rearrangement was detected involving the inversion of an approximately 59 cM block of the short arm of sorghum chromosome 3. Several small-scale changes in gene collinearity were detected, indicating that single genes and/or small clusters of genes have moved since the divergence of sorghum and rice. Additionally, the alignment of the sorghum physical map to the rice genome sequence allowed sequence-assisted assembly of an approximately 1.6 Mbp sorghum BAC contig. This streamlined approach to high-resolution genome alignment and map building will yield important information about the relationships between rice and sorghum genes and genomic segments and ultimately enhance our understanding of cereal genome structure and evolution.

Base Sequence↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

Comparison of the periplasmic receptors for L-arabinose, D-glucose/D-galactose, and D-ribose. Structural and Functional Similarity.

The primary sequence of the receptor for L-arabinose or Ara-binding protein (ABP) composed of 306 residues is very different from the D-glucose/D-galactose-binding protein (GGBP) which consists of 309 residues. Nevertheless, superimpositioning of the well-refined high resolution structures of ABP in complex with D-galactose and the GGBP in complex with D-glucose shows very similar structures; 220 of the residues (or about 70%) have a root mean square deviation of 2.0 A. From the superpositioning, nine pairs of continuous segments (consisting of 8-51 residues), mainly alpha-helices and beta-strands that form the core of the two lobes of the bilobate proteins were found to exhibit strong sequence homology. The equivalenced structures and aligned sequences show that many of the polar, as well as aromatic residues, in the sugar-binding sites located in the cleft between the two lobes are highly conserved. Surprisingly, however, the exact mode of binding of the D-galactose in ABP is totally different from that of the D-glucose in GGBP. Using the structurally aligned sequences of the ABP and GGBP as a template, we have matched the sequence of the ribose-binding protein (RBP) which consists of 271 residues with the ABP/GGBP pair. Although the nine aligned segments of all three proteins show little sequence identity, they have significant homology. Four additional segments of RBP were matched only with GGBP, leading to the alignment of about 90% of the RBP sequence with the GGBP sequence. Many of the conserved residues in the binding sites of ABP and GGBP matched with similar residues in RBP. Additional observations indicate that the GGBP/RBP pair is more closely related than the ABP/RBP or ABP/GGBP pair. All three binding proteins, which may have diverged from a common ancestor, serve as primary receptors for bacterial high affinity active transport systems. Moreover, GGBP and RBP, but not ABP, also act as receptors for chemotaxis. An exposed site located in one domain, which includes Gly74, for interacting with the trg transmembrane signal transducer that is involved in triggering chemotaxis has been located in the structure of GGBP (Vyas, N.K., Vyas, M.N., and Quiocho, F.A. (1988) Science 242, 1290-1295). Whereas the site is absent in the structure of ABP, it is strongly predicted to be present in RBP which shares the same trg transducer with GGBP. The knowledge-based alignment of RBP further revealed two possible additional peripheral chemotactic sites that show high structural and sequence similarity between GGBP and RBP only. At least one of these sites, together with the one proven to exist in the other domain, could be used by the signal transducer with which both binding proteins interact in a way which the substrate-loaded "closed cleft" structure could be discriminated from the unliganded "open cleft" form by the transducer.

Amino Acid Sequence↗

Protein topology and stability define the space of allowed sequences.

We describe a new approach to explore and quantify the sequence space associated with a given protein structure. A set of sequences are optimized for a given target structure, using all-atom models and a physical energy function. Specificity of the sequence for its target is ensured by using the random energy model, which keeps the amino acid composition of the sequence constant. The designed sequences provide a multiple sequence alignment that describes the sequence space compatible with the structure of interest; here the size of this space is estimated by using an information entropy measure. In parallel, multiple alignments of naturally occurring sequences can be derived by using either sequence or structure alignments. We compared these 3 independent multiple sequence alignments for 10 different proteins, ranging in size from 56 to 310 residues. We observed that the subset of the sequence space derived by using our design procedure is similar in size to the sequence spaces observed in nature. These results suggest that the volume of sequence space compatible with a given protein fold is defined by the length of the protein as well as by the topology (i.e., geometry of the polypeptide chain) and the stability (i.e., free energy of denaturation) of the fold.

Amino Acid Sequence↗

Recco: recombination analysis using cost optimization.

MOTIVATION: Recombination plays an important role in the evolution of many pathogens, such as HIV or malaria. Despite substantial prior work, there is still a pressing need for efficient and effective methods of detecting recombination and analyzing recombinant sequences. RESULTS: We introduce Recco, a novel fast method that, given a multiple sequence alignment, scores the cost of obtaining one of the sequences from the others by mutation and recombination. The algorithm comes with an illustrative visualization tool for locating recombination breakpoints. We analyze the sequence alignment with respect to all choices of the parameter alpha weighting recombination cost against mutation cost. The analysis of the resulting cost curve yields additional information as to which sequence might be recombinant. On random genealogies Recco is comparable in its power of detecting recombination with the algorithm Geneconv (Sawyer, 1989). For specific relevant recombination scenarios Recco significantly outperforms Geneconv.

Algorithms↗

A word-oriented approach to alignment validation.

MOTIVATION: Multiple sequence alignment at the level of whole proteomes requires a high degree of automation, precluding the use of traditional validation methods such as manual curation. Since evolutionary models are too general to describe the history of each residue in a protein family, there is no single algorithm/model combination that can yield a biologically or evolutionarily optimal alignment. We propose a 'shotgun' strategy where many different algorithms are used to align the same family, and the best of these alignments is then chosen with a reliable objective function. We present WOOF, a novel 'word-oriented' objective function that relies on the identification and scoring of conserved amino acid patterns (words) between pairs of sequences. RESULTS: Tests on a subset of reference protein alignments from BAliBASE showed that WOOF tended to rank the (manually curated) reference alignment highest among 1060 alternative (automatically generated) alignments for a majority of protein families. Among the automated alignments, there was a strong positive relationship between the WOOF score and similarity to the reference alignment. The speed of WOOF and its independence from explicit considerations of three-dimensional structure make it an excellent tool for analyzing large numbers of protein families. AVAILABILITY: On request from the authors.

Algorithms↗

Secondary structure model for the ITS-2 precursor rRNA of strongyloid nematodes of equids: implications for phylogenetic inference.

In order to maximise the positional homology in the primary sequence alignment of the second internal transcribed spacer for 30 species of equine strongyloid nematodes, the secondary structures of the precursor ribosomal RNA were predicted using an approach combining an energy minimisation method and comparative sequence analysis. The results indicated that a common secondary structure model of the second internal transcribed spacer of these nematodes was maintained despite significant interspecific differences (2-56%) in primary sequences. The secondary structure model was then used to refine the primary second internal transcribed spacer sequence alignment. The 'manual' and 'structure' alignments were both subjected to phylogenetic analysis to compare the effect of using different sequence alignments on phylogenetic inference. The topologies of the phylogenetic trees inferred from the manual second internal transcribed spacer alignment were usually different to those derived from the structure second internal transcribed spacer alignment. The results suggested that the positional homology in the second internal transcribed spacer primary sequence alignment was maximised when the secondary structure model was taken into consideration.

Animals↗

Subset partitioning of the ribosomal DNA small subunit and its effects on the phylogeny of the Anopheles punctulatus group.

A phylogenetic study, based on maximum parsimony, of ten species in the Anopheles punctulatus group of malaria vectors from the south-west Pacific was performed using structural and similarity-based DNA sequence alignments of the nuclear small ribosomal subunit (SSU = 18S). The structural alignment proved to be more informative than a computer generated similarity-based alignment. Analyses involving the full structural sequence alignment (2169 bp) and the helical regions (1547 bp) resolved a single tree of the same topology, while analyses using the similarity based alignment could not resolve the group. Studies on the three structural domains of the nuclear rDNA SSU identified domain 2 (769 bp) as the only region informative at the sibling-species level and resulted in the same tree as the full structural sequence and helical regions. The main conclusions of these studies were that the An. punctulatus group formed two clades: a Farauti clade containing members displaying an all black scaled proboscis (An. farauti 1-3 and 5-7) and a Punctulatus clade containing members that display some degree of white scaling on the proboscis (An. farauti 4, An. punctulatus and An. species near punctulatus). Anopheles koliensis can display either proboscis morphology and was positioned basal to the Farauti Clade. These results do not fully concord with those derived from the mitochondrial COII gene.

Animals↗

Multiple alignment of sequences on parallel computers.

A software package that allows one to carry out multiple alignment of protein and nucleic acid sequences of almost unlimited length and number of sequences is developed on C-DAC parallel computer--a transputer-based machine. The farming approach is used for data parallelization. The speed gains are almost linear when the number of transputers is increased from 4 to 64. The software is used to carry out multiple alignment of 100 sequences each of alpha-chain and beta-chain of hemoglobin and 83 cytochrome c sequences. The signature sequence of cytochrome c was found to be PGTKMXF. The single parameter, multiple alignment score, S, has been used to categorize proteins in different subfamilies and groups.

Algorithms↗

Feature matching and segmentation in motion perception.

We examined the role of feature matching in motion perception. The stimulus sequence was constructed from a vertical, 1 cycle deg-1 sinusoidal grating divided into horizontal strips of equal height, where alternate strips moved leftward and rightward. The initial relative phase of adjacent strips was either 0 degree (aligned) or 90 degrees (non-aligned) and the motion was sampled at 90 degrees phase steps. A blank interstimulus interval (ISI) of 0-117 ms was introduced between each 33 ms presentation of the stimulus frames. The observers had to identify the direction of motion of the central strip. Motion was perceived correctly at short ISIs, but at longer ISIs performance was much better for the non-aligned sequence than the aligned sequence. This difference in performance may reflect a role for feature correspondence and grouping of features in motion perception at longer ISIs. In the aligned sequence half the frames consisted of a single coherent vertical grating, while the interleaved frames contained short strips. We argue that to achieve feature matching over time, the long edge and bar features must be broken up perceptually (segmented) into shorter elements before these short segments can appear to move in opposite directions. This idea correctly predicted that overlaying narrow, stationary, black horizontal lines at the junctions of the grating strips would improve performance in the aligned condition. The results support the view that, in addition to motion energy, feature analysis and feature tracking play an important role in motion perception.

Humans↗

Confidence measures for protein fold recognition.

MOTIVATION: We present an extensive evaluation of different methods and criteria to detect remote homologs of a given protein sequence. We investigate two associated problems: first, to develop a sensitive searching method to identify possible candidates and, second, to assign a confidence to the putative candidates in order to select the best one. For searching methods where the score distributions are known, p-values are used as confidence measure with great success. For the cases where such theoretical backing is absent, we propose empirical approximations to p-values for searching procedures. RESULTS: As a baseline, we review the performances of different methods for detecting remote protein folds (sequence alignment and threading, with and without sequence profiles, global and local). The analysis is performed on a large representative set of protein structures. For fold recognition, we find that methods using sequence profiles generally perform better than methods using plain sequences, and that threading methods perform better than sequence alignment methods. In order to assess the quality of the predictions made, we establish and compare several confidence measures, including raw scores, z-scores, raw score gaps, z-score gaps, and different methods of p-value estimation. We work our way from the theoretically well backed local scores towards more explorative global and threading scores. The methods for assessing the statistical significance of predictions are compared using specificity--sensitivity plots. For local alignment techniques we find that p-value methods work best, albeit computationally cheaper methods such as those based on score gaps achieve similar performance. For global methods where no theory is available methods based on score gaps work best. By using the score gap functions as the measure of confidence we improve the more powerful fold recognition methods for which p-values are unavailable. AVAILABILITY: The benchmark set is available upon request.

Amino Acid Sequence↗

Study and prediction of secondary structure for membrane proteins.

In this paper we present a novel approach to membrane protein secondary structure prediction based on the statistical stepwise discriminant analysis method. A new aspect of our approach is the possibility to derive physical-chemical properties that may affect the formation of membrane protein secondary structure. The certain physical-chemical properties of protein chains can be used to clarify the formation of the secondary structure types under consideration. Another aspect of our approach is that the results of multiple sequence alignment, or the other kinds of sequence alignment, are not used in the frame of the method. Using our approach, we predicted the formation of three main secondary structure types (alpha-helix, beta-structure and coil) with high accuracy, that is Q(3) = 76%. Predicting the formation of alpha-helix and non-alpha-helix states we reached the accuracy which was measured as Q(2) = 86%. Also we have identified certain protein chain properties that affect the formation of membrane protein secondary structure. These protein properties include hydrophobic properties of amino acid residues, presence of Gly, Ala and Val amino acids, and the location of protein chain end.

Amino Acid Sequence↗