PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

JavaShade: multiple sequence alignment box-and-shading on the World Wide Web.

RESULTS: JavaShade is a multiple sequence alignment box-and-shade tool for generating high quality printed output that uses a variety of methods for boxing and shading, allowing the most appropriate functions to be chosen for displaying the most meaningful positions in an alignment. AVAILABILITY: JavaShade is available from the WWW at http://industry.ebi.ac.uk/JavaShade

Algorithms↗

Greene SCPrimer: a rapid comprehensive tool for designing degenerate primers from multiple sequence alignments.

Polymerase chain reaction (PCR) is widely applied in clinical and environmental microbiology. Primer design is key to the development of successful assays and is often performed manually by using multiple nucleic acid alignments. Few public software tools exist that allow comprehensive design of degenerate primers for large groups of related targets based on complex multiple sequence alignments. Here we present a method for designing such primers based on tree building followed by application of a set covering algorithm, and demonstrate its utility in compiling Multiplex PCR primer panels for detection and differentiation of viral pathogens.

Algorithms↗

An efficient Z-score algorithm for assessing sequence alignments.

We describe an alternative method for scoring of the pairwise alignment of two biological sequences. Designed to overcome the bias due to the composition of the alignment, it measures the distance (in standard deviations) between the given alignment and the mean value of all other alignments that can be obtained by a permutation of either sequence. We demonstrate that the standard deviation can be calculated efficiently. By concentrating upon the ungapped case, the mean and standard deviation can be calculated exactly and in two steps, the first being O(N) time, where N is the length of the sequence, the second in a fixed number of calculations, i.e., in O(1) time. We argue that this statistic is a more consistent measure than a similarity score based upon a standard scoring matrix. Even in the ungapped case, the statistic proves in many cases to be more accurate than the commonly used (FASTA) (Pearson and Lipman, 1988) gapped Z-score in which the sequence is matched against a random sample of the database. We demonstrate the use of the POZ-score as a secondary filter which screens out several well-known types of false positive, reducing the amount of manual screening to be done by the biologist.

Algorithms↗

A comprehensive comparison of multiple sequence alignment programs.

In recent years improvements to existing programs and the introduction of new iterative algorithms have changed the state-of-the-art in protein sequence alignment. This paper presents the first systematic study of the most commonly used alignment programs using BAliBASE benchmark alignments as test cases. Even below the 'twilight zone' at 10-20% residue identity, the best programs were capable of correctly aligning on average 47% of the residues. We show that iterative algorithms often offer improved alignment accuracy though at the expense of computation time. A notable exception was the effect of introducing a single divergent sequence into a set of closely related sequences, causing the iteration to diverge away from the best alignment. Global alignment programs generally performed better than local methods, except in the presence of large N/C-terminal extensions and internal insertions. In these cases, a local algorithm was more successful in identifying the most conserved motifs. This study enables us to propose appropriate alignment strategies, depending on the nature of a particular set of sequences. The employment of more than one program based on different alignment techniques should significantly improve the quality of automatic protein sequence alignment methods. The results also indicate guidelines for improvement of alignment algorithms.

Algorithms↗

A comparative study of available software for high-accuracy homology modeling: from sequence alignments to structural models.

An open question in protein homology modeling is, how well do current modeling packages satisfy the dual criteria of quality of results and practical ease of use? To address this question objectively, we examined homology-built models of a variety of therapeutically relevant proteins. The sequence identities across these proteins range from 19% to 76%. A novel metric, the difference alignment index (DAI), is developed to aid in quantifying the quality of local sequence alignments. The DAI is also used to construct the relative sequence alignment (RSA), a new representation of global sequence alignment that facilitates comparison of sequence alignments from different methods. Comparisons of the sequence alignments in terms of the RSA and alignment methodologies are made to better understand the advantages and caveats of each method. All sequence alignments and corresponding 3D models are compared to their respective structure-based alignments and crystal structures. A variety of protein modeling software was used. We find that at sequence identities >40%, all packages give similar (and satisfactory) results; at lower sequence identities (<25%), the sequence alignments generated by Profit and Prime, which incorporate structural information in their sequence alignment, stand out from the rest. Moreover, the model generated by Prime in this low sequence identity region is noted to be superior to the rest. Additionally, we note that DSModeler and MOE, which generate reasonable models for sequence identities >25%, are significantly more functional and easier to use when compared with the other structure-building software.

Algorithms↗

A fast and sensitive multiple sequence alignment algorithm.

A two-step multiple alignment strategy is presented that allows rapid alignment of a set of homologous sequences and comparison of pre-aligned groups of sequences. Examples are given demonstrating the improvement in the quality of alignments when comparing entire groups instead of single sequences. The modular design of computer programs based on this algorithm allows for storage of aligned sequences and successive alignment of any number of sequences.

Algorithms↗

Building multiple sequence alignments with a flavor of HSSP alignments.

Homology-derived secondary structure of proteins (HSSP) is a well-known database of multiple sequence alignments (MSAs) which merges information of protein sequences and their three-dimensional structures. It is available for all proteins whose structure is deposited in the PDB. It is also used by STING and (Java)Protein Dossier to calculate and present relative entropy as a measure of the degree of conservation for each residue of proteins whose structure has been solved and deposited in the PDB. However, if the STING and (Java)Protein Dossier are to provide support for analysis of protein structures modeled in computers or being experimentally solved but not yet deposited in the PDB, then we need a new method for building alignments having a flavor of HSSP alignments (myMSAr). The present study describes a new method and its corresponding databank (SH2QS--database of sequences homologue to the query [structure-having] sequence). Our main interest in making myMSAr was to measure the degree of residue conservation for a given query sequence, regardless of whether it has a corresponding structure deposited in the PDB. In this study, we compare the measurement of residue conservation provided by corresponding alignments produced by HSSP and SH2QS. As a case study, we also present two biologically relevant examples, the first one highlighting the equivalence of analysis of the degree of residue conservation by using HSSP or SH2QS alignments, and the second one presenting the degree of residue conservation for a structure modeled in a computer, which , as a consequence, does not have an alignment reported by HSSP.

Amino Acid Sequence↗

Do aligned sequences share the same fold?

Sequence comparison remains a powerful tool to assess the structural relatedness of two proteins. To develop a sensitive sequence-based procedure for fold recognition, we performed an exhaustive global alignment (with zero end gap penalties) between sequences of protein domains with known three-dimensional folds. The subset of 1.3 million alignments between sequences of structurally unrelated domains was used to derive a set of analytical functions that represent the probability of structural significance for any sequence alignment at a given sequence identity, sequence similarity and alignment score. Analysis of overlap between structurally significant and insignificant alignments shows that sequence identity and sequence similarity measures are poor indicators of structural relatedness in the "twilight zone", while the alignment score allows much better discrimination between alignments of structurally related and unrelated sequences for a wide variety of alignment settings. A fold recognition benchmark was used to compare eight different substitution matrices with eight sets of gap penalties. The best performing matrices were Gonnet and Blosum50 with normalized gap penalties of 2.4/0.15 and 2.0/0.15, respectively, while the positive matrices were the worst performers. The derived functions and parameters can be used for fold recognition via a multilink chain of probability weighted pairwise sequence alignments.

Databases as Topic↗

Generalized affine gap costs for protein sequence alignment.

Based on the observation that a single mutational event can delete or insert multiple residues, affine gap costs for sequence alignment charge a penalty for the existence of a gap, and a further length-dependent penalty. From structural or multiple alignments of distantly related proteins, it has been observed that conserved residues frequently fall into ungapped blocks separated by relatively nonconserved regions. To take advantage of this structure, a simple generalization of affine gap costs is proposed that allows nonconserved regions to be effectively ignored. The distribution of scores from local alignments using these generalized gap costs is shown empirically to follow an extreme value distribution. Examples are presented for which generalized affine gap costs yield superior alignments from the standpoints both of statistical significance and of alignment accuracy. Guidelines for selecting generalized affine gap costs are discussed, as is their possible application to multiple alignment.

Algorithms↗

Generating models of surgical procedures using UMLS concepts and multiple sequence alignment.

Surgical procedures can be viewed as a process composed of a sequence of steps performed on, by, or with the patient's anatomy. This sequence is typically the pattern followed by surgeons when generating surgical report narratives for documenting surgical procedures. This paper describes a methodology for semi-automatically deriving a model of conducted surgeries, utilizing a sequence of derived Unified Medical Language System (UMLS) concepts for representing surgical procedures. A multiple sequence alignment was computed from a collection of such sequences and was used for generating the model. These models have the potential of being useful in a variety of informatics applications such as information retrieval and automatic document generation.

Abstracting and Indexing↗

The size distribution of insertions and deletions in human and rodent pseudogenes suggests the logarithmic gap penalty for sequence alignment.

The size distributions of deletions, insertions, and indels (i.e., insertions or deletions) were studied, using 78 human processed pseudogenes and other published data sets. The following results were obtained: (1) Deletions occur more frequently than do insertions in sequence evolution; none of the pseudogenes studied shows significantly more insertions than deletions. (2) Empirically, the size distributions of deletions, insertions, and indels can be described well by a power law, i.e., fk = Ck-b, where fk is the frequency of deletion, insertion, or indel with gap length k, b is the power parameter, and C is the normalization factor. (3) The estimates of b for deletions and insertions from the same data set are approximately equal to each other, indicating that the size distributions for deletions and insertions are approximately identical. (4) The variation in the estimates of b among various data sets is small, indicating that the effect of local structure exists but only plays a secondary role in the size distribution of deletions and insertions. (5) The linear gap penalty, which is most commonly used in sequence alignment, is not supported by our analysis; rather, the power law for the size distribution of indels suggests that an appropriate gap penalty is wk = a + b ln k, where a is the gap creation cost and blnk is the gap extension cost. (6) The higher frequency of deletion over insertion suggests that the gap creation cost of insertion (ai) should be larger than that of deletion (ad); that is, ai - ad = ln R, where R is the frequency ratio of deletions to insertions.

Animals↗

Yeast TOR (DRR) proteins: amino-acid sequence alignment and identification of structural motifs.

The yeast TOR1 (DRR1) and TOR2 (DRR2) proteins are putative targets of the immunosuppressive drug rapamycin (Rm), defined by dominant drug-resistance mutations. They share a large C-terminal domain that exhibits sequence similarity to the 110-kDa subunit of phosphatidylinositol (PI) 3-kinases. In this report, we present an amino acid (aa) sequence alignment of TOR1 (DRR1) and TOR2 (DRR2) and identify conserved and nonconserved motifs within the N-terminal domain that are indicative of possible nuclear localization. We also show that the mutations responsible for Rm resistance in four independent drr2dom alleles alter the identical aa (Ser1975-->Arg) previously identified in drr1dom mutants (Ser1972-->Arg or Asn). Models for TOR (DRR) protein function are discussed.

Alleles↗

Bayesian coestimation of phylogeny and sequence alignment.

BACKGROUND: Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. RESULTS: We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. CONCLUSION: Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

Algorithms↗

Fast and sensitive multiple sequence alignments on a microcomputer.

A strategy is described for the rapid alignment of many long nucleic acid or protein sequences on a microcomputer. The program described can handle up to 100 sequences of 1200 residues each. The approach is based on progressively aligning sequences according to the branching order in an initial phylogenetic tree. The results obtained using the package appear to be as sensitive as those from any other available method.

Algorithms↗

Secondary structure prediction of beta-subunits of the gonadotropin-thyrotropin family from its aligned sequences using environment-dependent amino-acid substitution tables and conformational propensities.

The secondary structures of beta-subunits of the glycoprotein hormone family, LH (luteinizing hormone), CG (chorionic gonadotropin), FSH (follicle stimulating hormone), TSH (thyroid stimulating hormone), and GTH I/GTH II (two types of fish gonadotropins), are predicted by comparing an amino-acid substitution pattern at equivalent sites in their aligned sequences with environment-dependent amino-acid substitution tables and conformational propensities calculated from other protein families whose three-dimensional structures are known. According to the prediction results, together with other structural information obtained from experiments, the following points come up as important structural features of the beta-subunits of this family; The regions assigned to regular secondary structures (one alpha-helix and three beta-strands) are considered to constitute a core of the beta-subunits. They involve interaction sites with carbohydrate and alpha-subunit. Out of the six disulfide bonds formed in the beta-subunit, four are located together on one side of the core, and the other two on the opposite side. The two regions assumed to be a receptor binding region from experiments (therefore, species-specific regions) are predicted as loops located on the same side of the beta-subunit in this study. Some of the predicted loops are rich in proline residues. While the positions of proline residues are conserved in the family generally, there are hormone- or species-specific ones in the loop that is assumed to take part in receptor binding. The possible importance of proline residues in hormone or species specificity is discussed. (After submitting the manuscript the X-ray crystal structure of human CG was published. In order to evaluate the prediction, the original manuscript is kept intact and a comparison has been made between the prediction results and the crystal structure in an appendix).

Amino Acid Sequence↗

The partition matrix: exploring variable phylogenetic signals along nucleotide sequence alignments.

The partition matrix is a graphical tool for comparative analysis of nucleotide sequences following alignment. It is particularly useful for investigating the divergent phylogenies of sequence regions undergoing reticulate evolution. A partition matrix is generated by determining the consistency of the parsimoniously informative sites in a set of aligned sequences with the binary partitions inferred from the sequences. Since the linear order of sites is maintained, the matrix can be used to assess whether the distribution of sites either supporting or conflicting with particular partitions changes along the length of the alignment. The usefulness of the matrix in allowing visual identification of differences in evolutionary history among regions depends on the order in which partitions are shown; several suitable ordering schemes are proposed. We demonstrate the use of the partition matrix in interpreting the evolution of the pseudoautosomal boundary region on the sex chromosome of catarrhine primates. Its routine use should help to avoid attempts to derive single phylogenies from sequences whose evolution has been reticulate and to identify the gene conversion or recombination events underlying the reticulation. The method is relatively fast. It is exploratory, and it can form the basis for more formal analysis, which we discuss.

Animals↗

Models of the primary and secondary structure for the 12S rRNA of birds: a guideline for sequence alignment.

Models of the primary and secondary structure for the 12S ribosomal RNA (rRNA) gene of birds is presented based on a comparison of 100 species. Preliminary higher-order structures were delimited following the model for vertebrates. Paired regions were refined following a complementary base-pairing criterion and compensatory mutations were considered as a further confirmation of their existence. The model shows 40 stems, 20 internal loops and 17 external loops, arranged in the typical four domains suggested for the small subunit rRNA. The higher-order structures recovered in the model were used to build a multiple sequence alignment appropriate for phylogenetic analysis. The phylogeny recovered from this alignment was compared with trees inferred from alignments assembled using different alignment parameters in the program ClustalW. The alignment based on the secondary structure is sensitive to positional covariation of stems. Nonetheless, the phylogeny recovered with this method resulted in relationships that are more congruent with non-molecular data than those inferred from alternative alignments.

Amino Acid Sequence↗

Significant improvement in accuracy of multiple protein sequence alignments by iterative refinement as assessed by reference to structural alignments.

The relative performances of four strategies for aligning a large number of protein sequences were assessed by referring to corresponding structural alignments of 54 independent families. Multiple sequence alignment of a family was constructed by a given method from the sequences of known structures and their homologues, and the subset consisting of the sequences of known structures was extracted from the whole alignment and compared with the structural counterpart in a residue-to-residue fashion. Gap-opening and -extension penalties were optimized for each family and method. Each of the four multiple alignment methods gave significantly more accurate alignments than the conventional pairwise method. In addition, a clear difference in performance was detected among three of the four multiple alignment methods examined. The currently most popular progressive method ranked worst among the four, and the randomized iterative strategy that optimizes the sum-of-pairs score ranked next worst. The two best-performing strategies, one of which was newly developed, both pursue an optimal weighted sum-of-pairs score, where the pair weights were introduced to correct for uneven representations of subgroups in a family. The new method uses doubly nested iterations to make alignment, phylogenetic tree and pair weights mutually consistent. Most importantly, the improvement in accuracy of alignments obtained by these iterative methods over pairwise or progressive method tends to increase with decreasing average sequence identity, implying that iterative refinement is more effective for the generally difficult alignment of remotely related sequences. Four well-known amino acid substitution matrices were also tested in combination with the various methods. However, the effects of substitution matrices were found to be minor in the framework of multiple alignment, and the same order of relative performance of the alignment methods was observed with any of the matrices.

Algorithms↗