PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Genomic divergences among cattle, dog and human estimated from large-scale alignments of genomic sequences.

BACKGROUND: Approximately 11 Mb of finished high quality genomic sequences were sampled from cattle, dog and human to estimate genomic divergences and their regional variation among these lineages. RESULTS: Optimal three-way multi-species global sequence alignments for 84 cattle clones or loci (each >50 kb of genomic sequence) were constructed using the human and dog genome assemblies as references. Genomic divergences and substitution rates were examined for each clone and for various sequence classes under different functional constraints. Analysis of these alignments revealed that the overall genomic divergences are relatively constant (0.32-0.37 change/site) for pairwise comparisons among cattle, dog and human; however substitution rates vary across genomic regions and among different sequence classes. A neutral mutation rate (2.0-2.2 x 10(-9) change/site/year) was derived from ancestral repetitive sequences, whereas the substitution rate in coding sequences (1.1 x 10(-9) change/site/year) was approximately half of the overall rate (1.9-2.0 x 10(-9) change/site/year). Relative rate tests also indicated that cattle have a significantly faster rate of substitution as compared to dog and that this difference is about 6%. CONCLUSION: This analysis provides a large-scale and unbiased assessment of genomic divergences and regional variation of substitution rates among cattle, dog and human. It is expected that these data will serve as a baseline for future mammalian molecular evolution studies.

Animals↗

Fundamentals of massive automatic pairwise alignments of protein sequences: theoretical significance of Z-value statistics.

MOTIVATION: Different automatic methods of sequence alignments are routinely used as a starting point for homology searches and function inference. Confidence in an alignment probability is one of the major fundamentals of massive automatic genome-scale pairwise comparisons, for clustering of putative orthologs and paralogs, sequenced genome annotation or multiple-genomic tree constructions. Extreme value distribution based on the Karlin-Altschul model, usually advised for large-scale comparisons are not always valid, particularly in the case of comparisons of non-biased with nucleotide-biased genomes (such that of Plasmodium falciparum). Z-values estimates based on Monte Carlo technics, can be calculated experimentally for any alignment output, whatever the method used. Empirically, a Z-value higher than approximately 8 is supposed reasonable to assess that an alignment score is significant, but this arbitrary figure was never theoretically justified. RESULTS: In this paper, we used the Bienaymé-Chebyshev inequality to demonstrate a theorem of the upper limit of an alignment score probability (or P-value). This theorem implies that a computed Z-value is a statistical test, a single-linkage clustering criterion and that 1/Z-value(2) is an upper limit to the probability of an alignment score whatever the actual probability law is. Therefore, this study provides the missing theoretical link between a Z-value cut-off used for an automatic clustering of putative orthologs and/or paralogs, and the corresponding statistical risk in such genome-scale comparisons (using non-biased or biased genomes).

Algorithms↗

A space-efficient algorithm for aligning large genomic sequences.

SUMMARY: In the segment-by-segment approach to sequence alignment, pairwise and multiple alignments are generated by comparing gap-free segments of the sequences under study. This method is particularly efficient in detecting local homologies, and it has been used to identify functional regions in large genomic sequences. Herein, an algorithm is outlined that calculates optimal pairwise segment-by-segment alignments in essentially linear space. AVAILABILTIY: The program is available at the Bielefeld Bioinformatics Server (BiBiServ) at http://bibiserv.techfak. uni-bielefeld.de/dialign/

Algorithms↗

DIVAA: analysis of amino acid diversity in multiple aligned protein sequences.

MOTIVATION: Multiple alignments of proteins are an effective way of identifying conserved amino acids that provide clues to functional relationships among proteins. Quantitation of the abundances of amino acids found at each position in a sequence motif can provide a basis for understanding the structural and functional constraints at each point. Distribution of information across a motif has been used previously, but the non-intuitive nature of the analysis has limited its impact. RESULTS: Here, we introduce a quantitative measure of amino acid sequence diversity (DIVAA) that has a simple, intuitive meaning. Diversity, as a measure of sequence conservation or variation, is inextricably linked to the probability of selecting identical pairs from a distribution. We demonstrate its utility through the analysis of four populations: ATP-binding P-loops, hypervariable domains of kappa light chains, signal sequences, and the N- and C- termini of proteins. DIVAA provides a simple means to generate hypotheses concerning the contribution of individual residues to the functional and evolutionary relationships among proteins. AVAILABILITY: Access to DIVAA software is available at RELIC (http://relic.bio.anl.gov).

Algorithms↗

A probabilistic measure for alignment-free sequence comparison.

MOTIVATION: Alignment-free sequence comparison methods are still in the early stages of development compared to those of alignment-based sequence analysis. In this paper, we introduce a probabilistic measure of similarity between two biological sequences without alignment. The method is based on the concept of comparing the similarity/dissimilarity between two constructed Markov models. RESULTS: The method was tested against six DNA sequences, which are the thrA, thrB and thrC genes of the threonine operons from Escherichia coli K-12 and from Shigella flexneri; and one random sequence having the same base composition as thrA from E.coli. These results were compared with those obtained from CLUSTAL W algorithm (alignment-based) and the chaos game representation (alignment-free). The method was further tested against a more complex set of 40 DNA sequences and compared with other existing sequence similarity measures (alignment-free). AVAILABILITY: All datasets and computer codes written in MATLAB are available upon request from the first author.

Algorithms↗

Inverse protein folding by the residue pair preference profile method: estimating the correctness of alignments of structurally compatible sequences.

The residue pair preference profile (R3P) method is an inverse folding method that combines environmental profiles and pair preference profiles. The method uses statistical preferences for residue pairs which score the likelihood of finding a profiled residue to be paired with a residue within its local environment. All pairs are characterized by their dihedral angles, secondary structure and number of neighboring residues as a function of residue type. Each residue pair preference is expressed for all 20 amino acids of the profiled residue and is weighted by the compatibility of the environment residue with its own local environment. The R3P method produces an initial profile-sequence alignment which is then refined by converting the initial profile into a profile of a target sequence threaded into the structure of the initial profile. We have tested this method by evaluating alignments of sequences with known 3-D structures using structural superposition alignments as reference. R3P-sequence alignments are > or = 50% correct on average for sequences whose 3-D structure pairs superimpose with an r.m.s. deviation of < or = 1.97 A. The average improvement in correctness during this iterative refinement is 14%. The R3P-sequence alignments are compared with sequence-sequence and 3-D profile-sequence alignments. When all three methods are combined, on average > or = 50% of the alignments are correct for pairs of 3-D structures that superimpose within 2.12 A. A 3-D model of HisA is predicted with the combined method.

Aldose-Ketose Isomerases↗

Assignment of enzyme substrate specificity by principal component analysis of aligned protein sequences: an experimental test using DNA glycosylase homologs.

We have studied the relationship between amino acid sequence and substrate specificity in a DNA glycosylase family by characterizing experimentally the specificity of four new members of the family. We show that principal component analysis (PCA) of the sequence family correctly predicts the substrate specificity of one of the novel homologs even though conventional sequence analysis methods fail to group this homolog with other sequences of the same specificity. PCA also suggested, correctly, that another homolog characterized previously differs in its specificity from those sequences with which it clusters by conventional criteria. These results suggest that principal component analysis of sequence families can be a useful tool in annotating genome sequences when there is ambiguity concerning which subfamily a new homolog belongs to. Published 2000 Wiley-Liss, Inc.

Amino Acid Sequence↗

Kohonen map as a visualization tool for the analysis of protein sequences: multiple alignments, domains and segments of secondary structures.

The method of Kohonen maps, a special form of neural networks, was applied as a visualization tool for the analysis of protein sequence similarity. The procedure converts sequence (domains, aligned sequences, segments of secondary structure) into a characteristic signal matrix. This conversion depends on the property or replacement score vector selected by the user. Similar sequences have small distance in the signal space. The trained Kohonen network is functionally equivalent to an unsupervised non-linear cluster analyzer. Protein families, or aligned sequences, or segments of similar secondary structure, aggregate as clusters, and their proximity may be inspected on a color screen or on paper. Pull-down menus permit access to background information in the established text-oriented way.

Amino Acid Sequence↗

Enumerating suboptimal alignments of multiple biological sequences efficiently.

The multiple sequence alignment problem is very applicable and important in various fields in molecular biology. Because the optimal alignment that maximizes the score is not always biologically most significant, providing many suboptimal alignments as alternatives for the optimal one is very useful. As for the alignment of two sequences, this suboptimal problem is well-studied, but for the alignment of multiple sequences, it has been considered impossible to investigate such suboptimal alignments because of the enormous size of the problem. The optimal multiple alignment can be obtained with A* algorithm, and an efficient algorithm for the k shortest paths problem on general graphs is discovered recently. We extend these algorithms for computation of set of all aligned groups of residues in optimal and suboptimal alignments, and for enumeration of suboptimal alignments. The suboptimal alignments are numerous. Thus we discuss what kind of suboptimal alignment is unnecessary to enumerate, and propose an efficient technique to enumerate only necessary alignments. The practicality of these algorithms are demonstrated through experiments. Moreover, the property of suboptimal alignments of multiple sequences are also examined through experiments.

Algorithms↗

Scores for sequence searches and alignments.

Every sequence comparison method requires a set of scores. For aligning protein sequences, substitution scores are based on models of amino acid conservation and properties, and matrices of these scores have substantially improved in recent years. Position-specific scoring matrices provide representations of sequence families that are capable of detecting subtle similarities. Comprehensive evaluations can effectively guide the choice of scores for sequence alignment and searching applications, including those that aid in the prediction of protein structures.

Amino Acid Sequence↗

Molecular cloning of a human thyrotropin receptor cDNA fragment. Use of highly degenerate, inosine containing primers derived from aligned amino acid sequences of a homologous family of glycoprotein hormone receptors.

Autoantibodies to the thyrotropin (TSH) hormone receptor (TSH-R) are present in the sera of patients with thyroid autoimmune disease which are pathogenetic leading to hyperthyroidism of Graves' disease. Considerable interest has been focused on the cloning of the human TSH-R, which has until very recently, proven exceedingly difficult due to the very low receptor level expression on thyroid cells. We have used polymerase chain reaction and highly degenerate, inosine containing oligonucleotides derived from sequence alignments of the transmembrane regions 2 and 7 of a number of G-binding protein receptors including the lutropin/choriogonadotropin (LH/CG) receptors to amplify various cDNAs from human thyroid cDNA. Sequencing analysis of 27 different clones revealed that they fall into eight different groups. The very recent publication of the complete nucleotide sequence of the human TSH-R revealed that one of the groups (GT1) containing seven clones which had been sequenced belong to the human TSH-receptor. The sequence of all 7 GT1 clones was identical and in complete concordance with transmembrane regions 2 and 7 of the published TSH-R sequence. Our results show that by designing oligonucleotides to common transmembrane regions of G-binding proteins where the primers are biased in their sequence to the LH/CG receptors it is possible to amplify the TSH-R receptor sequence.

Amino Acid Sequence↗

Molecular identification of enterovirus by analyzing a partial VP1 genomic region with different methods.

VP1 is the most suitable region for use in the identification of enterovirus. Although VP1 sequencing methods may vary, it is necessary to agree on a common strategy of sequence analysis. Identification of a strain type may be achieved by three different approaches: pairwise sequence alignment, multiple-sequence alignment, and phylogenetic inference. Other methods are also available, but they are not simple enough to be performed at a virology laboratory. The performances of these methods were evaluated with nucleotide and protein sequences obtained from 32 original samples, 8 enterovirus isolates, and 64 GenBank sequences. Pairwise sequence alignment methods had very different results. The DNASTAR package identified only 28.8% of enterovirus strains, while the Genetics Computer Group package identified 50.0 or 72.1% of enterovirus strains when nucleotide or amino acid sequences were analyzed, respectively. Multiple-sequence alignment methods identified 94.2% (Clustal W program) or 92.3% (Pileup program) of the enterovirus strains, while the phylogenetic method increased this rate to 99.0%. Comparative evaluation of these analysis methods showed that the Clustal W program (version 1.81), a freely available multiple-sequence alignment program, presented one of the best performances when used with the correct criteria. Other commercial and expensive programs did not achieve the same performances, making them less suitable for molecular typing of enteroviruses. Finally, although phylogenetic inference is the most demanding method in terms of knowledge of the user, it remained the best option analyzed.

Algorithms↗

Improved alignment of nucleosome DNA sequences using a mixture model.

DNA sequences that are present in nucleosomes have a preferential approximately 10 bp periodicity of certain dinucleotide signals, but the overall sequence similarity of the nucleosomal DNA is weak, and traditional multiple sequence alignment tools fail to yield meaningful alignments. We develop a mixture model that characterizes the known dinucleotide periodicity probabilistically to improve the alignment of nucleosomal DNAs. We assume that a periodic dinucleotide signal of any type emits according to a probability distribution around a series of 'hot spots' that are equally spaced along nucleosomal DNA with 10 bp period, but with a 1 bp phase shift across the middle of the nucleosome. We model the three statistically most significant dinucleotide signals, AA/TT, GC and TA, simultaneously, while allowing phase shifts between the signals. The alignment is obtained by maximizing the likelihood of both Watson and Crick strands simultaneously. The resulting alignment of 177 chicken nucleosomal DNA sequences revealed that all 10 distinct dinucleotides are periodic, however, with only two distinct phases and varying intensity. By Fourier analysis, we show that our new alignment has enhanced periodicity and sequence identity compared with center alignment. The significance of the nucleosomal DNA sequence alignment is evaluated by comparing it with that obtained using the same model on non-nucleosomal sequences.

Algorithms↗

A comparison of scoring functions for protein sequence profile alignment.

MOTIVATION: In recent years, several methods have been proposed for aligning two protein sequence profiles, with reported improvements in alignment accuracy and homolog discrimination versus sequence-sequence methods (e.g. BLAST) and profile-sequence methods (e.g. PSI-BLAST). Profile-profile alignment is also the iterated step in progressive multiple sequence alignment algorithms such as CLUSTALW. However, little is known about the relative performance of different profile-profile scoring functions. In this work, we evaluate the alignment accuracy of 23 different profile-profile scoring functions by comparing alignments of 488 pairs of sequences with identity < or =30% against structural alignments. We optimize parameters for all scoring functions on the same training set and use profiles of alignments from both PSI-BLAST and SAM-T99. Structural alignments are constructed from a consensus between the FSSP database and CE structural aligner. We compare the results with sequence-sequence and sequence-profile methods, including BLAST and PSI-BLAST. RESULTS: We find that profile-profile alignment gives an average improvement over our test set of typically 2-3% over profile-sequence alignment and approximately 40% over sequence-sequence alignment. No statistically significant difference is seen in the relative performance of most of the scoring functions tested. Significantly better results are obtained with profiles constructed from SAM-T99 alignments than from PSI-BLAST alignments. AVAILABILITY: Source code, reference alignments and more detailed results are freely available at http://phylogenomics.berkeley.edu/profilealignment/

Algorithms↗

[Research on the recombinant plasmid pDJH2 of L. interrogans serovar lai: sequencing and alignment with other known bacterial Omp sequence].

The Leptospira whole cell vaccine (LWCV) currently used in China is safe and effective, out the immunity following vaccination with two doses of the fluid medium vaccine is of low order. The duration of immunity conferred by this vaccine is rather short, six months or at most one year. Therefore, it is necessary to develop new generation vaccines against Leptospirosis for the developing world. In this paper we report the sequencing of the insert fragment of pDJH2 from genomic DNA of L. interrogans sevovar lai strain 017 and its alignment with other bacterial omp sequences. A genomic library of Leptospira interrogaans serovar lai strain 017 was constructed with the plasmid vector pUC18. A recombinant plasmid designated pJDH2 was screened from the genomic library. Inserted fragment of pDH2 is 1.9 kb by gel electrophoresis. Immunization/protection was studied in BALB/c mice model. The results showed highly significant difference between pDJH2 and pUC18 (control). Inserted fragment of pDJH2 DNA sequencing was performed by Dr Yan Zhengxin (Max-Planck-Institut for Biology. Tubingen, Germany). Insert fragment was cloned into pBluescript II KS-(stratagene) and sequenced by using AB1 (Applied Bio Systems, Model 373A). Two open reading frames of 565 and 662 nucleotides were identified. There were identifiable initiation codons, terminators, Shine-Dalgano ribosome combining site, Pribnow boxes and Sextama boxes within the 2 sequenced regions. Nucleotide sequences were analysed using Gene Work, a suit of computer program developed by Department of Biochemistry St. Jude Children's Research Hospital Memphis. U.S.A. The results of formatted alignment showed the predicted nucleotide sequence of ORF1 of the serovar lai had significant similarity with ORF2 (49.36%). L. kirschneri ompL1 (49.26%), Borrelia burgdoferi omp (48.97%), Treponema phagedenis omp (47.3%); Salmonella typhimurium ompC(46.87%), Yersinia enterocolitica ompH (46.7%), Leptospira borgpeterseni pfap (46.3%), and Serratia marcescens omp (43.3%). The close relationship of the pDJH2 ORF1 and ORF2 nucleotide sequences from Leptospira kirschneri ompL 1 is apparent. Whether the recombinant pDJH2 will prove useful for vaccine development remains to be tested.

Animals↗

MUSCA: An Algorithm for Constrained Alignment of Multiple Data Sequences.

Given a set of N sequences, the Multiple Sequence Alignment problem is to align these N sequences, possibly with gaps, that brings out the best commonality of the N sequences. MUSCA is a two-stage approach to the alignment problem by identifying two relatively simpler sub-problems whose solutions are used to obtain the alignment of the sequences. We first discover motifs in the N sequences and then extract an appropriate subset of compatible motifs to obtain a good alignment. The motifs of interest to us are the irredundant motifs which are only polynomial in the input size. In practice, however, the number is much smaller (sub-linear). Notice that this step aids in a direct N-wise alignment, as opposed to composing the alignments from lower order (say pairwise) alignments and the solution is also independent of the order of the input sequences; hence the algorithm works very well while dealing with a large number of sequences. The second part of the problem that deals with obtaining a good alignment is solved using a graph-theoretic approach that computes an induced subgraph satisfying certain simple constraints. We reduce a version of this problem to that of solving an instance of a set covering problem, thus offer the best possible approximate solution to the problem (provided P not equalNP). Our experimental results, while being preliminary, indicate that this approach is efficient, particularly on large numbers of long sequences, and, gives good alignments when tested on biological data such as DNA and protein sequences. We introduce the the notion of an alignment number K (2 </= K </= N), a user-controlled parameter, that lends a useful flexibility to the aligning program: this additional requirement constrains the alignment to have at least K sequences agree on a character, whenever possible, in the alignment. The usefulness of the alignment number is corroborated by the users who view this as a natural constraint while dealing with a large number of sequences.

Journal Article↗

Detection of unrelated proteins in sequences multiple alignments by using predicted secondary structures.

MOTIVATION: Multiple sequence alignments are essential tools for establishing the homology relations between proteins. Essential amino acids for the function and/or the structure are generally conserved, thus providing key arguments to help in protein characterization. However for distant proteins, it is more difficult to establish, in a reliable way, the homology relations that may exist between them. In this article, we show that secondary structure prediction is a valuable way to validate protein families at low identity rate. RESULTS: We show that the analysis of the secondary structures compatibility is a reliable way to discard non-related proteins in low identity multiple alignment. AVAILABILITY: This validation is possible through our NPS@ server (http://npsa-pbil.ibcp.fr)

Algorithms↗