PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

GALA, a database for genomic sequence alignments and annotations.

We have developed a relational database to contain whole genome sequence alignments between human and mouse with extensive annotations of the human sequence. Complex queries are supported on recorded features, both directly and on proximity among them. Searches can reveal a wide variety of relationships, such as finding all genes expressed in a designated tissue that have a highly conserved noncoding sequence 5' to the start site. Other examples are finding single nucleotide polymorphisms that occur in conserved noncoding regions upstream of genes and identifying CpG islands that overlap the 5' ends of divergently transcribed genes. The database is available online at http://globin.cse.psu.edu/ and http://bio.cse.psu.edu/.

5' Untranslated Regions↗

The correlation error and finite-size correction in an ungapped sequence alignment.

MOTIVATION: The BLAST program for comparing two sequences assumes independent sequences in its random model. The resulting random alignment matrices have correlations across their diagonals. Analytic formulas for the BLAST p-value essentially neglect these correlations and are equivalent to a random model with independent diagonals. Progress on the independent diagonals model has been surprisingly rapid, but the practical magnitude of the correlations it neglects remains unknown. In addition, BLAST uses a finite-size correction that is particularly important when either of the sequences being compared is short. Several formulas for the finite-size correction have now been given, but the corresponding errors in the BLAST p-values have not been quantified. As the lengths of compared sequences tend to infinity, it is also theoretically unknown whether the neglected correlations vanish faster than the finite-size correction. RESULTS: Because we required certain analytic formulas, our study restricted its computer experiments to ungapped sequence alignment. We expect some of our conclusions to extend qualitatively to gapped sequence alignment, however. With this caveat, the finite-size correction appeared to vanish faster than the neglected correlations. Although the finite-size correction underestimated the BLAST p-value, it improved the approximation substantially for all but very short sequences. In practice, the Altschul-Gish finite-size correction was superior to Spouge's. The independent diagonals model was always within a factor of 2 of the true BLAST p-value, although fitting p-value parameters from it probably is unwise. CONTACT: spouge@ncbi.nlm.nih.gov

Database Management Systems↗

Integrated tools for structural and sequence alignment and analysis.

We have developed new computational methods for displaying and analyzing members of protein superfamilies. These methods (MinRMS, AlignPlot and MSFviewer) integrate sequence and structural information and are implemented as separate but cooperating programs to our Chimera molecular modeling system. Integration of multiple sequence alignment information and three-dimensional structural representations enable researchers to generate hypotheses about the sequence-structure relationship. Structural superpositions can be generated and easily tuned to identify similarities around important characteristics such as active sites or ligand binding sites. Information related to the release of Chimera, MinRMS, AlignPlot and MSFviewer can be obtained at http:¿www.cgl.ucsf.edu/chimera.

Amino Acid Sequence↗

Towards an automatic method of predicting protein structure by homology: an evaluation of suboptimal sequence alignments.

A major problem in predicting protein structure by homology modelling is that the sequence alignment from which the model is built may not be the best one in terms of the correct equivalencing of residues assessed by structural or functional criteria. A useful strategy is to generate and examine a number of suboptimal alignments as better alignments can often be found away from the optimal. A procedure to filter rapidly suboptimal alignments based on measurement of core volumes and packing pair potentials is investigated. The approach is benchmarked on three pairs of sequences which are non-trivial to align correctly, namely two immunoglobulin domains, plastocyanin with azurin and two distant globin sequences. It is shown to be useful to reduce a large ensemble of possible alignments down to a few which correspond more closely to the correct (structure based) alignment.

Algorithms↗

Effects of sequence alignment and structural domains of ribosomal DNA on phylogeny reconstruction for the protozoan family sarcocystidae.

Finding correct species relationships using phylogeny reconstruction based on molecular data is dependent on several empirical and technical factors. These include the choice of DNA sequence from which phylogeny is to be inferred, the establishment of character homology within a sequence alignment, and the phylogeny algorithm used. Nevertheless, sequencing and phylogeny tools provide a way of testing certain hypotheses regarding the relationship among the organisms for which phenotypic characters demonstrate conflicting evolutionary information. The protozoan family Sarcocystidae is one such group for which molecular data have been applied phylogenetically to resolve questionable relationships. However, analyses carried out to date, particularly based on small-subunit ribosomal DNA, have not resolved all of the relationships within this family. Analysis of more than one gene is necessary in order to obtain a robust species signal, and some DNA sequences may not be appropriate in terms of their phylogenetic information content. With this in mind, we tested the informativeness of our chosen molecule, the large-subunit ribosomal DNA (lsu rDNA), by using subdivisions of the sequence in phylogenetic analysis through PAUP, fastDNAml, and neighbor joining. The segments of sequence applied correspond to areas of higher nucleotide variation in a secondary-structure alignment involving 21 taxa. We found that subdivision of the entire lsu rDNA is inappropriate for phylogenetic analysis of the Sarcocystidae. There are limited informative nucleotide sites in the lsu rDNA for certain clades, such as the one encompassing the subfamily Toxoplasmatinae. Consequently, the removal of any segment of the alignment compromises the final tree topology. We also tested the effect of using two different alignment procedures (CLUSTAL W and the structure alignment using DCSE) and three different tree-building methods on the final tree topology. This work shows that congruence between different methods in the formation of clades may be a feature of robust topology; however, a sequence alignment based on primary structure may not be comparing homologous nucleotides even though the expected topology is obtained. Our results support previous findings showing the paraphyly of the current genera Sarcocystis and Hammondia and again bring to question the relationships of Sarcocystis muris, Isospora felis, and Neospora caninum. In addition, results based on phylogenetic analysis of the structure alignment suggest that Sarcocystis zamani and Sarcocystis singaporensis, which have reptilian definitive hosts, are monophyletic with Sarcocystis species using mammalian definitive hosts if the genus Frenkelia is synonymized with Sarcocystis.

Animals↗

Visualization of near-optimal sequence alignments.

MOTIVATION: Mathematically optimal alignments do not always properly align active site residues or well-recognized structural elements. Most near-optimal sequence alignment algorithms display alternative alignment paths, rather than the conventional residue-by-residue pairwise alignment. Typically, these methods do not provide mechanisms for finding effectively the most biologically meaningful alignment in the potentially large set of options. RESULTS: We have developed Web-based software that displays near optimal or alternative alignments of two protein or DNA sequences as a continuous moving picture. A WWW interface to a C++ program generates near optimal alignments, which are sent to a Java Applet, which displays them in a series of alignment frames. The Applet aligns residues so that consistently aligned regions remain at a fixed position on the display, while variable regions move. The display can be stopped to examine alignment details.

Algorithms↗

On the complexity of multiple sequence alignment.

We study the computational complexity of two popular problems in multiple sequence alignment: multiple alignment with SP-score and multiple tree alignment. It is shown that the first problem is NP-complete and the second is MAX SNP-hard. The complexity of tree alignment with a given phylogeny is also considered.

Algorithms↗

Effects of sequence alignment on the phylogeny of Sarcocystis deduced from 18S rDNA sequences.

The family Sarcocystidae contains a wide variety of parasitic protozoa, some of which are important pathogens of livestock and humans. The taxonomic relationships between two of the genera in this family (Toxoplasma and Sarcocystis) have been debated for a number of years and remain controversial. Recent studies, from comparisons of 18S rDNA-sequence data, have suggested that Sarcocystis is paraphyletic, although a hypothesis supporting monophyly of Sarcocystis could not be rejected. The present study shows that the phylogenetically informative nucleotide positions within the 18S rDNA are primarily located in the regions that make up the helices in the secondary structure of the 18S rRNA. A phylogenetic analysis of 18S rDNA-sequence data aligned by secondary structure constraints, or a subset of the data corresponding to all nucleotides found in the helices, provide unambiguous evidence supporting monophyly of Sarcocystis.

Animals↗

Match-Box_server: a multiple sequence alignment tool placing emphasis on reliability.

MOTIVATION: The Match-Box software comprises protein sequence alignment tools based on strict statistical thresholds of similarity between protein segments. The method circumvents the gap penalty requirement: gaps being the result of the alignment and not a governing parameter of the procedure. The reliable conserved regions outlined by Match-Box are particularly relevant for homology modelling of protein structures, prediction of essential residues for site-directed mutagenesis and oligonucleotide design for cloning homologous genes by polymerase chain reaction (PCR). RESULTS: The method produces reliable results, as assessed by tests performed on protein families of known structures and of low sequence similarity. A reliability score is computed in relation to a threshold of similarity progressively raised to extend the aligned regions to their maximal length, up to the significance limit of matching segments. The score obtained at each position is printed below the sequences and allows a discriminant reading of each aligned region. AVAILABILITY: Sequences may be submitted to a Web server at http://www.fundp.ac.be/sciences/biologie/bms/+ ++matchbox_submit.html or sent by e-mail to matchbox/biq.fundp.ac.be (help available by just mailing help).

Algorithms↗

Structure-based sequence alignment of elongation factors Tu and G with related GTPases involved in translation.

The G domain and domain II in the crystal structure of Thermus thermophilus elongation factor G (EF-G) were compared with the homologous domains in Thermus aquaticus elongation factor Tu (EF-Tu). Sequence alignment derived from the structural superposition was used to define conserved sequence elements in domain II. These elements and previously known conserved sequence elements in the G domain were used to guide the alignment of the sequences of Sulfolobus acidocaldarius elongation factor 2, human elongation factor 2, and Escherichia coli initiation factor 2 and release factor 3 to the aligned sequences of EF-G and EF-Tu. This alignment, which deviates from previously published alignments, has evolutionary implications and leads to alternative interpretations of biochemical data concerning the interaction of elongation factors with the alpha-sarcin/ricin region of the ribosome. A single conserved sequence motif in domain II was identified and used to further characterize the GTPase subfamily of translation factors and related proteins. It was shown that the motif is found in most if not all the members of the family. Apparently, the common characteristic of these GTPases is an extensive consensus structural unit that possibly accounts for a similar interaction with the ribosome and is composed of two domains homologous to the G domain and domain II in EF-Tu and EF-G.

Amino Acid Sequence↗

Multiple sequence alignment with user-defined constraints at GOBICS.

Most multi-alignment methods are fully automated, i.e. they are based on a fixed set of mathematical rules. For various reasons, such methods may fail to produce biologically meaningful alignments. Herein, we describe a semi-automatic approach to multiple sequence alignment where biological expert knowledge can be used to influence the alignment procedure. The user can specify parts of the sequences that are biologically related to each other; our software program uses these sites as anchor points and creates a multiple alignment respecting these user-defined constraints. By using known functionally, structurally or evolutionarily related positions of the input sequences as anchor points, our method can produce alignments that reflect the true biological relationships among the input sequences more accurately than fully automated procedures can do.

Algorithms↗

Combining transcriptome data with genomic and cDNA sequence alignments to make confident functional assignments for Aspergillus nidulans genes.

Whole genome sequencing of several filamentous ascomycetes is complete or in progress; these species, such as Aspergillus nidulans, are relatives of Saccharomyces cerevisiae. However, their genomes are much larger and their gene structure more complex, with genes often containing multiple introns. Automated annotation programs can quickly identify open reading frames for hypothetical genes, many of which will be conserved across large evolutionary distances, but further information is required to confirm functional assignments. We describe a comparative and functional genomics approach using sequence alignments and gene expression data to predict the function of Aspergillus nidulans genes. By highlighting examples of discrepancies between the automated genome annotation and cDNA or EST sequencing, we demonstrate that the greater complexity of gene structure in filamentous fungi demands independent data on gene expression and the gene sequence be used to make confident functional assignments.

Aspergillus nidulans↗

The HSSP database of protein structure-sequence alignments.

HSSP (homology-derived structures of proteins) is a derived database merging structural (2-D and 3-D) and sequence information (1-D). For each protein of known 3D structure from the Protein Data Bank, the database has a file with all sequence homologues, properly aligned to the PDB protein. Homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of sequence aligned sequence families, but it is also a database of implied secondary and tertiary structures.

Amino Acid Sequence↗

Sequence alignment kernel for recognition of promoter regions.

UNLABELLED: In this paper we propose a new method for recognition of prokaryotic promoter regions with startpoints of transcription. The method is based on Sequence Alignment Kernel, a function reflecting the quantitative measure of match between two sequences. This kernel function is further used in Dual SVM, which performs the recognition. Several recognition methods have been trained and tested on positive data set, consisting of 669 sigma70-promoter regions with known transcription startpoints of Escherichia coli and two negative data sets of 709 examples each, taken from coding and non-coding regions of the same genome. The results show that our method performs well and achieves 16.5% average error rate on positive & coding negative data and 18.6% average error rate on positive & non-coding negative data. AVAILABILITY: The demo version of our method is accessible from our website http://mendel.cs.rhul.ac.uk/

Algorithms↗

Conservation of structure detected in two trypanosome surface glycoproteins by amino acid sequence alignment.

The predominant molecule exposed to antibody on the surface of Trypanosoma brucei is a glycoprotein of about 60 000 molecular weight which varies in amino acid sequence. The complete sequences of two such variable surface glycoproteins (VSGs) from randomly isolated, different antigenic types of trypanosomes were compared by amino acid sequence alignment. Homologous sequences were found distributed over various regions of the VSGs. Particularly good homology was observed between residues 16-34, 91-115, 177-194 and 254-345 from the N-terminus, in addition to the known conserved region close to the C-terminus. Homology was also demonstrated in the corresponding regions of the cDNA sequences by matrix analysis.

Amino Acid Sequence↗

Evolutionary profiles from the QR factorization of multiple sequence alignments.

We present an algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of the homologous group. The method, based on the multidimensional QR factorization of numerically encoded multiple sequence alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. We observe a general trend that these smaller, more evolutionarily balanced profiles have comparable and, in many cases, better performance in database searches than conventional profiles containing hundreds of sequences, constructed in an iterative and computationally intensive procedure. For more diverse families or superfamilies, with sequence identity <30%, structural alignments, based purely on the geometry of the protein structures, provide better alignments than pure sequence-based methods. Merging the structure and sequence information allows the construction of accurate profiles for distantly related groups. These structure-based profiles outperformed other sequence-based methods for finding distant homologs and were used to identify a putative class II cysteinyl-tRNA synthetase (CysRS) in several archaea that eluded previous annotation studies. Phylogenetic analysis showed the putative class II CysRSs to be a monophyletic group and homology modeling revealed a constellation of active site residues similar to that in the known class I CysRS.

Algorithms↗

Gaps in structurally similar proteins: towards improvement of multiple sequence alignment.

An algorithm was developed to locally optimize gaps from the FSSP database. Over 2 million gaps were identified from all versus all FSSP structure comparisons, and datasets of non-identical gaps and flanking regions comprising between 90,000 and 135,000 sequence fragments were extracted for statistical analysis. Relative to background frequencies, gaps were enriched in residue types with small side chains and high turn propensity (D, G, N, P, S), and were depleted in residue types with hydrophobic side chains (C, F, I, L, V, W, Y). In contrast, regions flanking a gap exhibited opposite trends in amino acid frequencies, i.e., enrichment in hydrophobic residues and a high degree of secondary structure. Log-odds scores of residue type as a function of position in or around a gap were derived from the statistics. Three simple experiments demonstrated that these scores contained significant predictive information. First, regions where gaps were observed in single sequences taken from HOMSTRAD structure-based multiple sequence alignments generally scored higher than regions where gaps were not observed. Second, given the correct pairwise-aligned cores, the actual positions of gaps could be reproduced from sequence more accurately using the structurally-derived statistics than by using random pairwise alignments. Finally, revision of the Clustal-W residue-specific gap opening parameters with this new information improved the agreement of Clustal-W alignments with the structure-based alignments. At least three applications for these results are envisioned: improvement of gap penalties in pairwise (or multiple) sequence alignment, prediction of regions of single sequences likely (or unlikely) to contain indels, and more accurate placement of gaps in automated pairwise structure alignment.

Algorithms↗