PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

PROANAL version 2: multifunctional program for analysis of multiple protein sequence alignments and for studying the structure--activity relationships in protein families.

A new version of the program PROANAL is described. A multiple linear regression analysis of the protein structure--activity relationship allows one to investigate the combinations of protein sites and factors influencing the activity. The program also provides the possibility to seek out protein sites, conservative or variable in variations of physicochemical characteristics, and regions with high or low values of these characteristics. PROANAL2 may be useful in the simulation of protein-engineering experiments and in the search of a number of protein regions such as functional sites, secondary structures, solvent-exposed regions, T- and B-cell antigenic determinants, etc.

Algorithms↗

A statistical theory of sequence alignment with gaps.

A statistical theory of local alignment algorithms with gaps is presented. Both the linear and logarithmic phases, as well as the phase transition separating the two phases, are described in a quantitative way. Markov sequences without mutual correlations are shown to have scale-invariant alignment statistics. Deviations from scale invariance indicate the presence of mutual correlations detectable by alignment algorithms. Conditions are obtained for the optimal detection of a class of mutual sequence correlations.

Algorithms↗

Ribosomal RNA as molecular barcodes: a simple correlation analysis without sequence alignment.

MOTIVATION: We explored the feasibility of using unaligned rRNA gene sequences as DNA barcodes, based on correlation analysis of composition vectors (CVs) derived from nucleotide strings. We tested this method with seven rRNA (including 12, 16, 18, 26 and 28S) datasets from a wide variety of organisms (from archaea to tetrapods) at taxonomic levels ranging from class to species. RESULT: Our results indicate that grouping of taxa based on CV analysis is always in good agreement with the phylogenetic trees generated by traditional approaches, although in some cases the relationships among the higher systemic groups may differ. The effectiveness of our analysis might be related to the length and divergence among sequences in a dataset. Nevertheless, the correct grouping of sequences and accurate assignment of unknown taxa make our analysis a reliable and convenient approach in analyzing unaligned sequence datasets of various rRNAs for barcoding purposes. AVAILABILITY: The newly designed software (CVTree 1.0) is publicly available at the Composition Vector Tree (CVTree) web server http://cvtree.cbi.pku.edu.cn.

Algorithms↗

MAVL/StickWRLD: analyzing structural constraints using interpositional dependencies in biomolecular sequence alignments.

The increasing availability of structurally aligned protein families has made it possible to use statistical methods to discover regions of interpositional dependencies of residue identity. Such dependencies amongst residues often have structural or functional implications, and their discovery can supply valuable constraints that assist in the refinement of measured, or predicted molecular structure assignments. Multiple Alignment Variation Linker (MAVL) and StickWRLD [W. Ray (2004) Nucleic Acids Res., 32, W59-W63] were developed to analyze and visualize nucleic acid and protein alignments, to discover and illuminate position/location relationships to the user. The original system analyzed users' data from a web-form submission and presented the user with a static VRML diagram describing their data. We are pleased to report that MAVL/StickWRLD has been completely redesigned and rewritten. MAVL/StickWRLD now functions as a platform-independent Java applet, with real-time dynamic controls that enable much more intuitive exploration and interaction with the data. The system has also been upgraded to enable visualization of a range of aggregate residue properties, and an extensive database of pre-computed StickWRLD diagrams based on PFAM families is now available directly from the interface. The Java StickWRLD applet is available via the WWW at http://www.microbial-pathogenesis.org/stickwrld/.

Computer Graphics↗

Gene recognition via spliced sequence alignment.

Gene recognition is one of the most important problems in computational molecular biology. Previous attempts to solve this problem were based on statistics, and applications of combinatorial methods for gene recognition were almost unexplored. Recent advances in large-scale cDNA sequencing open a way toward a new approach to gene recognition that uses previously sequenced genes as a clue for recognition of newly sequenced genes. This paper describes a spliced alignment algorithm and software tool that explores all possible exon assemblies in polynomial time and finds the multiexon structure with the best fit to a related protein. Unlike other existing methods, the algorithm successfully recognizes genes even in the case of short exons or exons with unusual codon usage; we also report correct assemblies for genes with more than 10 exons. On a test sample of human genes with known mammalian relatives, the average correlation between the predicted and actual proteins was 99%. The algorithm correctly reconstructed 87% of genes and the rare discrepancies between the predicted and real exon-intron structures were caused either by short (less than 5 amino acids) initial/terminal exons or by alternative splicing. Moreover, the algorithm predicts human genes reasonably well when the homologous protein is nonvertebrate or even prokaryotic. The surprisingly good performance of the method was confirmed by extensive simulations: in particular, with target proteins at 160 accepted point mutations (PAM) (25% similarity), the correlation between the predicted and actual genes was still as high as 95%.

Algorithms↗

A new interactive protein sequence alignment program and comparison of its results with widely used algorithms.

A computer program that allows interactive sequence comparison is described. It graphically displays a search matrix using residue physiochemical characteristics and multilength segmental comparisons. The user selects through a mousing device and screen pointer the sequence spans to be matched. The results of this method are compared with those of ALIGN and BESTFIT.

Algorithms↗

An analysis of the periodicity of conserved residues in sequence alignments of G-protein coupled receptors. Implications for the three-dimensional structure.

Twenty-three sequences from the family of G-protein coupled receptors have been aligned according to the 'historical alignment' procedure of Feng and Doolittle. Fourier transform analysis of this reveals that parts of five of the seven putative membrane-spanning regions exhibit a periodicity of conserved/nonconserved residues which is compatible with the periodicity of the alpha-helix. This would place the conserved residues on one side of the helix, which may face the inside of the proposed seven membered helical bundle.

Amino Acid Sequence↗

Pairwise sequence alignment using a PROSITE pattern-derived similarity score.

Existing methods for alignments are based on edition costs computed additionally position by position, according to a fixed substitution matrix: a substitution always has the same weight regardless of the position. Nevertheless the biologist favours a similarity according to his knowledge of the structure or the function of the sequences considered. In the particular case of proteins, we present a method consisting in integrating other information, such as patterns of the PROSITE databank, in the classical dynamic programming algorithm. The method consists in making an alignment by dynamic programming taking a decision not only letter by letter as in the Smith & Waterman algorithm but also by giving a reward when aligning patterns.

Algorithms↗

Computational complexity of multiple sequence alignment with SP-score.

It is shown that the multiple alignment problem with SP-score is NP-hard for each scoring matrix in a broad class M that includes most scoring matrices actually used in biological applications. The problem remains NP-hard even if sequences can only be shifted relative to each other and no internal gaps are allowed. It is also shown that there is a scoring matrix M(0) such that the multiple alignment problem for M(0) is MAX-SNP-hard, regardless of whether or not internal gaps are allowed.

Algorithms↗

The PCR-SSP Manager computer program: a tool for maintaining sequence alignments and automatically updating the specificities of PCR-SSP primers and primer mixes.

An emerging problem of molecular typing methods such as PCR amplification using sequence-specific primers (PCR-SSP) is that they frequently require updating as new alleles are constantly being described which potentially affect the specificity of every PCR-SSP reaction. PCR-SSP uses pairs of primers to detect cis-linked polymorphisms and thus each new allele described must be compared to each individual primer pair. Furthermore, sequence homology between the various loci for class I and class II means that, for example, new HLA-A sequences have to be compared with HLA-B and HLA-C primer mixes to rule out cross-locus amplification. We have developed a computer program known as SSP Manager which is capable of aligning HLA class I and class II sequences obtained from Internet-accessible databases such as GenBank. The program then updates all individual primer specificities held in its database before updating the specificities of all primer mixes. Sets of primer mixes can then be combined from the primer mix directory to create PCR-SSP typing trays which are subsequently analysed by the program. A report is generated which stipulates whether all known sequences are amplified and the reason for apparent failure to test for individual alleles, e.g. a lack of relevant sequence information. SSP Manager has the flexibility to cope with unusual sequences (deletions and insertions), primers with internal mismatches and primers with a deliberate mismatch. The program also has many tools for developing new primer mixes, such as the facility to search for novel reactions using Boolean operators. The organisation and operational use of the SSP Manager program is described and its uses are illustrated with an updated allele list for our previously described Phototyping PCR-SSP class I and class II typing set. The SSP Manager is available on request from the authors.

Antisense Elements (Genetics)↗

A measure of the similarity of sets of sequences not requiring sequence alignment.

Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classifications. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism--for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.

Animals↗

Multiple sequence alignment with user-defined anchor points.

BACKGROUND: Automated software tools for multiple alignment often fail to produce biologically meaningful results. In such situations, expert knowledge can help to improve the quality of alignments. RESULTS: Herein, we describe a semi-automatic version of the alignment program DIALIGN that can take pre-defined constraints into account. It is possible for the user to specify parts of the sequences that are assumed to be homologous and should therefore be aligned to each other. Our software program can use these sites as anchor points by creating a multiple alignment respecting these constraints. This way, our alignment method can produce alignments that are biologically more meaningful than alignments produced by fully automated procedures. As a demonstration of how our method works, we apply our approach to genomic sequences around the Hox gene cluster and to a set of DNA-binding proteins. As a by-product, we obtain insights about the performance of the greedy algorithm that our program uses for multiple alignment and about the underlying objective function. This information will be useful for the further development of DIALIGN. The described alignment approach has been integrated into the TRACKER software system.

Journal Article↗