PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

The use of tri5 gene sequences for PCR detection and taxonomy of trichothecene-producing species in the Fusarium section Sporotrichiella.

Purified DNA from isolates of Fusarium poae, Fusarium sporotrichioides, Fusarium kyushuense and Fusarium langsethiae was used as a template to amplify a 658-bp fragment from the trichodiene synthase (tri5) gene of these fungi with the gene-specific PCR primer pair Tox5-1/Tox5-2. Fragments obtained were isolated and sequenced. DNA sequence alignments revealed high similarity between the sequences derived from F. sporotrichioides and F. langsethiae (98.7%) and less similarity between the latter species and F. poae (90.9%). Phylogenetic analysis of the aligned sequences using the tri5 sequence of Fusarium pseudograminearum as an outgroup revealed clear separation between one group consisting of F. poae and F. kyushuense and another consisting of F. sporotrichioides and F. langsethiae. The two latter species could not be distinguished phylogenetically on the basis of their tri5 sequences. Taxon-specific reverse primers were designed from the aligned sequences and combined with the tri5 gene-specific forward primer Tox5-1. The new reverse primers enabled specific amplification of a fragment of approximately 400 bp from DNA isolated from F. sporotrichioides, F. poae, F. langsethiae and F. kyushuense, respectively. All primers were tested for cross-reactivity with DNA from 26 fungal species potentially capable of producing trichothecenes. Only the primer designed for F. langsethiae cross-reacted with F. sporotrichioides. PCR assays were applied in analysis of artificially and naturally infected samples of barley and oats. On artificially infected barley, species were selectively detected by the corresponding primers. In naturally infected oats, F. langsethiae was identified by the combination of two PCR assays designed for detection of F. sporotrichioides and F. langsethiae, respectively.

Base Sequence↗

Variable gap penalty for protein sequence-structure alignment.

The penalty for inserting gaps into an alignment between two protein sequences is a major determinant of the alignment accuracy. Here, we present an algorithm for finding a globally optimal alignment by dynamic programming that can use a variable gap penalty (VGP) function of any form. We also describe a specific function that depends on the structural context of an insertion or deletion. It penalizes gaps that are introduced within regions of regular secondary structure, buried regions, straight segments and also between two spatially distant residues. The parameters of the penalty function were optimized on a set of 240 sequence pairs of known structure, spanning the sequence identity range of 20-40%. We then tested the algorithm on another set of 238 sequence pairs of known structures. The use of the VGP function increases the number of correctly aligned residues from 81.0 to 84.5% in comparison with the optimized affine gap penalty function; this difference is statistically significant according to Student's t-test. We estimate that the new algorithm allows us to produce comparative models with an additional approximately 7 million accurately modeled residues in the approximately 1.1 million proteins that are detectably related to a known structure.

Algorithms↗

Improvement of TRANSFAC matrices using multiple local alignment of transcription factor binding site sequences.

This paper describes a novel approach to constructing Position-Specific Weight Matrices (PWMs) based on the transcription factor binding site (TFBS) data provide by the TRANSFAC database and comparison of the newly generated PWMs with the original TRANSFAC matrices. Multiple local sequence alignment was performed on the TFBSs of each transcription factor. Several different alignment programs were tested and their matrices were compared to the original TRANSFAC matrices. One of the alignment programs, GLAM, produced comparable matrices in terms of the average ranking of true positive sites across the whole test set of sequences.

Algorithms↗

Analysis for free: comparing programs for sequence analysis.

Programs to import, manage and align sequences and to analyse the properties of DNA, RNA and proteins are essential for every biological laboratory. This review describes two different freeware (BioEdit and pDRAW for MS Windows) and a commercial program (Sequencher for MS Windows and Apple MacOS). Bioedit and Sequencher offer functions such as sequence alignment and editing plus reading of sequence trace files. pDRAW is a very comfortable visualisation tool with a variety of analysis functions. While Sequencher impresses with a very user-friendly interface and easy-to-use tools, BioEdit offers the largest and most customisable variety of tools. The strength of pDRAW is drawing and analysis of single sequences for priming and restriction sites and virtual cloning. It has a database function for user-specific oligonucleotides and restriction enzymes.

Base Sequence↗

Secondary structure prediction for aligned RNA sequences.

Most functional RNA molecules have characteristic secondary structures that are highly conserved in evolution. Here we present a method for computing the consensus structure of a set aligned RNA sequences taking into account both thermodynamic stability and sequence covariation. Comparison with phylogenetic structures of rRNAs shows that a reliability of prediction of more than 80% is achieved for only five related sequences. As an application we show that the Early Noduline mRNA contains significant secondary structure that is supported by sequence covariation.

Algorithms↗

Probabilistic description of protein alignments for sequences and structures.

A number of equally optimal alignments inherently exist in the sequence and structure comparisons among proteins. To represent the sub-optimal alignments systematically, we have developed a method of generating probabilistic alignments for sequences and structures, by which the correspondence between pairs of residues is evaluated in a probabilistic manner. Our method uses the periodic boundary condition to avoid the entropy artifact favoring full-length matches. In the structure comparison, the environmental effects are incorporated by the mean-field approximation. We applied this method in comparisons of two pairs of proteins with internal symmetry; the first set were proteins of TIM-barrel fold and the second were beta-trefoil fold. These pairs are expected to have distinct sub-optimal alignments suitable for probabilistic description with the periodic boundary. It was shown that the sequence and structure alignments are consistent with each other and that the alignments with the highest probability represent circular permutation.

Algorithms↗

A simplified proof of the NP- and MAX SNP-hardness of multiple sequence tree alignment.

We give a simple proof which shows that the multiple sequence tree alignment problem from molecular biology is both NP-complete and MAX SNP-hard. Our proof of MAX SNP-hardness is simpler than that given previously by Wang and Jiang. These results suggest that it is unlikely that the multiple sequence tree alignment problem has polynomial-time algorithms that produce either optimal solutions or approximate solutions whose cost may be arbitrarily close to optimal.

Algorithms↗

MOSAIC: segmenting multiple aligned DNA sequences.

UNLABELLED: MOSAIC is a set of tools for the segmentation of multiple aligned DNA sequences into homogeneous zones. The segmentation is based on the distribution of mutational events along the alignment. As an example, the analysis of one repeated sequence belonging to the subtelomeric regions of the yeast genome is presented. AVAILABILITY: Free access from ftp://ftp.biomath.jussieu.fr/pub/papers/MOSAIC

Genome, Fungal↗

BayesFold: rational 2 degrees folds that combine thermodynamic, covariation, and chemical data for aligned RNA sequences.

BayesFold is a Web application that folds an alignment of closely related sequences and evaluates hypotheses about their shared structure. It uses Bayes's Theorem to combine information from several sources, including chemical mapping (if available), thermodynamic folding, and observed sequence variations. Its method provides a rational basis for integrating results, even when these methods conflict. On a gapped alignment of 86 tRNAPhe sequences each 77 bases long, BayesFold takes 31 sec to perform the calculations; the best structure contained 95% of the base pairs in the true structure, and the true structure was ranked second. Notably, similar results come from random samples of only 10 sequences from the alignment (running time 3 sec), suggesting that remarkably few sequences are required for good results. In contrast, folding single sequences with BayesFold produced structures 9.6 bp different, or with the Vienna package, 13.4 bp different, from the true structure. Similar results were obtained for other families of tRNAs. We especially recommend BayesFold for alignments of 3-50 closely related sequences, such as the sequence families frequently found in SELEX. In addition to providing a convenient way to explore the effects of each of the criteria on the plausibility of different structures, BayesFold also makes it easy to produce publication-quality secondary-structure graphics. The Web interface, available at http://bayes.colorado.edu/fold/, includes the flexibility to thread any of the sequences (or the consensus sequence) through any of the structures, including the one judged most probable.

Algorithms↗

Amino acid similarity matrices based on force fields.

MOTIVATION: We propose a general method for deriving amino acid substitution matrices from low resolution force fields. Unlike current popular methods, the approach does not rely on evolutionary arguments or alignment of sequences or structures. Instead, residues are computationally mutated and their contribution to the total energy/score is collected. The average of these values over each position within a set of proteins results in a substitution matrix. RESULTS: Example substitution matrices have been calculated from force fields based on different philosophies and their performance compared with conventional substitution matrices. Although this can produce useful substitution matrices, the methodology highlights the virtues, deficiencies and biases of the source force fields. It also allows a rather direct comparison of sequence alignment methods with the score functions underlying protein sequence to structure threading. AVAILABILITY: Example substitution matrices are available from http://www.rsc.anu.edu.au/~zsuzsa/suppl/matrices.html. SUPPLEMENTARY INFORMATION: The list of proteins used for data collection and the optimized parameters for the alignment are given as supplementary material at http://www.rsc.anu.edu.au/~zsuzsa/suppl/matrices.html.

Amino Acid Substitution↗

GeneSilico protein structure prediction meta-server.

Rigorous assessments of protein structure prediction have demonstrated that fold recognition methods can identify remote similarities between proteins when standard sequence search methods fail. It has been shown that the accuracy of predictions is improved when refined multiple sequence alignments are used instead of single sequences and if different methods are combined to generate a consensus model. There are several meta-servers available that integrate protein structure predictions performed by various methods, but they do not allow for submission of user-defined multiple sequence alignments and they seldom offer confidentiality of the results. We developed a novel WWW gateway for protein structure prediction, which combines the useful features of other meta-servers available, but with much greater flexibility of the input. The user may submit an amino acid sequence or a multiple sequence alignment to a set of methods for primary, secondary and tertiary structure prediction. Fold-recognition results (target-template alignments) are converted into full-atom 3D models and the quality of these models is uniformly assessed. A consensus between different FR methods is also inferred. The results are conveniently presented on-line on a single web page over a secure, password-protected connection. The GeneSilico protein structure prediction meta-server is freely available for academic users at http://genesilico.pl/meta.

Internet↗

SequenceEditingAligner: a multiple sequence editor and aligner.

Here we present the SequenceEditingAligner system for editing multiple, aligned genetic sequences. This is an interactive multi-window color system that displays more than 3500 nucleotides or amino acids. The system handles nucleic acid or protein sequences with or without secondary structure data. More than 300 sequences, each more than 1500 elements in length, may be analyzed together. With the system scientists can classify elements, align sequences, edit them, find consensus patterns, and simultaneously generate oligomer frequency histograms and other statistics.

Algorithms↗

Subtilases: the superfamily of subtilisin-like serine proteases.

Subtilases are members of the clan (or superfamily) of subtilisin-like serine proteases. Over 200 subtilases are presently known, more than 170 of which with their complete amino acid sequence. In this update of our previous overview (Siezen RJ, de Vos WM, Leunissen JAM, Dijkstra BW, 1991, Protein Eng 4:719-731), details of more than 100 new subtilases discovered in the past five years are summarized, and amino acid sequences of their catalytic domains are compared in a multiple sequence alignment. Based on sequence homology, a subdivision into six families is proposed. Highly conserved residues of the catalytic domain are identified, as are large or unusual deletions and insertions. Predictions have been updated for Ca(2+)-binding sites, disulfide bonds, and substrate specificity, based on both sequence alignment and three-dimensional homology modeling.

Amino Acid Sequence↗

Aligning two sequences within a specified diagonal band.

We describe an algorithm for aligning two sequences within a diagonal band that requires only O(NW) computation time and O(N) space, where N is the length of the shorter of the two sequences and W is the width of the band. The basic algorithm can be used to calculate either local or global alignment scores. Local alignments are produced by finding the beginning and end of a best local alignment in the band, and then applying the global alignment algorithm between those points. This algorithm has been incorporated into the FASTA program package, where it has decreased the amount of memory required to calculate local alignments from O(NW) to O(N) and decreased the time required to calculate optimized scores for every sequence in a protein sequence database by 40%. On computers with limited memory, such as the IBM-PC, this improvement both allows longer sequences to be aligned and allows optimization within wider bands, which can include longer gaps.

Algorithms↗

Sigma: multiple alignment of weakly-conserved non-coding DNA sequence.

BACKGROUND: Existing tools for multiple-sequence alignment focus on aligning protein sequence or protein-coding DNA sequence, and are often based on extensions to Needleman-Wunsch-like pairwise alignment methods. We introduce a new tool, Sigma, with a new algorithm and scoring scheme designed specifically for non-coding DNA sequence. This problem acquires importance with the increasing number of published sequences of closely-related species. In particular, studies of gene regulation seek to take advantage of comparative genomics, and recent algorithms for finding regulatory sites in phylogenetically-related intergenic sequence require alignment as a preprocessing step. Much can also be learned about evolution from intergenic DNA, which tends to evolve faster than coding DNA. Sigma uses a strategy of seeking the best possible gapless local alignments (a strategy earlier used by DiAlign), at each step making the best possible alignment consistent with existing alignments, and scores the significance of the alignment based on the lengths of the aligned fragments and a background model which may be supplied or estimated from an auxiliary file of intergenic DNA. RESULTS: Comparative tests of sigma with five earlier algorithms on synthetic data generated to mimic real data show excellent performance, with Sigma balancing high "sensitivity" (more bases aligned) with effective filtering of "incorrect" alignments. With real data, while "correctness" can't be directly quantified for the alignment, running the PhyloGibbs motif finder on pre-aligned sequence suggests that Sigma's alignments are superior. CONCLUSION: By taking into account the peculiarities of non-coding DNA, Sigma fills a gap in the toolbox of bioinformatics.

Algorithms↗

UTR reconstruction and analysis using genomically aligned EST sequences.

Untranslated regions (UTR) play important roles in the posttranscriptional regulation of mRNA processing. There is a wealth of UTR-related information to be mined from the rapidly accumulating EST collections. A computational tool, UTR-extender, has been developed to infer UTR sequences from genomically aligned ESTs. It can completely and accurately reconstruct 72% of the 3' UTRs and 15% of the 5' UTRs when tested using 908 functionally cloned transcripts. In addition, it predicts extensions for 11% of the 5' UTRs and 28% of the 3' UTRs. These extension regions are validated by examining splicing frequencies and conservation levels. We also developed a method called polyadenylation site scan (PASS) to precisely map polyadenylation sites in human genomic sequences. A PASS analysis of 908 genic regions estimates that 40-50% of human genes undergo alternative polyadenylation. Using EST redundancy to assess expression levels, we also find that genes with short 3' UTRs tend to be highly expressed.

Algorithms↗

[Membrane probability profile construction based on amino acids sequences multiple alignment].

Prediction of membrane segments in sequences of membrane proteins is well known and important problem. Accuracy of the solution of this problem by methods that don't use homology search in additional data bank can be improved. There is a lack of testing data in this area because of small amount of real structures of membrane proteins. In this work, we create a testing set of structural alignments of membrane proteins, in which positioning of the membrane segments reflects agreement of known 3D-structures of proteins in the alignment. We propose a method for predicting position of membrane segments in multiple alignment based on forward-backward algorithm from HMM theory. This method not only allows to predict positions of membrane segments but also forms probability membrane profile, which can be used in multiple alignment methods that take into account secondary structure information about sequences. Method is implemented in computer program available on the World-Wide Web site http://bioinf.fbb.msu.ru/fwdbck/. Proposed method provides results better than MEMSAT method, which is nearly only tool for prediction of membrane segments in multiple alignments without additional homology search.

Animals↗

INTERALIGN: interactive alignment editor for distantly related protein sequences.

SUMMARY: Improving and ascertaining the quality of a multiple sequence alignment is a very challenging step in protein sequence analysis. This is particularly the case when dealing with sequences in the 'twilight zone', i.e. sharing < 30% identity. Here we describe INTERALIGN, a dedicated user-friendly alignment editor including a view of secondary structures and a synchronized display of carbon alpha traces of corresponding protein structures. Profile alignment, using CLUSTALW, is implemented to improve the alignment of a sequence of unknown structure with the visually optimized structural alignment as compared with a standard multiple sequence alignment. Tree-based ordering further helps in identifying the structure closest to a given sequence.

Algorithms↗