PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Prediction of genetic structure in eukaryotic DNA using reference point logistic regression and sequence alignment.

MOTIVATION: Current software tools are moderately effective in predicting genetic structure (exons, introns, intergenic regions, and complete genes) from raw DNA sequence data. Improvements in accuracy and speed are needed to deal with the increasing volume of data from large scale sequencing projects. RESULTS: We present a two-stage computer program to predict genetic structure in eukaryotic DNA. The first stage makes use of a novel statistical technique, called reference point logistic (RPL) regression, to calculate scores for potential functional sites. These site scores are combined with interval content, length, and state scores, via a Generalized Hidden Markov Model, to determine a combined score for each possible parse of a given DNA sequence into exons, introns, and intergenic regions. An optimal parse is found using a dynamic programming algorithm. In the second stage, protein sequence alignment methods are applied to improve the accuracy of the initial parse. Computation in the first stage of the program is very fast (1 s on a 360 MHz CPU for a 16 kb sequence) and its predictive accuracy typically matches or exceeds the best results reported for other methods (Sensitivity = 0.93 and Specificity = 0.93 for the Burset/Guigótest set). Computation in the second stage is slower, but the final predictions are more accurate (Sn = 0.97, Sp = 0.97). The program (called GRPL) can handle partial, single, and multi-gene sequences. The program is also capable of predicting the genetic structure of vertebrate, invertebrate, and plant DNA with nearly equal accuracy. Statistical techniques have also been introduced to model the effects of varying C+G content in a continuous manner and to control overfitting of parameters for smaller training sets. AVAILABILITY: An academic implementation of GRPL, compiled for SUN workstations, is available by anonymous ftp from snipe.pharmacy. ualberta.ca/pub. The training and test sets used in this work, together with supplementary material, can be found at the same location. A commercial implementation is available as a component of GeneTool (BioTools Inc., http://biotools.com).

Algorithms↗

On the dependence structure of sequence alignment scores calculated with multiple scoring matrices.

A common practice in protein sequence alignment is to try several scoring matrices until "something interesting'' is found. This leads to a multiple testing problem making p- and E-values hard to interpret. We focus on local alignment and propose to use logistic copula functions to model explicitly the dependence structure of scores obtained using different scoring matrices. By doing this, we obtain p-value correction factors when using more than one scoring matrix on the same sequences. Furthermore the parameter of the logistic copula can be interpreted as measure of dependence, providing insight concerning the relatedness of the scores from different matrices.

Journal Article↗

Evolution of the cytochrome P450 superfamily: sequence alignments and pharmacogenetics.

The evolution of the cytochrome P450 (CYP) superfamily is described, with particular reference to major events in the development of biological forms during geological time. It is noted that the currently accepted timescale for the elaboration of the P450 phylogenetic tree exhibits close parallels with the evolution of terrestrial biota. Indeed, the present human P450 complement of xenobiotic-metabolizing enzymes may have originated from coevolutionary 'warfare' between plants and animals during the Devonian period about 400 million years ago. A number of key correspondences between the evolution of P450 system and the course of biological development over time, point to a mechanistic molecular biology of evolution which is consistent with a steady increase in atmospheric oxygenation beginning over 2000 million years ago, whereas dietary changes during more recent geological time may provide one possible explanation for certain species differences in metabolism. Alignment between P450 protein sequences within the same family or subfamily, together with across-family comparisons, aid the rationalization of drug metabolism specificities for different P450 isoforms, and can assist in an understanding of genetic polymorphisms in P450-mediated oxidations at the molecular level. Moreover, the variation in P450 regulatory mechanisms and inducibilities between different mammalian species are likely to have important implications for current procedures of chemical safety evaluation, which rely on pure genetic strains of laboratory bred rodents for the testing of compounds destined for human exposure.

Amino Acid Sequence↗

FootPrinter3: phylogenetic footprinting in partially alignable sequences.

FootPrinter3 is a web server for predicting transcription factor binding sites by using phylogenetic footprinting. Until now, phylogenetic footprinting approaches have been based either on multiple alignment analysis (e.g. PhyloVista, PhastCons), or on motif-discovery algorithms (e.g. FootPrinter2). FootPrinter3 integrates these two approaches, making use of local multiple sequence alignment blocks when those are available and reliable, but also allowing finding motifs in unalignable regions. The result is a set of predictions that joins the advantages of alignment-based methods (good specificity) to those of motif-based methods (good sensitivity, even in the presence of highly diverged species). FootPrinter3 is thus a tool of choice to exploit the wealth of vertebrate genomes being sequenced, as it allows taking full advantage of the sequences of highly diverged species (e.g. chicken, zebrafish), as well as those of more closely related species (e.g. mammals). The FootPrinter3 web server is available at: http://www.mcb.mcgill.ca/~blanchem/FootPrinter3.

Animals↗

In unison: regularization of protein secondary structure predictions that makes use of multiple sequence alignments.

We present a method whose purpose is to post-process the fuzzy results of secondary structure prediction methods that use multiple sequence alignments, in order to obtain 'realistic' secondary structures, i.e., secondary structure elements whose length is greater than or equal to some predefined minimum length. This regularization helps with interpretation of the secondary structure prediction.

Algorithms↗

An algorithm to determine protein sequence alignment by utilizing data obtained from a peptide mixture and individual peptides.

With the aim of limiting peptide purification steps and unambiguously ascertaining protein sequences, we have designed and implemented on a personal computer an algorithm to determine sequence alignment by utilizing data obtained from automatic Edman degradation performed on a single peptide mixture and individual peptides. The protein under study is digested by two different hydrolysis methods and fragments are just isolated from one mixture and sequenced, while the second mixture is submitted unfractionated to sequence analysis. The algorithm provides for the exact alignment of the individual peptides using the mixture data for the overlapping. We report an example of application of this approach by utilizing experimental data obtained from a protein of known sequence.

Algorithms↗

Dynamic use of multiple parameter sets in sequence alignment.

The level of conservation between two homologous sequences often varies among sequence regions; functionally important domains are more conserved than the remaining regions. Thus, multiple parameter sets should be used in alignment of homologous sequences with a stringent parameter set for highly conserved regions and a moderate parameter set for weakly conserved regions. We describe an alignment algorithm to allow dynamic use of multiple parameter sets with different levels of stringency in computation of an optimal alignment of two sequences. The algorithm dynamically considers various candidate alignments, partitions each candidate alignment into sections, and determines the most appropriate set of parameter values for each section of the alignment. The algorithm and its local alignment version are implemented in a computer program named GAP4. The local alignment algorithm in GAP4, that in its predecessor GAP3, and an ordinary local alignment program SIM were evaluated on 257,716 pairs of homologous sequences from 100 protein families. On 168,475 of the 257,716 pairs (a rate of 65.4%), alignments from GAP4 were more statistically significant than alignments from GAP3 and SIM.

Algorithms↗

Comprehensive aligned sequence construction for automated design of effective probes (CASCADE-P) using 16S rDNA.

MOTIVATION: Prokaryotic organisms have been identified utilizing the sequence variation of the 16S rRNA gene. Variations steer the design of DNA probes for the detection of taxonomic groups or specific organisms. The long-term goal of our project is to create probe arrays capable of identifying 16S rDNA sequences in unknown samples. This necessitated the authentication, categorization and alignment of the >75 000 publicly available '16S' sequences. Preferably, the entire process should be computationally administrated so the aligned collection could periodically absorb 16S rDNA sequences from the public records. A complete multiple sequence alignment would provide a foundation for computational probe selection and facilitates microbial taxonomy and phylogeny. RESULTS: Here we report the alignment and similarity clustering of 62 662 16S rDNA sequences and an approach for designing effective probes for each cluster. A novel alignment compression algorithm, NAST (Nearest Alignment Space Termination), was designed to produce the uniform multiple sequence alignment referred to as the prokMSA. From the prokMSA, 9020 Operational Taxonomic Units (OTUs) were found based on transitive sequence similarities. An automated approach to probe design was straightforward using the prokMSA clustered into OTUs. As a test case, multiple probes were computationally picked for each of the 27 OTUs that were identified within the Staphylococcus Group. The probes were incorporated into a customized microarray and were able to correctly categorize Staphylococcus aureus and Bacillus anthracis into their correct OTUs. Although a successful probe picking strategy is outlined, the main focus of creating the prokMSA was to provide a comprehensive, categorized, updateable 16S rDNA collection useful as a foundation for any probe selection algorithm.

Algorithms↗

A neural-network based method for prediction of gamma-turns in proteins from multiple sequence alignment.

In the present study, an attempt has been made to develop a method for predicting gamma-turns in proteins. First, we have implemented the commonly used statistical and machine-learning techniques in the field of protein structure prediction, for the prediction of gamma-turns. All the methods have been trained and tested on a set of 320 nonhomologous protein chains by a fivefold cross-validation technique. It has been observed that the performance of all methods is very poor, having a Matthew's Correlation Coefficient (MCC) </= 0.06. Second, predicted secondary structure obtained from PSIPRED is used in gamma-turn prediction. It has been found that machine-learning methods outperform statistical methods and achieve an MCC of 0.11 when secondary structure information is used. The performance of gamma-turn prediction is further improved when multiple sequence alignment is used as the input instead of a single sequence. Based on this study, we have developed a method, GammaPred, for gamma-turn prediction (MCC = 0.17). The GammaPred is a neural-network-based method, which predicts gamma-turns in two steps. In the first step, a sequence-to-structure network is used to predict the gamma-turns from multiple alignment of protein sequence. In the second step, it uses a structure-to-structure network in which input consists of predicted gamma-turns obtained from the first step and predicted secondary structure obtained from PSIPRED.

Databases, Protein↗

Potential for dramatic improvement in sequence alignment against structures of remote homologous proteins by extracting structural information from multiple structure alignment.

A novel method has been developed for acquiring the correct alignment of a query sequence against remotely homologous proteins by extracting structural information from profiles of multiple structure alignment. A systematic search algorithm combined with a group of score functions based on sequence information and structural information has been introduced in this procedure. A limited number of top solutions (15,000) with high scores were selected as candidates for further examination. On a test-set comprising 301 proteins from 75 protein families with sequence identity less than 30%, the proportion of proteins with completely correct alignment as first candidate was improved to 39.8% by our method, whereas the typical performance of existing sequence-based alignment methods was only between 16.1% and 22.7%. Furthermore, multiple candidates for possible alignment were provided in our approach, which dramatically increased the possibility of finding correct alignment, such that completely correct alignments were found amongst the top-ranked 1000 candidates in 88.3% of the proteins. With the assistance of a sequence database, completely correct alignment solutions were achieved amongst the top 1000 candidates in 94.3% of the proteins. From such a limited number of candidates, it would become possible to identify more correct alignment using a more time-consuming but more powerful method with more detailed structural information, such as side-chain packing and energy minimization, etc. The results indicate that the novel alignment strategy could be helpful for extending the application of highly reliable methods for fold identification and homology modeling to a huge number of homologous proteins of low sequence similarity. Details of the methods, together with the results and implications for future development are presented.

Algorithms↗

Phylo-VISTA: interactive visualization of multiple DNA sequence alignments.

MOTIVATION: The power of multi-sequence comparison for biological discovery is well established. The need for new capabilities to visualize and compare cross-species alignment data is intensified by the growing number of genomic sequence datasets being generated for an ever-increasing number of organisms. To be efficient these visualization algorithms must support the ability to accommodate consistently a wide range of evolutionary distances in a comparison framework based upon phylogenetic relationships. RESULTS: We have developed Phylo-VISTA, an interactive tool for analyzing multiple alignments by visualizing a similarity measure for multiple DNA sequences. The complexity of visual presentation is effectively organized using a framework based upon interspecies phylogenetic relationships. The phylogenetic organization supports rapid, user-guided interspecies comparison. To aid in navigation through large sequence datasets, Phylo-VISTA leverages concepts from VISTA that provide a user with the ability to select and view data at varying resolutions. The combination of multiresolution data visualization and analysis, combined with the phylogenetic framework for interspecies comparison, produces a highly flexible and powerful tool for visual data analysis of multiple sequence alignments. AVAILABILITY: Phylo-VISTA is available at http://www-gsd.lbl.gov/phylovista. It requires an Internet browser with Java Plug-in 1.4.2 and it is integrated into the global alignment program LAGAN at http://lagan.stanford.edu

Algorithms↗

SGP-1: prediction and validation of homologous genes based on sequence alignments.

Conventional methods of gene prediction rely on the recognition of DNA-sequence signals, the coding potential or the comparison of a genomic sequence with a cDNA, EST, or protein database. Reasons for limited accuracy in many circumstances are species-specific training and the incompleteness of reference databases. Lately, comparative genome analysis has attracted increasing attention. Several analysis tools that are based on human/mouse comparisons are already available. Here, we present a program for the prediction of protein-coding genes, termed SGP-1 (Syntenic Gene Prediction), which is based on the similarity of homologous genomic sequences. In contrast to most existing tools, the accuracy of depends little on species-specific properties such as codon usage or the nucleotide distribution. may therefore be applied to nonstandard model organisms in vertebrates as well as in plants, without the need for extensive parameter training. In addition to predicting genes in large-scale genomic sequences, the program may be useful to validate gene structure annotations from databases. To this end, SGP-1 output also contains comparisons between predicted and annotated gene structures in HTML format. The program can be accessed via a Web server at http://soft.ice.mpg.de/sgp-1. The source code, written in ANSI C, is available on request from the authors.

Algorithms↗

RNA secondary structure prediction from sequence alignments using a network of k-nearest neighbor classifiers.

We present a machine learning method (a hierarchical network of k-nearest neighbor classifiers) that uses an RNA sequence alignment in order to predict a consensus RNA secondary structure. The input to the network is the mutual information, the fraction of complementary nucleotides, and a novel consensus RNAfold secondary structure prediction of a pair of alignment columns and its nearest neighbors. Given this input, the network computes a prediction as to whether a particular pair of alignment columns corresponds to a base pair. By using a comprehensive test set of 49 RFAM alignments, the program KNetFold achieves an average Matthews correlation coefficient of 0.81. This is a significant improvement compared with the secondary structure prediction methods PFOLD and RNAalifold. By using the example of archaeal RNase P, we show that the program can also predict pseudoknot interactions.

Algorithms↗

A study on protein sequence alignment quality.

One of the most central methods in bioinformatics is the alignment of two protein or DNA sequences. However, so far large-scale benchmarks examining the quality of these alignments are scarce. On the other hand, recently several large-scale studies of the capacity of different methods to identify related sequences has led to new insights about the performance of fold recognition methods. To increase our understanding about fold recognition methods, we present a large-scale benchmark of alignment quality. We compare alignments from several different alignment methods, including sequence alignments, hidden Markov models, PSI-BLAST, CLUSTALW, and threading methods. For most methods, the alignment quality increases significantly at about 20% sequence identity. The difference in alignment quality between different methods is quite small, and the main difference can be seen at the exact positioning of the sharp rise in alignment quality, that is, around 15-20% sequence identity. The alignments are improved by using structural information. In general, the best alignments are obtained by methods that use predicted secondary structure information and sequence profiles obtained from PSI-BLAST. One interesting observation is that for different pairs many different methods create the best alignments. This finding implies that if a method that could select the best alignment method for each pair existed, a significant improvement of the alignment quality could be gained.

Computational Biology↗

Contextual multiple sequence alignment.

In a recently proposed contextual alignment model, efficient algorithms exist for global and local pairwise alignment of protein sequences. Preliminary results obtained for biological data are very promising. Our main motivation was to adopt the idea of context dependency to the multiple-alignment setting. To this aim the relaxation of the model was developed (we call this new model averaged contextual alignment) and a new family of amino acids substitution matrices are constructed. In this paper we present a contextual multiple-alignment algorithm and report the outcomes of experiments performed for the BAliBASE test set. The contextual approach turned out to give much better results for the set of sequences containing orphan genes.

Journal Article↗

CREMSA: compressed indexing of (ultra) large multiple sequence alignments.

MOTIVATION: Recent viral outbreaks motivate the systematic collection of pathogenic genomes in order to accelerate their study and monitor the apparition/spread of variants. Due to their limited length and temporal proximity of their sequencing, viral genomes are usually organized, and analyzed as oversized Multiple Sequence Alignments (MSAs). Such MSAs are largely ungapped, and mostly homogeneous on a column-wise level but not at a sequential level due to local variations, hindering the performances of sequential compression algorithms. RESULTS: In order to enable an efficient handling of MSAs, including subsequent statistical analyses, we introduce CREMSA (Column-wise Run-length Encoding for MSAs), a new index that builds on sparse bitvector representations to compress an existing or streamed MSA, all the while allowing for an expressive set of accelerated requests to query the alignment without prior decompression. Using CREMSA, a 65 GB MSA consisting of 1.9M SARS-CoV 2 genomes could be compressed into 22 MB using less than half a gigabyte of main memory, while executing access requests in the order of 100&#x2009;ns. Such a speed up enables a comprehensive analysis of covariation over this very large MSA. We further assess the impact of the sequence ordering on the compressibility of MSAs and propose a resorting strategy that, despite the proven NP-hardness of an optimal sort, induces greatly increased compression ratios at a marginal computational cost. AVAILABILITY AND IMPLEMENTATION: CREMSA is freely accessible at https://gitlab.univ-lille.fr/cremsa/cremsa. The Snakemake workflow for the benchmarks is available at: https://gitlab.univ-lille.fr/cremsa/bench. The data used in the paper is on Zenodo at https://zenodo.org/records/14698859 and https://zenodo.org/records/15100011.

SARS-CoV-2↗