PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

An IRP-like protein from Plasmodium falciparum binds to a mammalian iron-responsive element.

This study cloned and sequenced the complementary DNA (cDNA) encoding of a putative malarial iron responsive element-binding protein (PfIRPa) and confirmed its identity to the previously identified iron-regulatory protein (IRP)-like cDNA from Plasmodium falciparum. Sequence alignment showed that the plasmodial sequence has 47% identity with human IRP1. Hemoglobin-free lysates obtained from erythrocyte-stage P falciparum contain a protein that binds a consensus mammalian iron-responsive element (IRE), indicating that a protein(s) with iron-regulatory activity was present in the lysates. IRE-binding activity was found to be iron regulated in the electrophoretic mobility shift assays. Western blot analysis showed a 2-fold increase in the level of PfIRPa in the desferrioxamine-treated cultures versus control or iron-supplemented cells. Malarial IRP was detected by anti-PfIRPa antibody in the IRE-protein complex from P falciparum lysates. Immunofluorescence studies confirmed the presence of PfIRPa in the infected red blood cells. These findings demonstrate that erythrocyte P falciparum contains an iron-regulated IRP that binds a mammalian consensus IRE sequence, raising the possibility that the malaria parasite expresses transcripts that contain IREs and are iron-dependently regulated.

Aconitate Hydratase↗

Probing the structure of the Escherichia coli 10Sa RNA (tmRNA).

The conformation of the Escherichia coli 10Sa RNA (tmRNA) in solution was investigated using chemical and enzymatic probes. Single- and double-stranded domains were identified by hydrolysis of tmRNA in imidazole buffer and by lead(II)-induced cleavages. Ribonucleases T1 and S1 were used to map unpaired nucleotides and ribonuclease V1 was used to identify paired bases or stacked nucleotides. Specific atomic positions of bases were probed with dimethylsulfate, a carbodiimide, and diethylpyrocarbonate. Covariations, identified by sequence alignment with nine other tmRNA sequences, suggest the presence of several tertiary interactions, including pseudoknots. Temperature-gradient gel electrophoresis experiments showed structural transitions of tmRNA starting around 40 degrees C, and enzymatic probing performed at selected temperatures revealed the progressive melting of several predicted interactions. Based on these data, a secondary structure is proposed, containing two stems, four stem-loops, four pseudoknots, and an unstable structural domain, some connected by single-stranded A-rich sequence stretches. A tRNA-like domain, including an already reported acceptor branch, is supported by the probing data. A second structural domain encompasses the coding sequence, which extends from the top of one stem-loop to the top of another, with a 7-nt single-stranded stretch between. A third structural module containing pseudoknots connects and probably orients the tRNA-like domain and the coding sequence. Several discrepancies between the probing data and the phylogeny suggest that E. coli tmRNA undergoes a conformational change.

Alanine↗

SALSA: improved protein database searching by a new algorithm for assembly of sequence fragments into gapped alignments.

MOTIVATION: Optimal sequence alignment based on the Smith-Waterman algorithm is usually too computationally demanding to be practical for searching large sequence databases. Heuristic programs like FASTA and BLAST have been developed which run much faster, but at the expense of sensitivity. RESULTS: In an effort to approximate the sensitivity of an optimal alignment algorithm, a new algorithm has been devised for the computation of a gapped alignment of two sequences. After scanning for high-scoring words and extensions of these to form fragments of similarity, the algorithm uses dynamic programming to build an accurate alignment based on the fragments initially identified. The algorithm has been implemented in a program called SALSA and the performance has been evaluated on a set of test sequences. The sensitivity was found to be close to the Smith-Waterman algorithm, while the speed was similar to FASTA (ktup = 2). AVAILABILITY: Searches can be performed from the SALSA homepage at http://dna.uio.no/salsa/ using a wide range of databases. Source code and precompiled executables are also available. CONTACT: torbjorn.rognes@labmed.uio.no

Algorithms↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗

Cloning and characterization of a cytolytic and mosquitocidal delta-endotoxin from Bacillus thuringiensis subsp. jegathesan.

A cytolytic toxin gene encoding a 30.1-kDa Cyt2Bb1 toxin protein from B. thuringiensis subsp. jegathasan was cloned employing a limited-growth PCR screening method with forward and reverse oligonucleotide primers designed from N-terminal amino acid sequences of native and trypsin-cleaved protein, respectively. The expressed protein showed little cross-reactivity to the antibody raised against the Cyt1Aa protein. Unlike Cyt1Aa and Cyt2Aa expression, there was little or no visible crystal inclusion formation under microscopic observation. The amino acid sequence alignment indicated 31 and 66% identity to Cyt1Aa and Cyt2Aa, respectively. The sequence alignment for five known cytolytic proteins indicated three highly conserved regions, two in the loop regions between alpha-helices and beta-sheets and one in the loop region between beta-sheets 5 and 6. beta-Blocks 4 to 7 are also conserved, not only structurally but also among the amino acids in the hydrophobic faces. Mosquitocidal activity assays indicated that the Cyt2Bb toxin had less toxicity than Cyt1Aa and had about 600-times-lower toxicity than the wild-type whole toxin crystal. However, both the Cyt2Bb and the Cyt1Aa toxin showed comparable levels of hemolytic activity.

Aedes↗

PARSESNP: A tool for the analysis of nucleotide polymorphisms.

PARSESNP is a tool for the display and analysis of polymorphisms in genes. Using a reference DNA sequence, an exon/intron position model and a list of polymorphisms, it determines the effects of these polymorphisms on the expressed gene product, as well as the changes in restriction enzyme recognition sites. It shows the locations and effects of the polymorphisms in summary on a stylized graphic and in detail on a display of the protein sequence aligned with the DNA sequence. The addition of a homology model, in the form of an alignment of related protein sequences, allows for prediction of the severity of missense changes. PARSESNP is available on the World Wide Web at http://www.proweb.org/parsesnp/.

Base Sequence↗

Using iterative dynamic programming to obtain accurate pairwise and multiple alignments of protein structures.

We show how a basic pairwise alignment procedure can be improved to more accurately align conserved structural regions, by using variable, position-dependent gap penalties that depend on secondary structure and by taking the consensus of a number of suboptimal alignments. These improvements, which are novel for structural alignment, are direct analogs of what is possible with normal sequences alignment. They are feasible for us since our basic structural alignment procedure, unlike others, is so similar to normal sequence alignment. We further present preliminary results that show how our procedure can be generalized to produce a multiple alignment of a family of structures. Our approach is based on finding a "median" structure from doing all possible pairwise alignments and then aligning everything to it.

Amino Acid Sequence↗

eShadow: a tool for comparing closely related sequences.

Primate sequence comparisons are difficult to interpret due to the high degree of sequence similarity shared between such closely related species. Recently, a novel method, phylogenetic shadowing, has been pioneered for predicting functional elements in the human genome through the analysis of multiple primate sequence alignments. We have expanded this theoretical approach to create a computational tool, eShadow, for the identification of elements under selective pressure in multiple sequence alignments of closely related genomes, such as in comparisons of human-to-primate or mouse-to-rat DNA. This tool integrates two different statistical methods and allows for the dynamic visualization of the resulting conservation profile. eShadow also includes a versatile optimization module capable of training the underlying Hidden Markov Model to differentially predict functional sequences. This module grants the tool high flexibility in the analysis of multiple sequence alignments and in comparing sequences with different divergence rates. Here, we describe the eShadow comparative tool and its potential uses for analyzing both multiple nucleotide and protein alignments to predict putative functional elements.

Animals↗

Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis.

MOTIVATION: Nanopore sequencing by Oxford Nanopore Technologies (ONT) enables direct analysis of DNA and RNA by capturing raw electrical signals. Different nanopore chemistries have varied k-mer lengths, current levels, and standard deviations, which are stored in "k-mer models." In cases where official models are lacking or unsuitable for specific sequencing conditions, tailored k-mer models are crucial to ensure precise signal-to-sequence alignment, analysis and interpretation. The process of transforming raw signal data into nucleotide sequences, known as basecalling, is a fundamental step in nanopore sequencing. RESULTS: In this study, we leverage the move table produced by ONT's basecalling software to create a lightweight de novo k-mer model for RNA004 chemistry. We demonstrate the validity of our custom k-mer model by using it to guide signal-to-sequence alignment analysis, achieving high alignment rates (97.48%) compared to larger default models. Additionally, our 5-mer model exhibits similar performance as the default 9-mer models another analysis, such as detection of m6A RNA modifications. We provide our method, termed Poregen, as a generalizable approach for creation of custom, de novo k-mer models for nanopore signal data analysis. AVAILABILITY AND IMPLEMENTATION: Poregen is an open source package under an MIT license: https://github.com/hiruna72/poregen.

Nanopore Sequencing↗

Comparative modeling of the three-dimensional structure of type II antifreeze protein.

Type II antifreeze proteins (AFP), which inhibit the growth of seed ice crystals in the blood of certain fishes (sea raven, herring, and smelt), are the largest known fish AFPs and the only class for which detailed structural information is not yet available. However, a sequence homology has been recognized between these proteins and the carbohydrate recognition domain of C-type lectins. The structure of this domain from rat mannose-binding protein (MBP-A) has been solved by X-ray crystallography (Weis WI, Drickamer K, Hendrickson WA, 1992, Nature 360:127-134) and provided the coordinates for constructing the three-dimensional model of the 129-amino acid Type II AFP from sea raven, to which it shows 19% sequence identity. Multiple sequence alignments between Type II AFPs, pancreatic stone protein, MBP-A, and as many as 50 carbohydrate-recognition domain sequences from various lectins were performed to determine reliably aligned sequence regions. Successive molecular dynamics and energy minimization calculations were used to relax bond lengths and angles and to identify flexible regions. The derived structure contains two alpha-helices, two beta-sheets, and a high proportion of amino acids in loops and turns. The model is in good agreement with preliminary NMR spectroscopic analyses. It explains the observed differences in calcium binding between sea raven Type II AFP and MBP-A. Furthermore, the model proposes the formation of five disulfide bridges between Cys 7 and Cys 18, Cys 35 and Cys 125, Cys 69 and Cys 100, Cys 89 and Cys 111, and Cys 101 and Cys 117.(ABSTRACT TRUNCATED AT 250 WORDS)

Adaptation, Physiological↗

Flexible programs for the prediction of average amphipathicity of multiply aligned homologous proteins: application to integral membrane transport proteins.

Simple flexible programs (TREEMOMENT and PILEUPMOMENT) are described for depicting the average amphipathicity (hydrophobic moment) along multiply aligned sequences of a family of evolutionarily related proteins. The programs are applicable to any number of aligned sequences and can be set for any desired angle corresponding to a residue repeat unit in a protein secondary structural element such as 100 degrees per residue for an alpha-helix or 180 degrees per residue for a beta-strand. These programs can be used to identify amphipathic regions common to the members of a protein family. The use of these programs is exemplified by showing that some families of integral membrane transport proteins (i.e. permeases of the bacterial phosphotransferase system (PTS) and the anion exchangers of animals) exhibit strikingly amphipathic alpha-helical structures immediately preceding the first hydrophobic transmembrane segment of their membrane-embedded domain(s). Other families, such as the major facilitator superfamily of uniporters, symporters and antiporters, do not exhibit this structural feature. The amphipathic structures in PTS permeases have been implicated in membrane insertion during biogenesis.

Amino Acid Sequence↗

MulPSSM: a database of multiple position-specific scoring matrices of protein domain families.

Representation of multiple sequence alignments of protein families in terms of position-specific scoring matrices (PSSMs) is commonly used in the detection of remote homologues. A PSSM is generated with respect to one of the sequences involved in the multiple sequence alignment as a reference. We have shown recently that the use of multiple PSSMs corresponding to an alignment, with several sequences in the family used as reference, improves the sensitivity of the remote homology detection dramatically. MulPSSM contains PSSMs for a large number of sequence and structural families of protein domains with multiple PSSMs for every family. The approach involves use of a clustering algorithm to identify most distinct sequences corresponding to a family. With each one of the distinct sequences as reference, multiple PSSMs have been generated. The current release of MulPSSM contains approximately 33,000 and approximately 38,000 PSSMs corresponding to 7868 sequence and 2625 structural families. A RPS_BLAST interface allows sequence search against PSSMs of sequence or structural families or both. An analysis interface allows display and convenient navigation of alignments and domain hits. MulPSSM can be accessed at http://crick.mbu.iisc.ernet.in/~mulpssm.

Databases, Protein↗

Ribostral: an RNA 3D alignment analyzer and viewer based on basepair isostericities.

UNLABELLED: RNA atomic resolution structures have revealed the existance of different families of basepair interactions, each of which with its own isosteric sub-families. Ribostral (Ribonucleic Structural Aligner) is a user-friendly framework for analyzing, evaluating and viewing RNA sequence alignments with at least one available atomic resolution structure. It is the first of its kind that makes direct and easy- to-understand superposition of the isostericity matrices of basepairs observed in the structure onto sequence alignments, easily indicating allowed and unallowed substitutions at each BP position. Potential mistakes in the alignments can then be corrected using other sequence editing software. Ribostral has been developed and tested under Windows XP, and is capable of running on any PC or MAC platform with MATLAB 7.1 (SP3) or higher installed version. A stand-alone version is also available for the PC platform. AVAILABILITY: http://rna.bgsu.edu/ribostral.

Algorithms↗

Profile sequence analysis and database searches on a transputer machine connected to a Macintosh computer.

An implementation of Profilesearch (a technique to search for relationships between a protein sequence and multiply aligned sequences) for a parallel computer is described. The number-crunching machine, consisting of 21 T800 transputers, is connected to a Macintosh IIcx host computer. The program utilizes a standard Macintosh application as its user-interface, resulting in a transparent and user-friendly environment for addressing the parallel computer. The program is independent of the number of available processors and exceeds the speed of a VAXstation 3200 with only one transputer in operation, thus allowing cheap and fast database searches with a PC front-end. For a larger number of processors, the speed increase is approximately linear with no obvious symptoms of saturation with the available maximum of 21 transputers. The program and environment are useful to search quickly and easily for similarities between a single sequence or sequence set and individual sequences contained in a large database. The alignment is determined by typical dynamic programming techniques.

Amino Acid Sequence↗

The EMOTIF database.

The EMOTIF database is a collection of more than 170 000 highly specific and sensitive protein sequence motifs representing conserved biochemical properties and biological functions. These protein motifs are derived from 7697 sequence alignments in the BLOCKS+ database (released on June 23, 2000) and all 8244 protein sequence alignments in the PRINTS database (version 27.0) using the emotif-maker algorithm developed by Nevill-Manning et al. (Nevill-Manning,C.G., Wu,T.D. and Brutlag,D.L. (1998) Proc. Natl Acad. Sci. USA, 95, 5865-5871; Nevill-Manning,C.G., Sethi,K.S., Wu,T. D. and Brutlag,D.L. (1997) ISMB-97, 5, 202-209). Since the amino acids and the groups of amino acids in these sequence motifs represent critical positions conserved in evolution, search algorithms employing the EMOTIF patterns can identify and classify more widely divergent sequences than methods based on global sequence similarity. The emotif protein pattern database is available at http://motif.stanford.edu/emotif/.

Amino Acid Motifs↗

Improved alignment quality by combining evolutionary information, predicted secondary structure and self-organizing maps.

BACKGROUND: Protein sequence alignment is one of the basic tools in bioinformatics. Correct alignments are required for a range of tasks including the derivation of phylogenetic trees and protein structure prediction. Numerous studies have shown that the incorporation of predicted secondary structure information into alignment algorithms improves their performance. Secondary structure predictors have to be trained on a set of somewhat arbitrarily defined states (e.g. helix, strand, coil), and it has been shown that the choice of these states has some effect on alignment quality. However, it is not unlikely that prediction of other structural features also could provide an improvement. In this study we use an unsupervised clustering method, the self-organizing map, to assign sequence profile windows to "structural states" and assess their use in sequence alignment. RESULTS: The addition of self-organizing map locations as inputs to a profile-profile scoring function improves the alignment quality of distantly related proteins slightly. The improvement is slightly smaller than that gained from the inclusion of predicted secondary structure. However, the information seems to be complementary as the two prediction schemes can be combined to improve the alignment quality by a further small but significant amount. CONCLUSION: It has been observed in many studies that predicted secondary structure significantly improves the alignments. Here we have shown that the addition of self-organizing map locations can further improve the alignments as the self-organizing map locations seem to contain some information that is not captured by the predicted secondary structure.

Computational Biology↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

CLOURE: Clustal Output Reformatter, a program for reformatting ClustalX/ClustalW outputs for SNP analysis and molecular systematics.

We describe a program (and a website) to reformat the ClustalX/ClustalW outputs to a format that is widely used in the presentation of sequence alignment data in SNP analysis and molecular systematic studies. This program, CLOURE, CLustal OUtput REformatter, takes the multiple sequence alignment file (nucleic acid or protein) generated from Clustal as input files. The CLOURE-D format presents the Clustal alignment in a format that highlights only the different nucleotides/residues relative to the first query sequence. The program has been written in Visual Basic and will run on a Windows platform. The downloadable program, as well as a web-based server which has also been developed, can be accessed at http://imtech.res.in/~anand/cloure.html.

Internet↗