PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Combining the GOR V algorithm with evolutionary information for protein secondary structure prediction from amino acid sequence.

We have modified and improved the GOR algorithm for the protein secondary structure prediction by using the evolutionary information provided by multiple sequence alignments, adding triplet statistics, and optimizing various parameters. We have expanded the database used to include the 513 non-redundant domains collected recently by Cuff and Barton (Proteins 1999;34:508-519; Proteins 2000;40:502-511). We have introduced a variable size window that allowed us to include sequences as short as 20-30 residues. A significant improvement over the previous versions of GOR algorithm was obtained by combining the PSI-BLAST multiple sequence alignments with the GOR method. The new algorithm will form the basis for the future GOR V release on an online prediction server. The average accuracy of the prediction of secondary structure with multiple sequence alignment and full jack-knife procedure was 73.5%. The accuracy of the prediction increases to 74.2% by limiting the prediction to 375 (of 513) sequences having at least 50 PSI-BLAST alignments. The average accuracy of the prediction of the new improved program without using multiple sequence alignments was 67.5%. This is approximately a 3% improvement over the preceding GOR IV algorithm (Garnier J, Gibrat JF, Robson B. Methods Enzymol 1996;266:540-553; Kloczkowski A, Ting K-L, Jernigan RL, Garnier J. Polymer 2002;43:441-449). We have discussed alternatives to the segment overlap (Sov) coefficient proposed by Zemla et al. (Proteins 1999;34:220-223).

Algorithms↗

ConStruct: a tool for thermodynamic controlled prediction of conserved secondary structure.

A tool for prediction of conserved secondary structure of a set of homologous single-stranded RNAs is presented. For each RNA of the set the structure distribution is calculated and stored in a base pair probability matrix. Gaps, resulting from a multiple sequence alignment of the RNA set, are introduced into the individual probability matrices. These 'aligned' probability matrices are summed up to give a consensus probability matrix emphasizing the conserved structural elements of the RNA set. Because the multiple sequence alignment is independent of any structural constraints, such an alignment may result in introduction of gaps into the homologous probability matrices that disrupt a common consensus structure. By use of its graphical user interface the presented tool allows the removal of such misalignments, which are easily recognized, from the individual probability matrices by optimizing the sequence alignment with respect to a structural alignment. From the consensus probability matrix a consensus structure is extracted, which is viewable in three different graphical representations. The functionality of the tool is demonstrated using a small set of U7 RNAs, which are involved in 3'-end processing of histone mRNA precursors. Supplementary Material lists further results obtained. Advantages and drawbacks of the tool are discussed in comparison to several other algorithms.

Algorithms↗

Alignment of protein sequences by their profiles.

The accuracy of an alignment between two protein sequences can be improved by including other detectably related sequences in the comparison. We optimize and benchmark such an approach that relies on aligning two multiple sequence alignments, each one including one of the two protein sequences. Thirteen different protocols for creating and comparing profiles corresponding to the multiple sequence alignments are implemented in the SALIGN command of MODELLER. A test set of 200 pairwise, structure-based alignments with sequence identities below 40% is used to benchmark the 13 protocols as well as a number of previously described sequence alignment methods, including heuristic pairwise sequence alignment by BLAST, pairwise sequence alignment by global dynamic programming with an affine gap penalty function by the ALIGN command of MODELLER, sequence-profile alignment by PSI-BLAST, Hidden Markov Model methods implemented in SAM and LOBSTER, pairwise sequence alignment relying on predicted local structure by SEA, and multiple sequence alignment by CLUSTALW and COMPASS. The alignment accuracies of the best new protocols were significantly better than those of the other tested methods. For example, the fraction of the correctly aligned residues relative to the structure-based alignment by the best protocol is 56%, which can be compared with the accuracies of 26%, 42%, 43%, 48%, 50%, 49%, 43%, and 43% for the other methods, respectively. The new method is currently applied to large-scale comparative protein structure modeling of all known sequences.

Algorithms↗

A systematic search for RNA editing sites in pea chloroplasts: an editing event causes diversification from the evolutionarily conserved amino acid sequence.

RNA editing in higher plant chloroplasts involves C-to-U conversion at specific sites in the transcripts. To examine whether pea shares editing sites with other angiosperms, a systematic search for editing sites in pea chloroplast transcripts was performed. Based on amino acid sequence alignment, 451 RNA editing sites were predicted from 60 transcripts. Sequence analysis of amplified cDNAs for these potential editing sites revealed 19 true editing sites from 13 transcripts. Together with those reported previously, the total number of editing sites is 27 from 16 transcripts in pea chloroplasts. Twenty-two sites are conserved among other plant species, whereas five sites are unique to pea. Among the 27 editing sites, seven are partially edited. The most interesting is the ndhG site 1, which has led to the diversification of the evolutionarily conserved amino acid sequence. This observation suggests that some of the editing events cause the diversity of amino acid sequences, and hence, that prediction of editing sites based on amino acid sequence alignment has its own limitations.

Amino Acid Sequence↗

Multiple model approach--dealing with alignment ambiguities in protein modeling.

Sequence alignments for distantly homologous proteins are often ambiguous, which creates a weak link in structure prediction by homology. We address this problem by using several plausible alignments in a modeling procedure, obtaining many models of the target. All are subsequently evaluated by a threading algorithm. It is shown that this approach can identify best alignments and produce reasonable models, whose quality is now limited only by the extent of the structural similarity between the known and predicted protein. Using a similar approach structure prediction for the oxidized dimer of S100A1 protein, for which the structure is not known, is presented.

Amino Acid Sequence↗

ESPript/ENDscript: Extracting and rendering sequence and 3D information from atomic structures of proteins.

The fortran program ESPript was created in 1993, to display on a PostScript figure multiple sequence alignments adorned with secondary structure elements. A web server was made available in 1999 and ESPript has been linked to three major web tools: ProDom which identifies protein domains, PredictProtein which predicts secondary structure elements and NPS@ which runs sequence alignment programs. A web server named ENDscript was created in 2002 to facilitate the generation of ESPript figures containing a large amount of information. ENDscript uses programs such as BLAST, Clustal and PHYLODENDRON to work on protein sequences and such as DSSP, CNS and MOLSCRIPT to work on protein coordinates. It enables the creation, from a single Protein Data Bank identifier, of a multiple sequence alignment figure adorned with secondary structure elements of each sequence of known 3D structure. Similar 3D structures are superimposed in turn with the program PROFIT and a final figure is drawn with BOBSCRIPT, which shows sequence and structure conservation along the Calpha trace of the query. ESPript and ENDscript are available at http://genopole.toulouse.inra.fr/ESPript.

DNA-Binding Proteins↗

AliWABA: alignment on the web through an A-Bruijn approach.

UNLABELLED: Multiple sequence alignment programs are an invaluable tool in computational biology. A-Bruijn Alignment (ABA) is a method for multiple sequence alignment that represents an alignment as a directed graph and has proved useful in aligning nucleotide and amino acid sequences that are composed of repeated and shuffled subsequences. AliWABA is a web server that provides tools to generate alignments with ABA, visualize the resulting ABA graphs and extract subsequences from ABA graphs. AliWABA greatly simplifies the problem of analyzing multiple sequences for local similarities that may be reordered, as is common with the domain architectures of proteins. To facilitate the analysis of protein domains, AliWABA provides direct querying of the Conserved Domain Database. AVAILABILITY: http://aba.nbcr.net/

Algorithms↗

A workbench for multiple alignment construction and analysis.

Multiple sequence alignment can be a useful technique for studying molecular evolution, as well as for analyzing relationships between structure or function and primary sequence. We have developed for this purpose an interactive program, MACAW (Multiple Alignment Construction and Analysis Workbench), that allows the user to construct multiple alignments by locating, analyzing, editing, and combining "blocks" of aligned sequence segments. MACAW incorporates several novel features. (1) Regions of local similarity are located by a new search algorithm that avoids many of the limitations of previous techniques. (2) The statistical significance of blocks of similarity is evaluated using a recently developed mathematical theory. (3) Candidate blocks may be evaluated for potential inclusion in a multiple alignment using a variety of visualization tools. (4) A user interface permits each block to be edited by moving its boundaries or by eliminating particular segments, and blocks may be linked to form a composite multiple alignment. No completely automatic program is likely to deal effectively with all the complexities of the multiple alignment problem; by combining a powerful similarity search algorithm with flexible editing, analysis and display tools, MACAW allows the alignment strategy to be tailored to the problem at hand.

Algorithms↗

Homology model of human corticosteroid binding globulin: a study of its steroid binding ability and a plausible mechanism of steroid hormone release at the site of inflammation.

Corticosteroid binding globulin (CBG) and thyroxin binding globulin (TBG) both belong to the same SERPIN superfamily of serine-proteinase inhibitors but in the course of evolution CBG has adapted to its new role as a transport agent of insoluble hormones. CBG binds corticosteroids in plasma, delivering them to sites of inflammation to modify the inflammatory response. CBG is an effective drug carrier for genetic manipulation, and hence there is immense biological interest in the location of the hormone binding site. The crystal structure of human CBG (hCBG) has not been determined, but sequence alignment with other SERPINs suggests that it conforms as a whole to the tertiary structure shared by the superfamily. Human CBG shares 52.15% and 55.50% sequence similarity with alpha1-antitrypsin and alpha1-antichymotrypsin, respectively. Multiple sequence alignment among the three sequences shows 73 conserved regions. The molecular structures of alpha1-antitrypsin and alpha1-antichymotrypsin, the archetype of the SERPIN superfamily, obtained by X-ray diffraction methods are used to develop a homology model of hCBG. Energy minimization was applied to the model to refine the structure further. The homology model of hCBG contains 371 residues (His13 to Val383 ). The secondary structure comprises 11 helices, 15 turns and 11 sheets. The putative corticosteroid binding region is found to exist in a pocket between beta-sheets S4, S10, S11 and alpha helix H10. Both cortisol and aldosterone are docked to the elongated hydrophobic ligand binding pocket with the polar residues at the two extremities. A difference accessible surface area (DASA) study revealed that cortisol binds with the native hCBG more tightly than aldosterone. Cleavage at the Val379-Met380 peptide bond causes a deformation of hCBG (also revealed through a DASA study). This deformation could probably trigger the release of the bound hormone. Figure Stereoscopic view of the ribbon diagram of hCBG complexed with cortisol. The bound cortisol is shown in space filling model in blue. Helices and sheets are shown in red and magenta respectively. Turns are shown in yellow.

Adrenal Cortex Hormones↗

SSALN: an alignment algorithm using structure-dependent substitution matrices and gap penalties learned from structurally aligned protein pairs.

In template-based modeling of protein structures, the generation of the alignment between the target and the template is a critical step that significantly affects the accuracy of the final model. This paper proposes an alignment algorithm SSALN that learns substitution matrices and position-specific gap penalties from a database of structurally aligned protein pairs. In addition to the amino acid sequence information, secondary structure and solvent accessibility information of a position are used to derive substitution scores and position-specific gap penalties. In a test set of CASP5 targets, SSALN outperforms sequence alignment methods such as a Smith-Waterman algorithm with BLOSUM50 and PSI_BLAST. SSALN also generates better alignments than PSI_BLAST in the CASP6 test set. LOOPP server prediction based on an SSALN alignment is ranked the best for target T0280_1 in CASP6. SSALN is also compared with several threading methods and sequence alignment methods on the ProSup benchmark. SSALN has the highest alignment accuracy among the methods compared. On the Fischer's benchmark, SSALN performs better than CLUSTALW and GenTHREADER, and generates more alignments with accuracy >50%, >60% or >70% than FUGUE, but fewer alignments with accuracy >80% than FUGUE. All the supplemental materials can be found at http://www.cs.cornell.edu/ approximately jianq/research.htm.

Algorithms↗

CLAGen: a tool for clustering and annotating gene sequences using a suffix tree algorithm.

Most multiple gene sequence alignment methods rely on conventions regarding the score of a multiple alignment in pairwise fashion. Therefore, as the number of sequences increases, the runtime of sequencing expands exponentially. In order to solve the problem, this paper presents a multiple sequence alignment method using a linear-time suffix tree algorithm to cluster similar sequences at one time without pairwise alignment. After searching for common subsequences, cross-matching common subsequences were generated, and sometimes inexact matching was found. So, a procedure aimed at masking the inexact cross-matching pairs was suggested here. In addition, BLAST was combined with a clustering tool in order to annotate the clusters generated by suffix tree clustering. The proposed method for clustering and annotating genes consists of the following steps: (1) construction of a suffix tree; (2) searching and overlapping common subsequences; (3) grouping subsequence pairs; (4) masking cross-matching pairs; (5) clustering gene sequences; (6) annotating gene clusters by the BLAST search. The performance of the proposed system, CLAGen, was successfully evaluated with 42 gene sequences in a TCA cycle (a citrate cycle) of bacteria. The system generated 11 clusters and found the longest subsequences of each cluster, which are biologically significant.

Algorithms↗

Comparative modeling in CASP6 using consensus approach to template selection, sequence-structure alignment, and structure assessment.

Along with over 150 other groups we have tested our template-based protein structure prediction approach by submitting models for 30 target proteins to the sixth round of the Critical Assessment of Protein Structure Prediction Methods (CASP6, http://predictioncenter.org). Most of our modeled proteins fall into the comparative or homology modeling (CM) category, and some are fold recognition (FR) targets. The key feature of our structure prediction strategy in CASP6 was an attempt to optimally select structural templates and to make accurate sequence-structure alignments. Template selection was based mainly on consensus results of multiple sequence searches. Likewise, the consensus of multiple alignment variants (or lack of it) was used to initially delineate reliable and unreliable alignment regions. Structure evaluation approaches were then used to identify the correct sequence-structure mapping. Our results suggest that in many cases use of multiple templates is advantageous. Selecting correct alignments even within the context of a three-dimensional structure remains a challenge. Together with more effective energy evaluation methods the simultaneous relaxation/refinement of a "frozen" backbone inherited from the template is likely needed to see a clear progress in tackling this problem. Our analysis also suggests that human input has little to contribute to automatic methods in modeling high homology targets. On the other hand, human expertise can be very valuable in modeling distantly related proteins and critical in cases of unexpected evolutionary changes in protein structure.

Algorithms↗

The UniMarker (UM) method for synteny mapping of large genomes.

MOTIVATION: Synteny mapping, or detecting regions that are orthologous between two genomes, is a key step in studies of comparative genomics. For completely sequenced genomes, this is increasingly accomplished by whole-genome sequence alignment. However, such methods are computationally expensive, especially for large genomes, and require rather complicated post-processing procedures to filter out non-orthologous sequence matches. RESULTS: We have developed a novel method that does not require sequence alignment for synteny mapping of two large genomes, such as the human and mouse. In this method, the occurrence spectra of genome-wide unique 16mer sequences present in both the human and mouse genome are used to directly detect orthologous genomic segments. Being sequence alignment-free, the method is very fast and able to map the two mammalian genomes in one day of computing time on a single Pentium IV personal computer. The resulting human-mouse synteny map was shown to be in excellent agreement with those produced by the Mouse Genome Sequencing Consortium (MGSC) and by the Ensembl team; furthermore, the syntenic relationship of segments found only by our method was supported by BLASTZ sequence alignment.

Algorithms↗

Functional multimerization of human telomerase requires an RNA interaction domain in the N terminus of the catalytic subunit.

Functional human telomerase complexes are minimally composed of the human telomerase RNA (hTR) and a catalytic subunit (human telomerase reverse transcriptase [hTERT]) containing reverse transcriptase (RT)-like motifs. The N terminus of TERT proteins is unique to the telomerase family and has been implicated in catalysis, telomerase RNA binding, and telomerase multimerization, and conserved motifs have been identified by alignment of TERT sequences from multiple organisms. We studied hTERT proteins containing N-terminal deletions or substitutions to identify and characterize hTERT domains mediating telomerase catalytic activity, hTR binding, and hTERT multimerization. Using multiple sequence alignment, we identified two vertebrate-conserved TERT N-terminal regions containing vertebrate-specific residues that were required for human telomerase activity. We identified two RNA interaction domains, RID1 and RID2, the latter containing a vertebrate-specific RNA binding motif. Mutations in RID2 reduced the association of hTR with hTERT by 50 to 70%. Inactive mutants defective in RID2-mediated hTR binding failed to complement an inactive hTERT mutant containing an RT motif substitution to reconstitute activity. Our results suggest that functional hTERT complementation requires intact RID2 and RT domains on the same hTERT molecule and is dependent on hTR and the N terminus.

Amino Acid Sequence↗

Hydrogen-bond interactions of the primary donor of the photosynthetic purple sulfur bacterium Chromatium tepidum.

We have used near-infrared Fourier transform (pre)resonance Raman spectroscopy to determine the protein interactions with the bacteriochlorophyll (BChl) dimer constituting the primary electron donor, P, in the reaction center (RC) from the thermophilic purple sulfur bacterium Chromatium tepidum. In addition, we report the alignment of partial sequences of the L and M protein subunits of C. tepidum RCs in the vicinity of the primary donor with those of Rhodobacter sphaeroides and Rhodopseudomonas viridis. Taken together, these results enable us to propose the hydrogen-bonding pattern and the H-bond donors to the conjugated carbonyl groups of P. Selective excitation (1064-nm laser radiation) of the FT (pre)-resonance Raman spectra of P in its neutral (P degree) and oxidized (P degree +) states were obtained via their electronic absorption bands at 876 and 1240 nm, respectively. The P degree spectrum exhibits vibrational frequencies at 1608, 1616, 1633, and 1697 cm-1 which bleach upon P oxidation. The P degree + spectrum exhibits new bands at 1600, 1639, and 1719 cm-1. The 1608-cm-1 band, which downshifts to 1600 cm-1 upon oxidation, is assigned to a CaCm methine bridge stretching mode of the P dimer, indicating that each BChl molecule possesses a single axial ligand (His L181 and His M201, from the sequence alignment). The 1616- and 1633-cm-1 bands correspond to two H-bonded pi-conjugated acetyl carbonyl groups of each BChl molecule. with different H-bond strengths: the 1616-cm-1 band is assigned to the PL C2 acetyl group which is H-bonded to a histidine residue (His L176), while the 1633-cm-1 band is assigned to the PM C2 acetyl carbonyl, H-bonded to a tyrosine residue (Tyr M196). Both PL and PM C9 keto carbonyls are free from interactions and vibrate at the same frequency (1697 cm-1). Thus, the H-bond pattern of the primary donor of C. tepidum differs from that of Rb. sphaeroides in the extra H-bond to the PM C2 acetyl carbonyl group; that of PL is H-bonded to a histidine residue in both primary donors (His L168 in Rb. sphaeroides and His L176 in C. tepidum). The P degree/P degree + redox midpoint potentials were measured to be +497 and +526 mV for isolated C. tepidum RCs with and without the associated tetraheme cytochrome c subunit, respectively, and +502 mV for intracytoplasmic membranes. The positive charge localization was estimated to be 69% in favor of PL, indicating a more delocalized situation over the primary donor of C. tepidum than that of Rb. sphaeroides (estimated to be 80% on PL). These differences in physicochemical properties are discussed with respect to the proposed structural model for the microenvironment of the primary donor of C. tepidum.

Amino Acid Sequence↗

Motifs and structural fold of the cofactor binding site of human glutamate decarboxylase.

The pyridoxal-P binding sites of the two isoforms of human glutamate decarboxylase (GAD65 and GAD67) were modeled by using PROBE (a recently developed algorithm for multiple sequence alignment and database searching) to align the primary sequence of GAD with pyridoxal-P binding proteins of known structure. GAD's cofactor binding site is particularly interesting because GAD activity in the brain is controlled in part by a regulated interconversion of the apo- and holoenzymes. PROBE identified six motifs shared by the two GADs and four proteins of known structure: bacterial ornithine decarboxylase, dialkylglycine decarboxylase, aspartate aminotransferase, and tyrosine phenol-lyase. Five of the motifs corresponded to the alpha/beta elements and loops that form most of the conserved fold of the pyridoxal-P binding cleft of the four enzymes of known structure; the sixth motif corresponded to a helical element of the small domain that closes when the substrate binds. Eight residues that interact with pyridoxal-P and a ninth residue that lies at the interface of the large and small domains were also identified. Eleven additional conserved residues were identified and their functions were evaluated by examining the proteins of known structure. The key residues that interact directly with pyridoxal-P were identical in ornithine decarboxylase and the two GADs, thus allowing us to make a specific structural prediction of the cofactor binding site of GAD. The strong conservation of the cofactor binding site in GAD indicates that the highly regulated transition between apo- and holoGAD is accomplished by modifications in this basic fold rather than through a novel folding pattern.

Amino Acid Sequence↗

Sequence and structural analysis of cellular retinoic acid-binding proteins reveals a network of conserved hydrophobic interactions.

Proteins in the intracellular lipid-binding protein (iLBP) family show remarkably high structural conservation despite their low-sequence identity. A multiple-sequence alignment using 52 sequences of iLBP family members revealed 15 fully conserved positions, with a disproportionately high number of these (n=7) located in the relatively small helical region. The conserved positions displayed high structural conservation based on comparisons of known iLBP crystal structures. It is striking that the beta-sheet domain had few conserved positions, despite its high structural conservation. This observation prompted us to analyze pair-wise interactions within the beta-sheet region to ask whether structural information was encoded in interacting amino acid pairs. We conducted this analysis on the iLBP family member, cellular retinoic acid-binding protein I (CRABP I), whose folding mechanism is under study in our laboratory. Indeed, an analysis based on a simple classification of hydrophobic and polar amino acids revealed a network of conserved interactions in CRABP I that cluster spatially, suggesting a possible nucleation site for folding. Significantly, a small number of residues participated in multiple conserved interactions, suggesting a key role for these sites in the structure and folding of CRABP I. The results presented here correlate well with available experimental evidence on folding of CRABPs and their family members and suggest future experiments. The analysis also shows the usefulness of considering pair-wise conservation based on a simple classification of amino acids, in analyzing sequences and structures to find common core regions among homologues.

Amino Acid Sequence↗

Protein contact prediction using patterns of correlation.

We describe a new method for using neural networks to predict residue contact pairs in a protein. The main inputs to the neural network are a set of 25 measures of correlated mutation between all pairs of residues in two "windows" of size 5 centered on the residues of interest. While the individual pair-wise correlations are a relatively weak predictor of contact, by training the network on windows of correlation the accuracy of prediction is significantly improved. The neural network is trained on a set of 100 proteins and then tested on a disjoint set of 1033 proteins of known structure. An average predictive accuracy of 21.7% is obtained taking the best L/2 predictions for each protein, where L is the sequence length. Taking the best L/10 predictions gives an average accuracy of 30.7%. The predictor is also tested on a set of 59 proteins from the CASP5 experiment. The accuracy is found to be relatively consistent across different sequence lengths, but to vary widely according to the secondary structure. Predictive accuracy is also found to improve by using multiple sequence alignments containing many sequences to calculate the correlations.

Amino Acids↗