PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Method for prediction of protein function from sequence using the sequence-to-structure-to-function paradigm with application to glutaredoxins/thioredoxins and T1 ribonucleases.

The practical exploitation of the vast numbers of sequences in the genome sequence databases is crucially dependent on the ability to identify the function of each sequence. Unfortunately, current methods, including global sequence alignment and local sequence motif identification, are limited by the extent of sequence similarity between sequences of unknown and known function; these methods increasingly fail as the sequence identity diverges into and beyond the twilight zone of sequence identity. To address this problem, a novel method for identification of protein function based directly on the sequence-to-structure-to-function paradigm is described. Descriptors of protein active sites, termed "fuzzy functional forms" or FFFs, are created based on the geometry and conformation of the active site. By way of illustration, the active sites responsible for the disulfide oxidoreductase activity of the glutaredoxin/thioredoxin family and the RNA hydrolytic activity of the T1 ribonuclease family are presented. First, the FFFs are shown to correctly identify their corresponding active sites in a library of exact protein models produced by crystallography or NMR spectroscopy, most of which lack the specified activity. Next, these FFFs are used to screen for active sites in low-to-moderate resolution models produced by ab initio folding or threading prediction algorithms. Again, the FFFs can specifically identify the functional sites of these proteins from their predicted structures. The results demonstrate that low-to-moderate resolution models as produced by state-of-the-art tertiary structure prediction algorithms are sufficient to identify protein active sites. Prediction of a novel function for the gamma subunit of a yeast glycosyl transferase and prediction of the function of two hypothetical yeast proteins whose models were produced via threading are presented. This work suggests a means for the large-scale functional screening of genomic sequence databases based on the prediction of structure from sequence, then on the identification of functional active sites in the predicted structure.

Algorithms↗

Group specific PCR-detection of potential trichothecene-producing Fusarium-species in pure cultures and cereal samples.

A PCR based assay (Tox5 PCR) which analyses Fusarium species potentially producing trichothecenes was developed using a pair of primers derived from the DNA-sequence of the trichodiene synthase gene (tri5). The primer pair was tested using DNA isolated from a variety of strains representing 64 species and varieties of Fusarium as well as from other fungi, bacteria and cereals. A 658 bp PCR fragment was specifically amplified with DNA isolated from strains of species belonging to the Fusarium sections Discolor, Sporotrichiella, Arthrosporiella, Gibbosum, and "Dlaminia". PCR products obtained were sequenced. Alignment to tri5 sequences given in the literature revealed a high degree of homology. Results of the PCR developed correlated well with literature data on the trichothecene producing capabilities of the respective species. Potential trichothecene producing fusaria were detected in contaminated cereals and malts using the Tox5 PCR assay. Intensity of the signals produced were well correlated with the concentration of deoxynivalenol (DON) in samples of wheat.

Base Sequence↗

Identification and kinetics analysis of a novel heparin-binding site (KEDK) in human tenascin-C.

The interaction between tenascin-C (TN-C), a multi-subunit extracellular matrix protein, and heparin was examined using a surface plasmon resonance-based technique on a Biacore system. The aims of the present study were to examine the affinity of fibronectin type III repeats of TN-C fragments (TNIII) for heparin, to investigate the role of the TNIII4 domains in the binding of TN-C to heparin, and to delineate a sequence of amino acids within the TNIII4 domain, which mediates cooperative heparin binding. At a physiological salt concentration, and pH 7.4, TNIII3-5 binds to heparin with high affinity (K(D) = 30 nm). However, a major heparin-binding site in TNIII5 produces a modest affinity binding at a K(D) near 4 microm, and a second site in TNIII4 enhances the binding by several orders of magnitude, although it was far too weak to produce an observable binding of TNIII4 by itself. Moreover, mutagenesis of the KEDK sequence in the TNIII4 domain resulted in the significant reduction of heparin-binding affinity. In addition, residues in the KEDK sequences are conserved in TN-C throughout mammalian evolution. Thus the structure-based sequence alignment, mutagenesis, and sequence conservation data together reveal a KEDK sequence in TNIII4 suggestive of a minor heparin-binding site. Finally, we demonstrate that TNIII4 contains binding sites for heparin sulfate proteoglycan and enhances the heparin sulfate proteoglycan-dependent human gingival fibroblast adhesion to TNIII5, thus providing the biological significance of heparin-binding site of TNIII4. These results suggest that the heparin-binding sites may traverse TNIII4-5 and thus require KEDK in TNIII4 for optimal heparin-binding.

Amino Acid Sequence↗

Quantification of secondary structure prediction improvement using multiple alignments.

The use of multiple sequence alignments for secondary structure predictions is analysed. Seven different protein families, containing only sequences of known structure, were considered to provide a range of alignment and prediction conditions. Using alignments obtained by spatial superposition of main chain atoms in known tertiary protein structures allowed a mean of 8% in secondary structure prediction accuracy, when compared to those obtained from the individual sequences. Substitution of these alignments by those determined directly from an automated sequence alignment algorithm showed variations in the prediction accuracy which correlated with the quality of the multiple alignments and distance of the primary sequence. Secondary structure predictions can be reliably improved using alignments from an automatic alignment procedure with a mean increase of 6.8%, giving an overall prediction accuracy of 68.5%, if there is a minimum of 25% sequence identity between all sequences in a family.

Amino Acid Sequence↗

Divide-and-conquer multiple alignment with segment-based constraints.

A large number of methods for multiple sequence alignment are currently available. Recent benchmarking tests demonstrated that strengths and drawbacks of these methods differ substantially. Global strategies can be outperformed by approaches based on local similarities and vice versa, depending on the characteristics of the input sequences. In recent years, mixed approaches that include both global and local features have shown promising results. Herein, we introduce a new algorithm for multiple sequence alignment that integrates the global divide-and-conquer approach with the local segment-based approach, thereby combining the strengths of those two strategies.

Algorithms↗

A note on clustering the functionally-related paralogues and orthologues of proteins: a case of the FK506-binding proteins (FKBPs).

The expression patterns of 18 FK506-binding proteins (FKBPs) encoded in the human genome have been established whereas the functional significance of the numerous ORFs coding for FKBP-like sequences remains unknown. Nominal masses of the human FKBPs vary from 12 to 135 kDa. Some large FKBPs consist up to four repeats of the 12 kDa FK506-like binding domain (FKBD) whereas other large FKBPs contain one FKBD linked to different functional domains such as TPRs, leucine-zipper, calmodulin-binding domain etc. The genomes of other eukaryotic organisms, namely D. melanogaster, C. elegans, A. thaliana, S. pombe and S. cerevisiae encode different numbers of the FKBPs' paralogues some of which are orthologues to the human FKBPs. A library of novel algorithms was developed and used for computation of the level of conservation of the hydrophobicity and bulkiness profiles, and the amino acid compositions (AACs) of 247 aligned sequences of FKBPs. The pairwisely-compared hydrophobicity and bulkiness profiles for some combinations of the aligned sequences of the FKBDs yielded high values of the correlation coefficients (CCF). The AACs of some combinations of the aligned sequences of the FKBDs also differed to a low degree. The functionally-related orthologues and paralogues of the FKBPs were clustered by using the following criteria: 1 degrees apparent conservation of the crucial amino acid (AA) residues for peptidylprolyl cis/trans isomerase (PPIase) acitity and binding of some immunosuppressive drugs; 2 degrees convergence of the three mentioned above properties of the polypeptide chain; 3 degrees similarity in the sequence attributes pI and total hydrophobicity index (HI). The clustering method was used for setting up several hypotheses on the emergence of certain classes of the FKBPs in the eukaryotic kingdom.

Algorithms↗

ProMSED: protein multiple sequence editor for Windows 3.11/95.

MOTIVATION: Most protein sequence alignment algorithms give similar results on closely related proteins, while manual intervention may be needed for distantly related molecules. To correct the alignment, it is often necessary to repeat calculations on selected parts of the alignments and edit the alignment manually. Software implementing such interactive alignment procedures is of significance. RESULTS: This paper presents a new MS Windows application called ProMSED for both automatic and manual protein sequence alignment. The program reads main sequence formats and has a user-friendly interface. ProMSED performs automatic (ClustalV algorithm) alignments, alignment visualization and editing, and it allows sequences to be aligned interactively leaving previously aligned regions unchanged. Manual alignment and sequence analysis are facilitated by colouring schemes reflecting amino acid similarity of mutational and physicochemical properties. The interactive alignment of a diverged set of reverse transcriptases has located four out of six known conserved motifs. AVAILABILITY: ProMSED is available on request from the authors. DEMO is available from ftp://ftp.ebi.ac.uk/pub/ software/dos/promsed/ or ftp://iubio.bio.indiana.edu/molbio/ ibmpc/.

Algorithms↗

Enrichment of regulatory signals in conserved non-coding genomic sequence.

MOTIVATION: Whole genome shotgun sequencing strategies generate sequence data prior to the application of assembly methodologies that result in contiguous sequence. Sequence reads can be employed to indicate regions of conservation between closely related species for which only one genome has been assembled. Consequently, by using pairwise sequence alignments methods it is possible to identify novel, non-repetitive, conserved segments in non-coding sequence that exist between the assembled human genome and mouse whole genome shotgun sequencing fragments. Conserved non-coding regions identify potentially functional DNA that could be involved in transcriptional regulation. RESULTS: Local sequence alignment methods were applied employing mouse fragments and the assembled human genome. In addition, transcription factor binding sites were detected by aligning their corresponding positional weight matrices to the sequence regions. These methods were applied to a set of transcripts corresponding to 502 genes associated with a variety of different human diseases taken from the Online Mendelian Inheritance in Man database. Using statistical arguments we have shown that conserved non-coding segments contain an enrichment of transcription factor binding sites when compared to the sequence background in which the conserved segments are located. This enrichment of binding sites was not observed in coding sequence. Conserved non-coding segments are not extensively repeated in the genome and therefore their identification provides a rapid means of finding genes with related conserved regions, and consequently potentially related regulatory mechanism. Conserved segments in upstream regions are found to contain binding sites that are co-localized in a manner consistent with experimentally known transcription factor pairwise co-occurrences and afford the identification of novel co-occurring Transcription Factor (TF) pairs. This study provides a methodology and more evidence to suggest that conserved non-coding regions are biologically significant since they contain a statistical enrichment of regulatory signals and pairs of signals that enable the construction of regulatory models for human genes. CONTACT: samuel.levy@celera.com.

Algorithms↗

Automated 1H and 13C chemical shift prediction using the BioMagResBank.

A computer program has been developed to accurately and automatically predict the 1H and 13C chemical shifts of unassigned proteins on the basis of sequence homology. The program (called SHIFTY) uses standard sequence alignment techniques to compare the sequence of an unassigned protein against the BioMagResBank--a public database containing sequences and NMR chemical shifts of nearly 200 assigned proteins [Seavey et al. (1991) J Biomol. NMR, 1, 217-236]. From this initial sequence alignment, the program uses a simple set of rules to directly assign or transfer a complete set of 1H or 13C chemical shifts (from the previously assigned homologues) to the unassigned protein. This 'homologous assignment' protocol takes advantage of the simple fact that homologous proteins tend to share both structural similarity and chemical shift similarity. SHIFTY has been extensively tested on more than 25 medium-sized proteins. Under favorable circumstances, this program can predict the 1H or 13C chemical shifts of proteins with an accuracy far exceeding any other method published to date. With the exponential growth in the number of assigned proteins appearing in the literature (now at a rate of more than 150 per year), we believe that SHIFTY may have widespread utility in assigning individual members in families of related proteins, an endeavor that accounts for a growing portion of the protein NMR work being done today.

Algorithms↗

Tree decomposition based fast search of RNA structures including pseudoknots in genomes.

Searching genomes for RNA secondary structure with computational methods has become an important approach to the annotation of non-coding RNAs. However, due to the lack of efficient algorithms for accurate RNA structure-sequence alignment, computer programs capable of fast and effectively searching genomes for RNA secondary structures have not been available. In this paper, a novel RNA structure profiling model is introduced based on the notion of a conformational graph to specify the consensus structure of an RNA family. Tree decomposition yields a small tree width t for such conformation graphs (e.g., t = 2 for stem loops and only a slight increase for pseudo-knots). Within this modelling framework, the optimal alignment of a sequence to the structure model corresponds to finding a maximum valued isomorphic subgraph and consequently can be accomplished through dynamic programming on the tree decomposition of the conformational graph in time O(k(t)N(2)), where k is a small parameter; and N is the size of the projiled RNA structure. Experiments show that the application of the alignment algorithm to search in genomes yields the same search accuracy as methods based on a Covariance model with a significant reduction in computation time. In particular; very accurate searches of tmRNAs in bacteria genomes and of telomerase RNAs in yeast genomes can be accomplished in days, as opposed to months required by other methods. The tree decomposition based searching tool is free upon request and can be downloaded at our site h t t p ://w.uga.edu/RNA-informatics/software/index.php.

Algorithms↗

Molecular and cellular characterization of CRP1, a Drosophila chromatin decondensation protein.

CRP1, a Drosophila nuclear protein that can catalyze decondensation of demembranated Xenopus sperm chromatin was cloned and its primary structure was deduced from cDNA sequence. Alignment of deduced amino acid sequence with published sequences of other proteins revealed strong homologies to Xenopus nucleoplasmin and NO38. CRP1 is encoded by one or several closely related genes found at a single locus, position 99A on the right arm of chromosome 3. CRP1 mRNA is expressed throughout Drosophila development; it is highest during oogenesis and early embryogenesis. mRNA levels correlate closely with levels of protein expression measured previously. Results of chemical crosslinking indicate that CRP1 is either tetrameric or pentameric; similar ambiguity was revealed by direct visualization using scanning transmission electron microscopy. Consistent with previously published results, parallel crosslinking studies of Xenopus nucleoplasmin suggested a pentameric structure. Scanning transmission electron microscopic examination after negative staining revealed that CRP1 and Xenopus nucleoplasmin are morphologically similar. CRP1 is able to substitute for nucleoplasmin in Xenopus egg extract-mediated sperm chromatin decondensation. In vitro, CRP1-induced decondensation is accompanied by direct binding of CRP1 to chromatin.

Amino Acid Sequence↗

Cloning and cDNA sequence analysis of Lys(49) and Asp(49) basic phospholipase A(2) myotoxin isoforms from Bothrops asper.

Snake venom myotoxic phospholipases A(2) contribute to much of the tissue damage observed during envenomation by Bothrops asper, the major cause of snake bites in Central America. Several myotoxic PLA(2)s have been identified, but their mechanism of myotoxicity is still unclear. To aid in the molecular characterization of these venom toxins, the complete open reading frames encoding two Lys(49) and one Asp(49) basic PLA(2) myotoxins from the Central American snake B. asper (terciopelo) were obtained by cDNA cloning from venom gland poly-adenylated RNA. The amino acid sequence deduced from the myotoxins II and III open reading frames corresponded in each case to one of the reported amino acid sequence isoforms. The sequence of a new myotoxin IV-like sequence (MT-IVa) contains conservative Val-->Leu(18) and Ala-->Val(23) substitutions when compared with the reported N-terminus of the native myotoxin IV, suggesting minor isoform variations among specimens of a single species. Sequence alignment studies indicated significant (>75% sequence identity) identities with other crotalid venom Lys(49) PLA(2)s, particularly bothropstoxin I/Ia isoforms of B. jararacussu and myotoxin II of B. asper.

Amino Acid Sequence↗

Cloning and sequencing of coat protein gene of an Indian potato leaf roll virus (PLRV) isolate and its similarity with other members of Luteoviridae.

An Indian strain of potato leaf roll virus (PLRV) was purified to generate complementary DNA corresponding to the coat protein (CP) gene. Virus cDNA was synthesized from purified viral RNA using oligo (dT)-anchor primer and virus specific primers. The viral sequence encoding the coat protein was specifically amplified by polymerase chain reaction (PCR), using specific primers bordering the CP gene. The unique amplified product thus obtained was A-T cloned into the pGEM-T Easy vector and the authenticity of the cloned gene verified by dot blot hybridization and sequence analysis. Run-way-transcripts of the cloned CP gene could detect PLRV in tissue imprints and tissue dilution. The nucleotide sequences and the deduced amino acid sequences were compared with the other PLRV isolates and found to be 97-99% identical at both the nucleotide and amino acid sequence level of other isolates. Multiple sequence alignment of deduced amino acid sequences revealed considerable homology to other luteoviruses. A nuclear localization signal located close to the N-terminus of the CP gene was predicted. This is the first report of PLRV coat protein sequence from an Indian strain.

Amino Acid Sequence↗

MatrixPlot: visualizing sequence constraints.

UNLABELLED: MatrixPlot is a program for making high-quality matrix plots, such as mutual information plots of sequence alignments and distance matrices of sequences with known three-dimensional coordinates. The user can add information about the sequences (e.g. a sequence logo profile) along the edges of the plot, as well as zoom in on any region in the plot. AVAILABILITY: MatrixPlot can be obtained on request, and can also be accessed online at http://www. cbs.dtu.dk/services/MatrixPlot. CONTACT: gorodkin@cbs.dtu.dk

Nucleic Acids↗

The structure of Rauvolfia serpentina strictosidine synthase is a novel six-bladed beta-propeller fold in plant proteins.

The enzyme strictosidine synthase (STR1) from the Indian medicinal plant Rauvolfia serpentina is of primary importance for the biosynthetic pathway of the indole alkaloid ajmaline. Moreover, STR1 initiates all biosynthetic pathways leading to the entire monoterpenoid indole alkaloid family representing an enormous structural variety of approximately 2000 compounds in higher plants. The crystal structures of STR1 in complex with its natural substrates tryptamine and secologanin provide structural understanding of the observed substrate preference and identify residues lining the active site surface that contact the substrates. STR1 catalyzes a Pictet-Spengler-type reaction and represents a novel six-bladed beta-propeller fold in plant proteins. Structure-based sequence alignment revealed a common repetitive sequence motif (three hydrophobic residues are followed by a small residue and a hydrophilic residue), indicating a possible evolutionary relationship between STR1 and several sequence-unrelated six-bladed beta-propeller structures. Structural analysis and site-directed mutagenesis experiments demonstrate the essential role of Glu-309 in catalysis. The data will aid in deciphering the details of the reaction mechanism of STR1 as well as other members of this enzyme family.

Amino Acid Sequence↗