PubMed Health⌕ Search

PubMed · 10869018

Predicting the oxidation state of cysteines by multiple sequence alignment.

Abstract

MOTIVATION: Protein sequences found in databanks usually do not report post translational covalent modifications such as the oxidation state of cystein (Cys) residues. Accurate prediction of whether a functionally or structurally important Cys occurs in the oxidized or thiol form would be helpful for molecular biology experiments and structure prediction. RESULTS: A new method is presented for predicting the oxidation state of Cys residues based on multiple sequence alignments and on the observation that Cys tends to occur in the same oxidation state within the same protein. The prediction of the redox state of Cys performs above 82%. The oxidation state of Cys correlates with the cellular location of the given protein within the cell, but the correlation is not perfect (up to 70%). We also perform a statistical analysis of the different redox states of Cys found in secondary structures and buried positions, and of the secondary structures linked by disulfide bonds. The results suggest that the natural borderline lies between the different oxidation states of Cys rather than between the half cystines and cysteins. AVAILABILITY: A web server implementing the prediction method is available at http://guitar.rockefeller.edu/approximately andras/cyspred.html CONTACT: fisera@rockefeller.edu

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Fiser, I Simon. 2000. Predicting the oxidation state of cysteines by multiple sequence alignment.. https://doi.org/10.1093/bioinformatics%2F16.3.251

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Base-pair resolution conservation data improves cell type specific sequence-to-expression prediction.

MOTIVATION: Genomic sequence-to-activity models can decipher gene regulatory mechanisms and predict the functional impact of regulatory variants. However, current models struggle to integrate information from sequences outside promoters, especially information from cell type specific regulatory elements. RESULTS: Here, we propose incorporating base-pair resolution evolutionary conservation data into genomic sequence-to-expression predictors. We explore two training strategies-training from scratch or fine-tuning an existing sequence-only model with additional conservation input. We find that in both cases, base-pair resolution conservation data improves cell type specific sequence-to-expression prediction, with training from scratch yielding the greatest benefit. The improvement in cell type specific expression prediction can be attributed in part to the fact that models trained on sequence and conservation data learn to better recognize cell type specific regulatory elements than models trained on sequence alone. AVAILABILITY: Code is available at https://github.com/ni-lab/basenji-phyloP.

Conserved Sequence↗

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence↗

New functional identity for the DNA uptake sequence in transformation and its presence in transcriptional terminators.

The frequently occurring DNA uptake sequence (DUS), recognized as a 10-bp repeat, is required for efficient genetic transformation in the human pathogens Neisseria meningitidis and Neisseria gonorrhoeae. Genome scanning for DUS occurrences in three different species of Neisseria demonstrated that 76% of the nearly 2,000 neisserial DUS were found to have two semiconserved base pairs extending from the 5' end of DUS to constitute a 12-mer repeat. Plasmids containing sequential variants of the neisserial DUS were tested for their ability to transform N. meningitidis and N. gonorrhoeae, and the 12-mer was found to outperform the 10-mer DUS in transformation efficiency. Assessment of meningococcal uptake of DNA confirmed the enhanced performance of the 12-mer compared to the 10-mer DUS. An inverted repeat DUS was not more efficient in transformation than DNA species containing a single or direct repeat DUS. Genome-wide analysis revealed that half of the nearly 1,500 12-mer DUS are arranged as inverted repeats predicted to be involved in rho-independent transcriptional termination or attenuation. The distribution of the uptake signal sequence required for transformation in the Pasteurellaceae was also biased towards transcriptional terminators, although to a lesser extent. In addition to assessing the intergenic location of DUS, we propose that the 10-mer identity of DUS should be extended and recognized as a 12-mer DUS. The dual role of DUS in transformation and as a structural component on RNA affecting transcription makes this a relevant model system for assessing significant roles of repeat sequences in biology.

Conserved Sequence↗