PubMed Health⌕ Search

Biomedical subjects

Simon J Hubbard

Publications and source records attributed to Simon J Hubbard.

17 recordsLinked to original sources

Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.

Protein identification via peptide mass fingerprinting (PMF) remains a key component of high-throughput proteomics experiments in post-genomic science. Candidate protein identifications are made using bioinformatic tools from peptide peak lists obtained via mass spectrometry (MS). These algorithms rely on several search parameters, including the number of potential uncut peptide bonds matching the primary specificity of the hydrolytic enzyme used in the experiment. Typically, up to one of these "missed cleavages" are considered by the bioinformatics search tools, usually after digestion of the in silico proteome by trypsin. Using two distinct, nonredundant datasets of peptides identified via PMF and tandem MS, a simple predictive method based on information theory is presented which is able to identify experimentally defined missed cleavages with up to 90% accuracy from amino acid sequence alone. Using this simple protocol, we are able to "mask" candidate protein databases so that confident missed cleavage sites need not be considered for in silico digestion. We show that that this leads to an improvement in database searching, with two different search engines, using the PMF dataset as a test set. In addition, the improved approach is also demonstrated on an independent PMF data set of known proteins that also has corresponding high-quality tandem MS data, validating the protein identifications. This approach has wider applicability for proteomics database searching, and the program for predicting missed cleavages and masking Fasta-formatted protein sequence databases has been made available via http:// ispider.smith.man.ac uk/MissedCleave.

Algorithms↗

Differential expression of ion channel transcripts in atrial muscle and sinoatrial node in rabbit.

The aim of the study was to identify ion channel transcripts expressed in the sinoatrial node (SAN), the pacemaker of the heart. Functionally, the SAN can be divided into central and peripheral regions (center is adapted for pacemaking only, whereas periphery is adapted to protect center and drive atrial muscle as well as pacemaking) and the aim was to study expression in both regions. In rabbit tissue, the abundance of 30 transcripts (including transcripts for connexin, Na(+), Ca(2+), hyperpolarization-activated cation and K(+) channels, and related Ca(2+) handling proteins) was measured using quantitative PCR and the distribution of selected transcripts was visualized using in situ hybridization. Quantification of individual transcripts (quantitative PCR) showed that there are significant differences in the abundance of 63% of the transcripts studied between the SAN and atrial muscle, and cluster analysis showed that the transcript profile of the SAN is significantly different from that of atrial muscle. There are apparent isoform switches on moving from atrial muscle to the SAN center: RYR2 to RYR3, Na(v)1.5 to Na(v)1.1, Ca(v)1.2 to Ca(v)1.3 and K(v)1.4 to K(v)4.2. The transcript profile of the SAN periphery is intermediate between that of the SAN center and atrial muscle. For example, Na(v)1.5 messenger RNA is expressed in the SAN periphery (as it is in atrial muscle), but not in the SAN center, and this is probably related to the need of the SAN periphery to drive the surrounding atrial muscle.

Animals↗

Global translational responses to oxidative stress impact upon multiple levels of protein synthesis.

Global inhibition of protein synthesis is a common response to stress conditions. We have analyzed the regulation of protein synthesis in response to oxidative stress induced by exposure to H(2)O(2) in the yeast Saccharomyces cerevisiae. Our data show that H(2)O(2) causes an inhibition of translation initiation dependent on the Gcn2 protein kinase, which phosphorylates the alpha-subunit of eukaryotic initiation factor-2. Additionally, our data indicate that translation is regulated in a Gcn2-independent manner because protein synthesis was still inhibited in response to H(2)O(2) in a gcn2 mutant. Polysome analysis indicated that H(2)O(2) causes a slower rate of ribosomal runoff, consistent with an inhibitory effect on translation elongation or termination. Furthermore, analysis of ribosomal transit times indicated that oxidative stress increases the average mRNA transit time, confirming a post-initiation inhibition of translation. Using microarray analysis of polysome- and monosome-associated mRNA pools, we demonstrate that certain mRNAs, including mRNAs encoding stress protective molecules, increase in association with ribosomes following H(2)O(2) stress. For some candidate mRNAs, we show that a low concentration of H(2)O(2) results in increased protein production. In contrast, a high concentration of H(2)O(2) promotes polyribosome association but does not necessarily lead to increased protein production. We suggest that these mRNAs may represent an mRNA store that could become rapidly activated following relief of the stress condition. In summary, oxidative stress elicits complex translational reprogramming that is fundamental for adaptation to the stress.

Eukaryotic Initiation Factor-2↗

PepSeeker: a database of proteome peptide identifications for investigating fragmentation patterns.

Proteome science relies on bioinformatics tools to characterize proteins via their proteolytic peptides which are identified via characteristic mass spectra generated after their ions undergo fragmentation in the gas phase within the mass spectrometer. The resulting secondary ion mass spectra are compared with protein sequence databases in order to identify the amino acid sequence. Although these search tools (e.g. SEQUEST, Mascot, X!Tandem, Phenyx) are frequently successful, much is still not understood about the amino acid sequence patterns which promote/protect particular fragmentation pathways, and hence lead to the presence/absence of particular ions from different ion series. In order to advance this area, we have developed a database, PepSeeker (http://nwsr.smith.man.ac.uk/pepseeker), which captures this peptide identification and ion information from proteome experiments. The database currently contains >185,000 peptides and associated database search information. Users may query this resource to retrieve peptide, protein and spectral information based on protein or peptide information, including the amino acid sequence itself represented by regular expressions coupled with ion series information. We believe this database will be useful to proteome researchers wishing to understand gas phase peptide ion chemistry in order to improve peptide identification strategies. Questions can be addressed to j.selley@manchester.ac.uk.

Databases, Protein↗

Analysis of gene expression in operons of Streptomyces coelicolor.

BACKGROUND: Recent studies have shown that microarray-derived gene-expression data are useful for operon prediction. However, it is apparent that genes within an operon do not conform to the simple notion that they have equal levels of expression. RESULTS: To investigate the relative transcript levels of intra-operonic genes, we have used a Z-score approach to normalize the expression levels of all genes within an operon to expression of the first gene of that operon. Here we demonstrate that there is a general downward trend in expression from the first to the last gene in Streptomyces coelicolor operons, in contrast to what we observe in Escherichia coli. Combining transcription-factor binding-site prediction with the identification of operonic genes that exhibited higher transcript levels than the first gene of the same operon enabled the discovery of putative internal promoters. The presence of transcription terminators and abundance of putative transcriptional control sequences in S. coelicolor operons are also described. CONCLUSION: Here we have demonstrated a polarity of expression in operons of S. coelicolor not seen in E. coli, bringing caution to those that apply operon prediction strategies based on E. coli 'equal-expression' to divergent species. We speculate that this general difference in transcription behavior could reflect the contrasting lifestyles of the two organisms and, in the case of Streptomyces, might also be influenced by its high G+C content genome. Identification of putative internal promoters, previously thought to cause problems in operon prediction strategies, has also been enabled.

Bacterial Proteins↗

Global gene expression profiling reveals widespread yet distinctive translational responses to different eukaryotic translation initiation factor 2B-targeting stress pathways.

Global inhibition of protein synthesis is a hallmark of many cellular stress conditions. Even though specific mRNAs defy this (e.g., yeast GCN4 and mammalian ATF4), the extent and variation of such resistance remain uncertain. In this study, we have identified yeast mRNAs that are translationally maintained following either amino acid depletion or fusel alcohol addition. Both stresses inhibit eukaryotic translation initiation factor 2B, but via different mechanisms. Using microarray analysis of polysome and monosome mRNA pools, we demonstrate that these stress conditions elicit widespread yet distinct translational reprogramming, identifying a fundamental role for translational control in the adaptation to environmental stress. These studies also highlight the complex interplay that exists between different stages in the gene expression pathway to allow specific preordained programs of proteome remodeling. For example, many ribosome biogenesis genes are coregulated at the transcriptional and translational levels following amino acid starvation. The transcriptional regulation of these genes has recently been connected to the regulation of cellular proliferation, and on the basis of our results, the translational control of these mRNAs should be factored into this equation.

Amino Acids↗

Conservation of orientation and sequence in protein domain--domain interactions.

The repertoire of naturally occurring protein structures is usually characterised in structural terms at the domain level by their constituent folds. As structure is acknowledged to be an important stepping stone to the understanding of protein function, an appreciation of how individual domain interactions are built to form complete, functional protein structures is essential. A comprehensive study of protein domain interactions has been undertaken, covering all those observed in known structures, as well as those predicted to occur in 46 completed genome sequences from all three domains of life. In particular, we examine the promiscuity of protein domains characterised by SCOP superfamilies in terms of their interacting partners, the surface they use to form these interactions, and the relative orientations of their domain partners. Protein domains are shown to display a variety of behaviours, ranging from high promiscuity to absolute monogamy of domain surface employed, with both multiple and single domain partners. In addition, the conservation of sequence and volume at domain interface surfaces is observed to be significantly higher than at accessible surface in general, acting as a powerful potential predictor for domain interactions. We also examine the separation of interacting domains in protein sequence, showing that standard thresholds of 30 amino acid residues lead to a significant false positive rate, and an even more significant false negative rate of approximately 40%. These data suggest that there may be many more than the 2000 domain--domain interactions that have not yet been observed structurally, and we provide a top 30 hit-list of putative domain interactions which should be targeted.

Amino Acid Motifs↗

Transcriptome analysis for the chicken based on 19,626 finished cDNA sequences and 485,337 expressed sequence tags.

We present an analysis of the chicken (Gallus gallus) transcriptome based on the full insert sequences for 19,626 cDNAs, combined with 485,337 EST sequences. The cDNA data set has been functionally annotated and describes a minimum of 11,929 chicken coding genes, including the sequence for 2260 full-length cDNAs together with a collection of noncoding (nc) cDNAs that have been stringently filtered to remove untranslated regions of coding mRNAs. The combined collection of cDNAs and ESTs describe 62,546 clustered transcripts and provide transcriptional evidence for a total of 18,989 chicken genes, including 88% of the annotated Ensembl gene set. Analysis of the ncRNAs reveals a set that is highly conserved in chickens and mammals, including sequences for 14 pri-miRNAs encoding 23 different miRNAs. The data sets described here provide a transcriptome toolkit linked to physical clones for bioinformaticians and experimental biologists who wish to use chicken systems as a low-cost, accessible alternative to mammals for the analysis of vertebrate development, immunology, and cell biology.

Animals↗

Improved prediction for N-termini of alpha-helices using empirical information.

The prediction of the secondary structure of proteins from their amino acid sequences remains a key component of many approaches to the protein folding problem. The most abundant form of regular secondary structure in proteins is the alpha-helix, in which specific residue preferences exist at the N-terminal locations. Propensities derived from these observed amino acid frequencies in the Protein Data Bank (PDB) database correlate well with experimental free energies measured for residues at different N-terminal positions in alanine-based peptides. We report a novel method to exploit this data to improve protein secondary structure prediction through identification of the correct N-terminal sequences in alpha-helices, based on existing popular methods for secondary structure prediction. With this algorithm, the number of correctly predicted alpha-helix start positions was improved from 30% to 38%, while the overall prediction accuracy (Q3) remained the same, using cross-validated testing. Although the algorithm was developed and tested on multiple sequence alignment-based secondary structure predictions, it was also able to improve the predictions of start locations by methods that use single sequences to make their predictions. Furthermore, the residue frequencies at N-terminal positions of the improved predictions better reflect those seen at the N-terminal positions of alpha-helices in proteins. This has implications for areas such as comparative modeling, where a more accurate prediction of the N-terminal regions of alpha-helices should benefit attempts to model adjacent loop regions. The algorithm is available as a Web tool, located at http://rocky.bms.umist.ac.uk/elephant.

Databases, Protein↗

PEDRo: a database for storing, searching and disseminating experimental proteomics data.

BACKGROUND: Proteomics is rapidly evolving into a high-throughput technology, in which substantial and systematic studies are conducted on samples from a wide range of physiological, developmental, or pathological conditions. Reference maps from 2D gels are widely circulated. However, there is, as yet, no formally accepted standard representation to support the sharing of proteomics data, and little systematic dissemination of comprehensive proteomic data sets. RESULTS: This paper describes the design, implementation and use of a Proteome Experimental Data Repository (PEDRo), which makes comprehensive proteomics data sets available for browsing, searching and downloading. It is also serves to extend the debate on the level of detail at which proteomics data should be captured, the sorts of facilities that should be provided by proteome data management systems, and the techniques by which such facilities can be made available. CONCLUSIONS: The PEDRo database provides access to a collection of comprehensive descriptions of experimental data sets in proteomics. Not only are these data sets interesting in and of themselves, they also provide a useful early validation of the PEDRo data model, which has served as a starting point for the ongoing standardisation activity through the Proteome Standards Initiative of the Human Proteome Organisation.

Animals↗

SiteSeer: Visualisation and analysis of transcription factor binding sites in nucleotide sequences.

The regulation of gene expression is a fundamental process within every living cell, which allows organisms to manage the precise levels of functional gene products with high sensitivity. It is well established that specific DNA sequences located upstream of the transcriptional start site are important in facilitating the binding of regulatory proteins that control the transcription of the gene. Indeed, microarray-based studies have successfully mined the upstream regions of co-expressed genes and discovered over-represented sequences corresponding to known promoter sites. Here we describe a tool for the visualisation of mapped transcription factor binding sites in the upstream regions of either single or grouped eukaryotic genes, which allows users to examine the positions of known and user-defined sites (http://rocky.bms.umist.ac.uk/SiteSeer/). SiteSeer allows the user to map different sections of the TRANSFAC and SCPD databases (or a set of user-defined sites) onto nucleotide sequences. Additionally, users may restrict the analysis by expectation values for certain DNA words as well as by known binding sites specific to a given organism. We believe this tool will prove particularly valuable for biologists who wish to examine sets of co-expressed or functionally-related genes and those who wish to visualise the positions of promoter sequences and generate displays for publications.

5' Flanking Region↗

A systematic approach to modeling, capturing, and disseminating proteomics experimental data.

Both the generation and the analysis of proteome data are becoming increasingly widespread, and the field of proteomics is moving incrementally toward high-throughput approaches. Techniques are also increasing in complexity as the relevant technologies evolve. A standard representation of both the methods used and the data generated in proteomics experiments, analogous to that of the MIAME (minimum information about a microarray experiment) guidelines for transcriptomics, and the associated MAGE (microarray gene expression) object model and XML (extensible markup language) implementation, has yet to emerge. This hinders the handling, exchange, and dissemination of proteomics data. Here, we present a UML (unified modeling language) approach to proteomics experimental data, describe XML and SQL (structured query language) implementations of that model, and discuss capture, storage, and dissemination strategies. These make explicit what data might be most usefully captured about proteomics experiments and provide complementary routes toward the implementation of a proteome repository.

Database Management Systems↗

A comprehensive collection of chicken cDNAs.

Birds have played a central role in many biological disciplines, particularly ecology, evolution, and behavior. The chicken, as a model vertebrate, also represents an important experimental system for developmental biologists, immunologists, cell biologists, and geneticists. However, genomic resources for the chicken have lagged behind those for other model organisms, with only 1845 nonredundant full-length chicken cDNA sequences currently deposited in the EMBL databank. We describe a large-scale expressed-sequence-tag (EST) project aimed at gene discovery in chickens (http://www.chick.umist.ac.uk). In total, 339,314 ESTs have been sequenced from 64 cDNA libraries generated from 21 different embryonic and adult tissues. These were clustered and assembled into 85,486 contiguous sequences (contigs). We find that a minimum of 38% of the contigs have orthologs in other organisms and define an upper limit of 13,000 new chicken genes. The remaining contigs may include novel avian specific or rapidly evolving genes. Comparison of the contigs with known chicken genes and orthologs indicates that 30% include cDNAs that contain the start codon and 20% of the contigs represent full-length cDNA sequences. Using this dataset, we estimate that chickens have approximately 35,000 genes in total, suggesting that this number may be a characteristic feature of vertebrates.

Animals↗

Stable isotope labelling in vivo as an aid to protein identification in peptide mass fingerprinting.

Peptide mass fingerprinting (PMF) is a powerful technique for identification of proteins derived from in-gel digests by virtue of their matrix-assisted laser desorption/ionization-time of flight mass spectra. However, there are circumstances where the under-representation of peptides in the mass spectrum and the complexity of the source proteome mean that PMF is inadequate as an identification tool. In this paper, we show that identification is substantially enhanced by inclusion of composition data for a single amino acid. Labelling in vivo with a stable isotope labelled amino acid (in this paper, decadeuterated leucine) identifies the number of such amino acids in each digest fragment, and show a considerable gain in the ability of PMF to identify the parent protein. The method is tolerant to the extent of labelling, and as such, may be applicable to a range of single cell systems.

Deuterium↗

Comparative bioinformatic analysis of complete proteomes and protein parameters for cross-species identification in proteomics.

Peptide mass fingerprinting (PMF) remains the most amenable technique for protein identification in proteomics, using mass spectrometry as the primary analytical technique coupled with bioinformatics. This relies on the presence of the amino acid sequence of the protein in the current databanks. Despite this, it is desirable to be able to use the technique for organisms whose genomes are not yet fully sequenced and apply cross-species protein identification. In this study, we have re-examined the feasibility of such approaches by considering the extent of protein similarity between genome sequences using a data set of 29 complete bacterial and two eukaryotic genomes. A range of protein and peptide features are considered, including protein isoelectric focussing point, protein mass, and amino acid conservation. The effectiveness of PMF approaches has then been tested with a series of computer simulations with varying peptide number and mass accuracy for several cross-species tests. The results show that PMF alone is unsuitable in general for divergent species jumps, or when protein similarity is less than 70% identity. Despite this, there exists a considerable enrichment above random of tryptic peptide conservation and PMF promises to remain useful when combined with other data than just peptide masses for cross-species protein identification.

Algorithms↗

A critical assessment of the secondary structure alpha-helices and their termini in proteins.

Secondary structure prediction from amino acid sequence is a key component of protein structure prediction, with current accuracy at approximately 75%. We analysed two state-of-the-art secondary structure prediction methods, PHD and JPRED, comparing predictions with secondary structure assigned by the algorithms DSSP and STRIDE. The specific focus of our study was alpha-helix N-termini, as empirical free energy scales are available for residue preferences at N-terminal positions. Although these prediction methods perform well in general at predicting the alpha-helical locations and length distributions in proteins, they perform less well at predicting the correct helical termini. For example, although most predicted alpha-helices overlap a real alpha-helix (with relatively few completely missed or extra predicted helices), only one-third of JPRED and PHD predictions correctly identify the N-terminus. Analysis of neighbouring N-terminal sequences to predicted helical N-termini shows that the correct N-terminus is often within one or two residues. More importantly, the true N-terminal motif is, on average, more favourable as judged by our experimentally measured free energies. This suggests a simple, but powerful, strategy to improve secondary structure prediction using empirically derived energies to adjust the predicted output to a more favourable N-terminal sequence.

Algorithms↗