PubMed Health⌕ Search

Biomedical subjects

Marc R Wilkins

Publications and source records attributed to Marc R Wilkins.

7 recordsLinked to original sources

Guidelines for the next 10 years of proteomics.

In the last ten years, the field of proteomics has expanded at a rapid rate. A range of exciting new technology has been developed and enthusiastically applied to an enormous variety of biological questions. However, the degree of stringency required in proteomic data generation and analysis appears to have been underestimated. As a result, there are likely to be numerous published findings that are of questionable quality, requiring further confirmation and/or validation. This manuscript outlines a number of key issues in proteomic research, including those associated with experimental design, differential display and biomarker discovery, protein identification and analytical incompleteness. In an effort to set a standard that reflects current thinking on the necessary and desirable characteristics of publishable manuscripts in the field, a minimal set of guidelines for proteomics research is then described. These guidelines will serve as a set of criteria which editors of PROTEOMICS will use for assessment of future submissions to the Journal.

Biomarkers↗

Characterisation of organellar proteomes: a guide to subcellular proteomic fractionation and analysis.

Subcellular fractionation is being widely used to increase our understanding of the proteome. Fractionation is often coupled with 2-DE, thus allowing the visualisation of proteins and their subsequent identification and characterisation by MS. Whilst this strategy should be effective, to date, there has been little or no consideration given to differences in the mass, pI, hydropathy or abundance of proteins in the organelles and how analytical strategies can be tailored to match the idiosyncrasies of proteins in each particular compartment. To address this, we analysed 3962 Saccharomyces cerevisiae proteins, previously localised to one or more of 22 subcellular compartments. Different compartments showed significantly different distributions of protein pI and hydropathy. Mitochondrial and ER proteins showed the most dramatic differences to other organelles, in their protein pIs and hydropathy, respectively. We show that organelles can be clustered by similarities in these physicochemical protein characteristics. Interestingly, the distribution of protein abundance was also significantly different between many organelles. Our results show that to fully explore subcellular fractions of the proteome, specific analytical strategies should be employed. We outline strategies for all 22 subcellular compartments.

Cell Compartmentation↗

GlycoSuiteDB: a curated relational database of glycoprotein glycan structures and their biological sources. 2003 update.

GlycoSuiteDB is an annotated and curated relational database of glycan structures reported in the literature. It contains information on the glycan type, core type, linkages and anomeric configurations, mass, composition and the analytical methods used by the researchers to determine the glycan structure. Native and recombinant sources are detailed, including species, tissue and/or cell type, cell line, strain, life stage, disease, and if known the protein to which the glycan structures are attached. There are links to SWISS-PROT/TrEMBL and PubMed where applicable. Recent developments include the implementation of searching by 2D structure and substructure, disease and reference. The database is updated twice a year, and now contains over 7650 entries. Access to GlycoSuiteDB is available at http://www.glycosuite.com.

Animals↗

Identification of cellular changes associated with increased production of human growth hormone in a recombinant Chinese hamster ovary cell line.

A proteomics approach was used to identify the proteins potentially implicated in the cellular response concomitant with elevated production levels of human growth hormone in a recombinant Chinese hamster ovary (CHO) cell line following exposure to 0.5 mM butyrate and 80 microM zinc sulphate in the production media. This involved incorporation of two-dimensional (2-D) gel electrophoresis and protein identification by a combination of N-terminal sequencing, matrix-assisted laser desorption/ionisation-time of flight mass spectrometry, amino acid analysis and cross species database matching. From these identifications a CHO 2-D reference map and annotated database have been established. Metabolic labelling and subsequent autoradiography showed the induction of a number of cellular proteins in response to the media additives butyrate and zinc sulphate. These were identified as GRP75, enolase and thioredoxin. The chaperone proteins GRP78, HSP90, GRP94 and HSP70 were not up-regulated under these conditions.

Amino Acids↗

Unseen proteome: mining below the tip of the iceberg to find low abundance and membrane proteins.

Abundant and hydrophilic nonmembrane proteins with isoelectric points below pH 8 are the predominant proteins identified in most proteomics projects. In yeast, however, low-abundance proteins make up 80% of the predicted proteome, approximately 50% have pl's above pH 8 and 30% of the yeast ORFs are predicted to encode membrane proteins with at least 1 trans-membrane span. By applying highly solubilizing reagents and isoelectric fractionation to a membrane fraction of yeast we have a purified and identified 780 protein isoforms, representing 323 gene products, including 28% low abundance proteins and 49% membrane or membrane associated proteins. More importantly, considering the frequency and importance of co- and post-translational modifications, the separation of protein isoforms is essential and two-dimensional electrophoresis remains the only technique which offers sufficient resolution to address this at a proteomic level.

Amino Acid Sequence↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

Optimal replication and the importance of experimental design for gel-based quantitative proteomics.

Quantitative proteomic studies, based on two-dimensional gel electrophoresis, are commonly used to find proteins that are differentially expressed between samples or groups of samples. These proteins are of interest as potential diagnostic or prognostic biomarkers, or as proteins associated with a trait. The complexity of proteomic data poses many challenges, so while experiments may reveal proteins that are differentially expressed, these are often not significant when subjected to rigorous statistical analysis. However, this can be addressed through appropriate experimental design. A good experimental design considers the impact of different sources of variation, both analytical and biological, on the statistical importance of the results. The design should address the number of samples that must be analyzed and the number of replicate gels per sample, in the context of a particular minimum difference that one is seeking to achieve. In this study, we explore the ways to improve the quality of protein expression data from 2-DE gels, and describe an approach for defining the number of samples required and the number of gels per sample. It has been developed for the simplest of situations, two groups of samples with variation at two levels: between samples and between gels. This approach will also be useful as a guide for more complex designs involving more than two groups of samples. We describe some Internet-accessible tools that can assist in the design of proteomic studies.

Algorithms↗