PubMed Health⌕ Search

Biomedical subjects

Ron D Appel

Publications and source records attributed to Ron D Appel.

At least 19 recordsLinked to original sources

Guidelines for the next 10 years of proteomics.

In the last ten years, the field of proteomics has expanded at a rapid rate. A range of exciting new technology has been developed and enthusiastically applied to an enormous variety of biological questions. However, the degree of stringency required in proteomic data generation and analysis appears to have been underestimated. As a result, there are likely to be numerous published findings that are of questionable quality, requiring further confirmation and/or validation. This manuscript outlines a number of key issues in proteomic research, including those associated with experimental design, differential display and biomarker discovery, protein identification and analytical incompleteness. In an effort to set a standard that reflects current thinking on the necessary and desirable characteristics of publishable manuscripts in the field, a minimal set of guidelines for proteomics research is then described. These guidelines will serve as a set of criteria which editors of PROTEOMICS will use for assessment of future submissions to the Journal.

Biomarkers↗

Proteome informatics I: bioinformatics tools for processing experimental data.

Bioinformatics tools for proteomics, also called proteome informatics tools, span today a large panel of very diverse applications ranging from simple tools to compare protein amino acid compositions to sophisticated software for large-scale protein structure determination. This review considers the available and ready to use tools that can help end-users to interpret, validate and generate biological information from their experimental data. It concentrates on bioinformatics tools for 2-DE analysis, for LC followed by MS analysis, for protein identification by PMF, by peptide fragment fingerprinting and by de novo sequencing and for data quantitation with MS data. It also discloses initiatives that propose to automate the processes of MS analysis and enhance the quality of the obtained results.

Algorithms↗

Proteome informatics II: bioinformatics for comparative proteomics.

The present review attempts to cover the most recent initiatives directed towards representing, storing, displaying and processing protein-related data suited to undertake "comparative proteomics" studies. Data interpretation is brought into focus. Efforts invested into analysing and interpreting experimental data increasingly express the need for adding meaning. This trend is perceptible in work dedicated to determining ontologies, modelling interaction networks, etc. In parallel, technical advances in computer science are spurred by the development of the Web and the growing need to channel and understand massive volumes of data. Biology benefits from these advances as an application of choice for many generic solutions. Some examples of bioinformatics solutions are discussed and directions for on-going and future work conclude the review.

Algorithms↗

InSilicoSpectro: an open-source proteomics library.

We present a new proteomics open-source project, InSilicoSpectro, aimed at implementing recurrent computations that are necessary for proteomics data analysis. Illustrative examples are mass list file format conversions, protein sequence digestion, theoretical peptide and fragment mass computations, graphical display, matching with experimental data, isoelectric point estimation, and peptide retention time prediction. The project library is written in Perl, a widely used scripting language in bioinformatics, and it offers a unique framework of integrated objects to implement complex proteomics data analyses. For instance, only a few lines of code are required to digest a protein with fixed and variable modifications, label peptides with 18O, compute the fragmentation spectra and display their match with experimental spectra. We believe that InSilicoSpectro will be of great help to bioinformaticians, without detailed knowledge of proteomics specifics, and to mass spectrometrists with computer programming interest as well.

Amino Acid Sequence↗

MSight: an image analysis software for liquid chromatography-mass spectrometry.

Images obtained from high-throughput mass spectrometry (MS) contain information that remains hidden when looking at a single spectrum at a time. Image processing of liquid chromatography-MS datasets can be extremely useful for quality control, experimental monitoring and knowledge extraction. The importance of imaging in differential analysis of proteomic experiments has already been established through two-dimensional gels and can now be foreseen with MS images. We present MSight, a new software designed to construct and manipulate MS images, as well as to facilitate their analysis and comparison.

Chromatography, Liquid↗

Correlation of proteomic and transcriptomic profiles of Staphylococcus aureus during the post-exponential phase of growth.

A combined proteomic and transcriptomic analysis of Staphylococcus aureus strain N315 was performed to study a sequenced strain at the system level. Total protein and membrane protein extracts were prepared and analyzed using various proteomic workflows including: 2-DE, SDS-PAGE combined with microcapillary LC-MALDI-MS/MS, and multidimensional liquid chromatography. The presence of a protein was then correlated with its respective transcript level from S. aureus cells grown under the same conditions. Gene-expression data revealed that 97% of the 2'596 ORFs were detected during the post-exponential phase. At the protein level, 23% of these ORFs (591 proteins) were identified. Correlation of the two datasets revealed that 42% of the identified proteins (248 proteins) were amongst the top 25% of genes with highest mRNA signal intensities, and 69% of the identified proteins (406 proteins) were amongst the top 50% with the highest mRNA signal intensities. The fact that the remaining 31% of proteins were not strongly expressed at the RNA level indicates either that some low-abundance proteins were identified or that some transcripts or proteins showed extended half-lives. The most abundant classes identified with the combined proteomic and transcriptomic approach involved energy production, translational activities and nucleotide transport, reflecting an active metabolism. The simultaneous large-scale analysis of transcriptomes and proteomes enables a global and holistic view of the S. aureus biology, allowing the parallel study of multiple active events in an organism.

Bacterial Proteins↗

SWISS-2DPAGE, ten years later.

The SWISS-2DPAGE database was established in 1993 and is maintained collaboratively by the Swiss Institute of Bioinformatics (SIB) and the Biomedical Proteomics Research Group (BPRG) of the Geneva University Hospital. During these years, SWISS-2DPAGE underwent constant modification and improvement. Current content includes about 4000 identified spots corresponding to 1200 different protein entries in 36 reference maps from human, mouse, Arabidopsis thaliana, Dictyostelium discoideum, Escherichia coli, Saccharomyces cerevisiae and Staphylococcus aureus origins. With a high level of annotation and integration with other relevant databases, SWISS-2DPAGE is a reference source in the proteomics world. Queries to SWISS-2DPAGE database currently reach 1000 hits per day.

Animals↗

The molecular scanner: concept and developments.

Approaches aimed at deciphering the proteome have illustrated the need for relatively complex and highly sensitive methodologies. The major elements of proteome analysis, such as powerful protein separation and enzymatic processing, mass spectrometry and dedicated bioinformatics have been assembled in the development of the molecular scanner. This highly flexible and data-rich approach has combined the power of electrophoretic protein separation, the simultaneous digestion and transfer of proteins through an enzymatic membrane, the immediate use of the MALDI mass spectrometer to scan a collecting membrane, and the development of dedicated bioinformatics tools to perform protein identification and molecular imaging of the proteome. Clinical applications of the molecular scanner have also started to be developed for disease diagnosis in biological material.

Animals↗

ExPASy: The proteomics server for in-depth protein knowledge and analysis.

The ExPASy (the Expert Protein Analysis System) World Wide Web server (http://www.expasy.org), is provided as a service to the life science community by a multidisciplinary team at the Swiss Institute of Bioinformatics (SIB). It provides access to a variety of databases and analytical tools dedicated to proteins and proteomics. ExPASy databases include SWISS-PROT and TrEMBL, SWISS-2DPAGE, PROSITE, ENZYME and the SWISS-MODEL repository. Analysis tools are available for specific tasks relevant to proteomics, similarity searches, pattern and profile searches, post-translational modification prediction, topology prediction, primary, secondary and tertiary structure analysis and sequence alignment. These databases and tools are tightly interlinked: a special emphasis is placed on integration of database entries with related resources developed at the SIB and elsewhere, and the proteomics tools have been designed to read the annotations in SWISS-PROT in order to enhance their predictions. ExPASy started to operate in 1993, as the first WWW server in the field of life sciences. In addition to the main site in Switzerland, seven mirror sites in different continents currently serve the user community.

Databases, Protein↗

Popitam: towards new heuristic strategies to improve protein identification from tandem mass spectrometry data.

In recent years, proteomics research has gained importance due to increasingly powerful techniques in protein purification, mass spectrometry and identification, and due to the development of extensive protein and DNA databases from various organisms. Nevertheless, current identification methods from spectrometric data have difficulties in handling modifications or mutations in the source peptide. Moreover, they have low performance when run on large databases (such as genomic databases), or with low quality data, for example due to bad calibration or low fragmentation of the source peptide. We present a new algorithm dedicated to automated protein identification from tandem mass spectrometry (MS/MS) data by searching a peptide sequence database. Our identification approach shows promising properties for solving the specific difficulties enumerated above. It consists of matching theoretical peptide sequences issued from a database with a structured representation of the source MS/MS spectrum. The representation is similar to the spectrum graphs commonly used by de novo sequencing software. The identification process involves the parsing of the graph in order to emphasize relevant sections for each theoretical sequence, and leads to a list of peptides ranked by a correlation score. The parsing of the graph, which can be a highly combinatorial task, is performed by a bio-inspired algorithm called Ant Colony Optimization algorithm.

Algorithms↗

The Make 2D-DB II package: conversion of federated two-dimensional gel electrophoresis databases into a relational format and interconnection of distributed databases.

The Make 2D-DB tool has been previously developed to help build federated two-dimensional gel electrophoresis (2-DE) databases on one's own web site. The purpose of our work is to extend the strength of the first package and to build a more efficient environment. Such an environment should be able to fulfill the different needs and requirements arising from both the growing use of 2-DE techniques and the increasing amount of distributed experimental data.

Databases, Protein↗

Mass spectrometry-based proteomics: current status and potential use in clinical chemistry.

For some years now, scientists have been spending a lot of effort in developing methods to analyse and compare complex protein samples. One of the goals of such global analyses of what is known as proteomes is to discover specific protein markers--or fingerprints of protein markers--from various types of affected biological samples. Considering the battery of technologies currently available, mass spectrometry (MS) constitutes an essential tool in proteomics. We describe here the type of MS instrumentation that is currently dedicated to proteomics research. We also describe the major experimental workflows that are typically used in proteomics today, with a focus on those incorporating MS as a major analysis tool.

Chemistry, Clinical↗

Proteomics and its trends facing nature's complexity.

The complexity of nature is tremendous, particularly at the epigenetic level. Proteomic studies must therefore complement genomic discoveries to better understand biological processes. Because of the very large number of modified proteins and their great variability in physico-chemical properties, no single method can be used to analyze all of them. Mass spectrometry has demonstrated its superior ability to rapidly identify and partially characterize numerous proteins in low abundance and has become a central element in most proteomic projects. Studies of protein function are necessary to understand biological pathways and this is being tackled using several approaches such as two hybrid systems, phage technology or affinity methods. Finally, mathematical and bioinformatic developments will be essential to study nature's complex systems.

Computational Biology↗

Peptide mass fingerprinting peak intensity prediction: extracting knowledge from spectra.

Matrix-assisted laser desorption/ionization-time of flight mass spectrometry has become a valuable tool in proteomics. With the increasing acquisition rate of mass spectrometers, one of the major issues is the development of accurate, efficient and automatic peptide mass fingerprinting (PMF) identification tools. Current tools are mostly based on counting the number of experimental peptide masses matching with theoretical masses. Almost all of them use additional criteria such as isoelectric point, molecular weight, PTMs, taxonomy or enzymatic cleavage rules to enhance prediction performance. However, these identification tools seldom use peak intensities as parameter as there is currently no model predicting the intensities based on the physicochemical properties of peptides. In this work, we used standard datamining methods such as classification and regression methods to find correlations between peak intensities and the properties of the peptides composing a PMF spectrum. These methods were applied on a dataset comprising a series of PMF experiments involving 157 proteins. We found that the C4.5 method gave the more informative results for the classification task (prediction of the presence or absence of a peptide in a spectra) and M5' for the regression methods (prediction of the normalized intensity of a peptide peak). The C4.5 result correctly classified 88% of the theoretical peaks; whereas the M5' peak intensities had a correlation coefficient of 0.6743 with the experimental peak intensities. These methods enabled us to obtain decision and model trees that can be directly used for prediction and identification of PMF results. The work performed permitted to lay the foundations of a method to analyze factors influencing the peak intensity of PMF spectra. A simple extension of this analysis could lead to improve the accuracy of the results by using a larger dataset. Additional peptide characteristics or even PMF experimental parameters can also be taken into account in the datamining process to analyze their influence on the peak intensity. Furthermore, this datamining approach can certainly be extended to the tandem mass spectrometry domain or other mass spectrometry derived methods.

Acrylamide↗

Molecular scanner experiment with human plasma: improving protein identification by using intensity distributions of matching peptide masses.

The development of high throughput utilities to identify proteins is a major challenge in present research in the field of proteomics. One such utility, the molecular scanner, uses proteins separated by two-dimensional polyacrylamide gel electrophoresis that are digested in the gel and during transfer onto a collecting membrane. After adding a matrix, the membrane is inserted into a matrix-assisted laser desorption/ionization-time of flight mass spectrometer and a peptide mass fingerprint (PMF) is measured for every scanned site. Since the spacing between scanned sites is much smaller than the size of the most abundant protein spots, there is a certain redundancy in the data that was used in an earlier experiment with Escherichia coli [1] to improve mass calibration and PMF identification results. It was observed that the signal intensity of a peptide mass as a function of the position on the membrane showed similar patterns if peptides stemmed from the same protein. Taking account of these similarities a clustering algorithm was used to find lists of experimental masses with similar intensity distributions, which provided clearer identification of the corresponding proteins. Here, these methods are applied to a human plasma scan, where proteins were highly modified and less separated. The presence of very abundant proteins like albumin and immunoglobulins added another difficulty. The calibration of the initial PMFs was not satisfactory and masses had to be recalibrated. After discarding chemical noise, the membrane was partitioned into regions and for each region protein identification was carried out separately. A new scoring method was used, where the PMF score was multiplied by a factor that measures the similarity of matching peptides. This method proved to be more robust than the method developed in [1] if the region where a protein was found had an extended, nonspherical shape and strong overlap with regions of other proteins. Many proteins annotated on the SWISS-2D PAGE human plasma master gel could be clearly identified and many interesting properties were observed.

Algorithms↗

Hydrogen/deuterium exchange for higher specificity of protein identification by peptide mass fingerprinting.

Genome sequencing projects produce large amounts of information that could be translated into potential protein sequences. Such amounts of material continuously increase protein database sizes. At present, 22 times more protein sequences are available in the SWISS-PROT and TrEMBL databases than 8 years ago in SWISS-PROT. One of the methods of choice for protein identification makes use of specific endoproteolytic cleavage followed by matrix-assisted laser desorption/ionisation mass spectrometric (MALDI-MS) analysis of the digested product. Since 1993, when this technique was first demonstrated, the conditions required for a correct identification have changed dramatically. Whilst 4-5 peptides with an uncertainty of 2-3 Da were sufficient for a correct identification in 1993, 10-13 peptides with less than 60 ppm mass error are now required for human and E. coli proteins. This evolution is directly related to the continuous increase in protein database sizes, which causes an increase in the number of false positive matches in identification results. Use of an information complement deduced from the primary protein sequence, in the process of identification by peptide mass fingerprints, can help to increase confidence in the identification results. In this article, we propose the exchange of labile hydrogen atoms with deuterium atoms to provide an alternative information complement. The exchange reaction with optimised techniques has shown an average 95% of hydrogen/deuterium (H/D) exchange on tryptic peptides. This level of exchange was sufficient to single out one or more peptides from a list of potential candidate proteins due to the dependence of H/D exchange on the peptide primary structure. This technique also has clear advantages in the identification of small proteins where direct protein identification is impaired by the limited number of endoproteolytic peptides. Then, information related to primary sequence obtained with this technique could help to identify proteins with high confidence without any expensive tandem mass spectrometry instruments.

Amino Acid Sequence↗