PubMed Health⌕ Search

Biomedical subjects

Jimmy Eng

Publications and source records attributed to Jimmy Eng.

At least 19 recordsLinked to original sources

A combined dataset of human cerebrospinal fluid proteins identified by multi-dimensional chromatography and tandem mass spectrometry.

Human cerebrospinal fluid (CSF) is an important source for studying protein biomarkers of age-related neurodegenerative diseases. Before characterizing biomarkers unique to each disease, it is necessary to categorize CSF proteins systematically and extensively. However, the enormous complexity, great dynamic range of protein concentrations, and tremendous protein heterogeneity due to post-translational modification of CSF create significant challenges to the existing proteomics technologies for an in-depth, nonbiased profiling of the human CSF proteome. To circumvent these difficulties, in the last few years, we have utilized several different separation methodologies and mass spectrometric platforms that greatly enhanced the identification coverage and the depth of protein profiling of CSF to characterize CSF proteome. In total, 2594 proteins were identified in well-characterized pooled human CSF samples using stringent proteomics criteria. This report summarizes our efforts to comprehensively characterize the human CSF proteome to date.

Cerebrospinal Fluid Proteins↗

UniPep--a database for human N-linked glycosites: a resource for biomarker discovery.

There has been considerable recent interest in proteomic analyses of plasma for the purpose of discovering biomarkers. Profiling N-linked glycopeptides is a particularly promising method because the population of N-linked glycosites represents the proteomes of plasma, the cell surface, and secreted proteins at very low redundancy and provides a compelling link between the tissue and plasma proteomes. Here, we describe UniPep http://www.unipep.org--a database of human N-linked glycosites--as a resource for biomarker discovery.

Computational Biology↗

A suite of algorithms for the comprehensive analysis of complex protein mixtures using high-resolution LC-MS.

MOTIVATION: Comparing two or more complex protein mixtures using liquid chromatography mass spectrometry (LC-MS) requires multiple analysis steps to locate and quantitate natural peptides within a single experiment and to align and normalize findings across multiple experiments. RESULTS: We describe msInspect, an open-source application comprising algorithms and visualization tools for the analysis of multiple LC-MS experimental measurements. The platform integrates novel algorithms for detecting signatures of natural peptides within a single LC-MS measurement and combines multiple experimental measurements into a peptide array, which may then be mined using analysis tools traditionally applied to genomic array analysis. The platform supports quantitation by both label-free and isotopic labeling approaches. The software implementation has been designed so that many key components may be easily replaced, making it useful as a workbench for integrating other novel algorithms developed by a growing research community. AVAILABILITY: The msInspect software is distributed freely under an Apache 2.0 license. The software as well as a Zip file with all peptide feature files and scripts needed to generate the tables and figures in this article are available at http://proteomics.fhcrc.org/.

Algorithms↗

The PeptideAtlas project.

The completion of the sequencing of the human genome and the concurrent, rapid development of high-throughput proteomic methods have resulted in an increasing need for automated approaches to archive proteomic data in a repository that enables the exchange of data among researchers and also accurate integration with genomic data. PeptideAtlas (http://www.peptideatlas.org/) addresses these needs by identifying peptides by tandem mass spectrometry (MS/MS), statistically validating those identifications and then mapping identified sequences to the genomes of eukaryotic organisms. A meaningful comparison of data across different experiments generated by different groups using different types of instruments is enabled by the implementation of a uniform analytic process. This uniform statistical validation ensures a consistent and high-quality set of peptide and protein identifications. The raw data from many diverse proteomic experiments are made available in the associated PeptideAtlas repository in several formats. Here we present a summary of our process and details about the Human, Drosophila and Yeast PeptideAtlas builds.

Animals↗

Computational Proteomics Analysis System (CPAS): an extensible, open-source analytic system for evaluating and publishing proteomic data and high throughput biological experiments.

The open-source Computational Proteomics Analysis System (CPAS) contains an entire data analysis and management pipeline for Liquid Chromatography Tandem Mass Spectrometry (LC-MS/MS) proteomics, including experiment annotation, protein database searching and sequence management, and mining LC-MS/MS peptide and protein identifications. CPAS architecture and features, such as a general experiment annotation component, installation software, and data security management, make it useful for collaborative projects across geographical locations and for proteomics laboratories without substantial computational support.

Computational Biology↗

Challenges in deriving high-confidence protein identifications from data gathered by a HUPO plasma proteome collaborative study.

The Human Proteome Organization (HUPO) recently completed the first large-scale collaborative study to characterize the human serum and plasma proteomes. The study was carried out in different locations and used diverse methods and instruments to compare and integrate tandem mass spectrometry (MS/MS) data on aliquots of pooled serum and plasma from healthy subjects. Liquid chromatography (LC)-MS/MS data sets from 18 laboratories were matched to the International Protein Index database, and an initial integration exercise resulted in 9,504 proteins identified with one or more peptides, and 3,020 proteins identified with two or more peptides. This article uses a rigorous statistical approach to take into account the length of coding regions in genes, and multiple hypothesis-testing techniques. On this basis, we now present a reduced set of 889 proteins identified with a confidence level of at least 95%. We also discuss the importance of such an integrated analysis in providing an accurate representation of a proteome as well as the value such data sets contain for the high-confidence identification of protein matches to novel exons, some of which may be localized in alternatively spliced forms of known plasma proteins and some in previously nonannotated gene sequences.

Blood Proteins↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

A uniform proteomics MS/MS analysis platform utilizing open XML file formats.

The analysis of tandem mass (MS/MS) data to identify and quantify proteins is hampered by the heterogeneity of file formats at the raw spectral data, peptide identification, and protein identification levels. Different mass spectrometers output their raw spectral data in a variety of proprietary formats, and alternative methods that assign peptides to MS/MS spectra and infer protein identifications from those peptide assignments each write their results in different formats. Here we describe an MS/MS analysis platform, the Trans-Proteomic Pipeline, which makes use of open XML file formats for storage of data at the raw spectral data, peptide, and protein levels. This platform enables uniform analysis and exchange of MS/MS data generated from a variety of different instruments, and assigned peptides using a variety of different database search programs. We demonstrate this by applying the pipeline to data sets generated by ThermoFinnigan LCQ, ABI 4700 MALDI-TOF/TOF, and Waters Q-TOF instruments, and searched in turn using SEQUEST, Mascot, and COMET.

Archaeal Proteins↗

High throughput proteome screening for biomarker detection.

Mass spectrometry-based quantitative proteomics has become an important component of biological and clinical research. Current methods, while highly developed and powerful, are falling short of their goal of routinely analyzing whole proteomes mainly because the wealth of proteomic information accumulated from prior studies is not used for the planning or interpretation of present experiments. The consequence of this situation is that in every proteomic experiment the proteome is rediscovered. In this report we describe an approach for quantitative proteomics that builds on the extensive prior knowledge of proteomes and a platform for the implementation of the method. The method is based on the selection and chemical synthesis of isotopically labeled reference peptides that uniquely identify a particular protein and the addition of a panel of such peptides to the sample mixture consisting of tryptic peptides from the proteome in question. The platform consists of a peptide separation module for the generation of ordered peptide arrays from the combined peptide sample on the sample plate of a MALDI mass spectrometer, a high throughput MALDI-TOF/TOF mass spectrometer, and a suite of software tools for the selective analysis of the targeted peptides and the interpretation of the results. Applying the method to the analysis of the human blood serum proteome we demonstrate the feasibility of using mass spectrometry-based proteomics as a high throughput screening technology for the detection and quantification of targeted proteins in a complex system.

Automation↗

Proteomic analysis of synaptosomes using isotope-coded affinity tags and mass spectrometry.

Synaptosomes are isolated synapses produced by subcellular fractionation of brain tissue. They contain the complete presynaptic terminal, including mitochondria and synaptic vesicles, and portions of the postsynaptic side, including the postsynaptic membrane and the postsynaptic density (PSyD). A proteomic characterisation of synaptosomes isolated from mouse brain was performed employing the isotope-coded affinity tag (ICAT) method and tandem mass spectrometry (MS/MS). After isotopic labelling and tryptic digestion, peptides were fractionated by cation exchange chromatography and cysteine-containing peptides were isolated by affinity chromatography. The peptides were identified by microcapillary liquid chromatography-electrospray ionisation MS/MS (muLC-ESI MS/MS). In two experiments, peptides representing a total of 1131 database entries were identified. They are involved in different presynaptic and postsynaptic functions, including synaptic vesicle exocytosis for neurotransmitter release, vesicle endocytosis for synaptic vesicle recycling, as well as postsynaptic receptors and proteins constituting the PSyD. Moreover, a large number of soluble and membrane-bound molecules serving functions in synaptic signal transduction and metabolism were detected. The results provide an inventory of the synaptic proteome and confirm the suitability of the ICAT method for the assessment of synaptic structure, function and plasticity.

Animals↗

Overview of the HUPO Plasma Proteome Project: results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database.

HUPO initiated the Plasma Proteome Project (PPP) in 2002. Its pilot phase has (1) evaluated advantages and limitations of many depletion, fractionation, and MS technology platforms; (2) compared PPP reference specimens of human serum and EDTA, heparin, and citrate-anti-coagulated plasma; and (3) created a publicly-available knowledge base (www.bioinformatics.med.umich.edu/hupo/ppp; www.ebi.ac.uk/pride). Thirty-five participating laboratories in 13 countries submitted datasets. Working groups addressed (a) specimen stability and protein concentrations; (b) protein identifications from 18 MS/MS datasets; (c) independent analyses from raw MS-MS spectra; (d) search engine performance, subproteome analyses, and biological insights; (e) antibody arrays; and (f) direct MS/SELDI analyses. MS-MS datasets had 15 710 different International Protein Index (IPI) protein IDs; our integration algorithm applied to multiple matches of peptide sequences yielded 9504 IPI proteins identified with one or more peptides and 3020 proteins identified with two or more peptides (the Core Dataset). These proteins have been characterized with Gene Ontology, InterPro, Novartis Atlas, OMIM, and immunoassay-based concentration determinations. The database permits examination of many other subsets, such as 1274 proteins identified with three or more peptides. Reverse protein to DNA matching identified proteins for 118 previously unidentified ORFs. We recommend use of plasma instead of serum, with EDTA (or citrate) for anticoagulation. To improve resolution, sensitivity and reproducibility of peptide identifications and protein matches, we recommend combinations of depletion, fractionation, and MS/MS technologies, with explicit criteria for evaluation of spectra, use of search algorithms, and integration of homologous protein matches. This Special Issue of PROTEOMICS presents papers integral to the collaborative analysis plus many reports of supplementary work on various aspects of the PPP workplan. These PPP results on complexity, dynamic range, incomplete sampling, false-positive matches, and integration of diverse datasets for plasma and serum proteins lay a foundation for development and validation of circulating protein biomarkers in health and disease.

Algorithms↗

Quantitative proteomic analysis of age-related changes in human cerebrospinal fluid.

Identification of cerebrospinal fluid (CSF) biomarkers of the common age-related neurodegenerative diseases would be of great value to clinicians because of the difficulties in differential diagnoses of these diseases in clinical practice. Proteins are one class of potential biomarkers currently under investigation in the hope that different ensembles of proteins will aid in the diagnosis of these diseases, as well as in the assessment of progression and response to therapy. However, before undertaking a rational approach to CSF protein biomarkers of age-related neurodegeneration, we must first systematically identify CSF proteins and determine whether their levels change with normal aging. In this study, we used a powerful shotgun proteomic method, two-dimensional microcapillary liquid chromatography electrospray ionization tandem mass spectrometry, to identify proteins in human CSF. Additionally, using pooled CSF samples, we quantitatively compared the CSF proteome of younger adults with that of older adults using isotope-coded affinity tags (ICAT). From these studies we identified more than 300 proteins in CSF and found that there were 30 proteins with >20% change in concentrations between older and younger individuals. Finally, we validated changes in concentration for two of these proteins using Western blots in CSF from a separate set of individuals. These data not only expand substantially our current knowledge regarding human CSF proteins, but also supply the necessary information to appropriately interpret protein biomarkers of age-related neurodegenerative diseases.

Adult↗

Pancreatic cancer proteome: the proteins that underlie invasion, metastasis, and immunologic escape.

BACKGROUND & AIMS: Pancreatic cancer is a highly lethal disease that has seen little headway in diagnosis and treatment for the past few decades. The effective treatment of pancreatic cancer is critically relying on the diagnosis of the disease at an early stage, which still remains challenging. New experimental approaches, such as quantitative proteomics, have shown great potential for the study of cancer and have opened new opportunities to investigate crucial events underlying pancreatic tumorigenesis and to exploit this knowledge for early detection and better intervention. METHODS: To systematically study protein expression in pancreatic cancer, we used isotope-coded affinity tag technology and tandem mass spectrometry to perform quantitative proteomic profiling of pancreatic cancer tissues and normal pancreas. RESULTS: A total of 656 proteins were identified and quantified in 2 pancreatic cancer samples, of which 151 were differentially expressed in cancer by at least 2-fold. This study revealed numerous proteins that are newly discovered to be associated with pancreatic cancer, providing candidates for future early diagnosis biomarkers and targets for therapy. Several differentially expressed proteins were further validated by tissue microarray immunohistochemistry. Many of the differentially expressed proteins identified are involved in protein-driven interactions between the ductal epithelium and the extracellular matrix that orchestrate tumor growth, migration, angiogenesis, invasion, metastasis, and immunologic escape. CONCLUSIONS: Our study is the first application of isotope-coded affinity tag technology for proteomic analysis of human cancer tissue and has shown the value of this technology in identifying differentially expressed proteins in cancer.

Humans↗

The Pseudomonas aeruginosa proteome during anaerobic growth.

Isotope-coded affinity tag analysis and two-dimensional gel electrophoresis followed by tandem mass spectrometry were used to identify Pseudomonas aeruginosa proteins expressed during anaerobic growth. Out of the 617 proteins identified, 158 were changed in abundance during anaerobic growth compared to during aerobic growth, including proteins whose increased expression was expected based on their role in anaerobic metabolism. These results form the basis for future analyses of alterations in bacterial protein content during growth in various environments, including the cystic fibrosis airway.

Anaerobiosis↗

Quantitative proteomics of cerebrospinal fluid from patients with Alzheimer disease.

Biomarkers to assist in the diagnosis and medical management of Alzheimer disease (AD) are a pressing need. We have employed a proteomic approach, microcapillary liquid chromatography mass spectrometry of proteins labeled with isotope-coded affinity tags (ICAT), to quantify relative changes in the proteome of human cerebrospinal fluid (CSF) obtained from the lumbar cistern. Using CSF from well-characterized AD patients and age-matched controls at 2 different institutions, we quantified protein concentration ratios of 42% of the 390 CSF proteins that we have identified and found differences > or = 20% in over half of them. We confirmed our findings by western blot and validated this approach by quantifying relative levels of amyloid precursor protein and cathepsin B in 17 AD patients and 16 control individuals. Quantitative proteomics of CSF from AD patients compared to age-matched controls, as well as from other neurodegenerative diseases, will allow us to generate a roster of proteins that may serve as specific biomarker panels for AD and other geriatric dementias.

Adult↗

System-based proteomic analysis of the interferon response in human liver cells.

BACKGROUND: Interferons (IFNs) play a critical role in the host antiviral defense and are an essential component of current therapies against hepatitis C virus (HCV), a major cause of liver disease worldwide. To examine liver-specific responses to IFN and begin to elucidate the mechanisms of IFN inhibition of virus replication, we performed a global quantitative proteomic analysis in a human hepatoma cell line (Huh7) in the presence and absence of IFN treatment using the isotope-coded affinity tag (ICAT) method and tandem mass spectrometry (MS/MS). RESULTS: In three subcellular fractions from the Huh7 cells treated with IFN (400 IU/ml, 16 h) or mock-treated, we identified more than 1,364 proteins at a threshold that corresponds to less than 5% false-positive error rate. Among these, 54 were induced by IFN and 24 were repressed by more than two-fold, respectively. These IFN-regulated proteins represented multiple cellular functions including antiviral defense, immune response, cell metabolism, signal transduction, cell growth and cellular organization. To analyze this proteomics dataset, we utilized several systems-biology data-mining tools, including Gene Ontology via the GoMiner program and the Cytoscape bioinformatics platform. CONCLUSIONS: Integration of the quantitative proteomics with global protein interaction data using the Cytoscape platform led to the identification of several novel and liver-specific key regulatory components of the IFN response, which may be important in regulating the interplay between HCV, interferon and the host response to virus infection.

Cell Line, Tumor↗

Integrated genomic and proteomic analyses of gene expression in Mammalian cells.

Using DNA microarrays together with quantitative proteomic techniques (ICAT reagents, two-dimensional DIGE, and MS), we evaluated the correlation of mRNA and protein levels in two hematopoietic cell lines representing distinct stages of myeloid differentiation, as well as in the livers of mice treated for different periods of time with three different peroxisome proliferative activated receptor agonists. We observe that the differential expression of mRNA (up or down) can capture at most 40% of the variation of protein expression. Although the overall pattern of protein expression is similar to that of mRNA expression, the incongruent expression between mRNAs and proteins emphasize the importance of posttranscriptional regulatory mechanisms in cellular development or perturbation that can be unveiled only through integrated analyses of both proteins and mRNAs.

Animals↗