PubMed Health⌕ Search

Biomedical subjects

Eugene Kolker

Publications and source records attributed to Eugene Kolker.

At least 19 recordsLinked to original sources

A predictive model for identifying proteins by a single peptide match.

MOTIVATION: Tandem mass-spectrometry of trypsin digests, followed by database searching, is one of the most popular approaches in high-throughput proteomics studies. Peptides are considered identified if they pass certain scoring thresholds. To avoid false positive protein identification, > or = 2 unique peptides identified within a single protein are generally recommended. Still, in a typical high-throughput experiment, hundreds of proteins are identified only by a single peptide. We introduce here a method for distinguishing between true and false identifications among single-hit proteins. The approach is based on randomized database searching and usage of logistic regression models with cross-validation. This approach is implemented to analyze three bacterial samples enabling recovery 68-98% of the correct single-hit proteins with an error rate of < 2%. This results in a 22-65% increase in number of identified proteins. Identifying true single-hit proteins will lead to discovering many crucial regulators, biomarkers and other low abundance proteins. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

New metrics for comparative genomics.

The availability of genome sequences from a variety of organisms presents an opportunity to apply this sequence information to solving the key problems of molecular biology. One of the principal roadblocks on this path is the lack of appropriate descriptors and metrics that could succinctly represent the new knowledge stemming from the genomic data. Several new metrics have recently been used in comparative genome analysis, yet challenges remain in finding an appropriate language for the emerging discipline of systems biology.

Animals↗

Protein identification and expression analysis using mass spectrometry.

The identification and quantification of the proteins that a whole organism expresses under certain conditions is a main focus of high-throughput proteomics. Advanced proteomics approaches generate new biologically relevant data and potent hypotheses. A practical report of what proteome studies can and cannot accomplish in common laboratory settings is presented here. The review discusses the most popular tandem mass-spectrometry-based methods and focuses on how to produce reliable results. A step-by-step description of proteome experiments is given, including sample preparation, digestion, labeling, liquid chromatography, data processing, database searching and statistical analysis. The difficulties and bottlenecks of proteome analysis are addressed and the requirements for further improvements are discussed. Several diverse high-throughput proteomics-based studies of microorganisms are described.

Amino Acid Sequence↗

Experimental standards for high-throughput proteomics.

Proteome analysis, utilizing high-throughput proteomics approaches, involves studying proteins that a whole organism (or specific tissue or cellular compartment) expresses under certain conditions. Intrinsic difficulties of these studies, as well as the enormous volumes of data they typically produce, make the proteome analysis and interpretation very difficult. As with any high-throughput approach, proteomics experiments should be carefully designed, analyzed, and verified. In addition to computational standards,experimental standards--simple and complex mixtures of known proteins--for high-throughput proteomics have to be developed and utilized. This article discusses such experimental standards and their implementations.

Animals↗

Global profiling of Shewanella oneidensis MR-1: expression of hypothetical genes and improved functional annotations.

The gamma-proteobacterium Shewanella oneidensis strain MR-1 is a metabolically versatile organism that can reduce a wide range of organic compounds, metal ions, and radionuclides. Similar to most other sequenced organisms, approximately 40% of the predicted ORFs in the S. oneidensis genome were annotated as uncharacterized "hypothetical" genes. We implemented an integrative approach by using experimental and computational analyses to provide more detailed insight into gene function. Global expression profiles were determined for cells after UV irradiation and under aerobic and suboxic growth conditions. Transcriptomic and proteomic analyses confidently identified 538 hypothetical genes as expressed in S. oneidensis cells both as mRNAs and proteins (33% of all predicted hypothetical proteins). Publicly available analysis tools and databases and the expression data were applied to improve the annotation of these genes. The annotation results were scored by using a seven-category schema that ranked both confidence and precision of the functional assignment. We were able to identify homologs for nearly all of these hypothetical proteins (97%), but could confidently assign exact biochemical functions for only 16 proteins (category 1; 3%). Altogether, computational and experimental evidence provided functional assignments or insights for 240 more genes (categories 2-5; 45%). These functional annotations advance our understanding of genes involved in vital cellular processes, including energy conversion, ion transport, secondary metabolism, and signal transduction. We propose that this integrative approach offers a valuable means to undertake the enormous challenge of characterizing the rapidly growing number of hypothetical proteins with each newly sequenced genome.

Gene Expression Profiling↗

Charge state estimation for tandem mass spectrometry proteomics.

High-throughput protein analysis by tandem mass spectrometry produces anywhere from thousands to millions of spectra that are being used for peptide and protein identifications. Though each spectrum corresponds only to one charged peptide (ion) state, repetitive database searches of multiple charge states are typically conducted since the resolution of many common mass spectrometers is not sufficient to determine the charge state. The resulting database searches are both error-prone and time-consuming. We describe a straightforward, accurate approach on charge state estimation (CHASTE). CHASTE relies on fragment ion peak distributions, and by using reliable logistic regression models, combines different measurements to improve its accuracy. CHASTE's performance has been validated on data sets, comprised of known peptide dissociation spectra, obtained by replicate analyses of our earlier developed protein standard mixture using ion trap mass spectrometers at different laboratories. CHASTE was able to reduce number of needed database searches by at least 60% and the number of redundant searches by at least 90% virtually without any informational loss. This greatly alleviates one of the major bottlenecks in high throughput peptide and protein identifications. Thresholds and parameter estimates can be tailored to specific analysis situations, pipelines, and instrumentations. CHASTE was implemented in Java GUI-based and command-line-based interfaces.

Computer Graphics↗

Randomized sequence databases for tandem mass spectrometry peptide and protein identification.

Tandem mass spectrometry (MS/MS) combined with database searching is currently the most widely used method for high-throughput peptide and protein identification. Many different algorithms, scoring criteria, and statistical models have been used to identify peptides and proteins in complex biological samples, and many studies, including our own, describe the accuracy of these identifications, using at best generic terms such as "high confidence." False positive identification rates for these criteria can vary substantially with changing organisms under study, growth conditions, sequence databases, experimental protocols, and instrumentation; therefore, study-specific methods are needed to estimate the accuracy (false positive rates) of these peptide and protein identifications. We present and evaluate methods for estimating false positive identification rates based on searches of randomized databases (reversed and reshuffled). We examine the use of separate searches of a forward then a randomized database and combined searches of a randomized database appended to a forward sequence database. Estimated error rates from randomized database searches are first compared against actual error rates from MS/MS runs of known protein standards. These methods are then applied to biological samples of the model microorganism Shewanella oneidensis strain MR-1. Based on the results obtained in this study, we recommend the use of use of combined searches of a reshuffled database appended to a forward sequence database as a means providing quantitative estimates of false positive identification rates of peptides and proteins. This will allow researchers to set criteria and thresholds to achieve a desired error rate and provide the scientific community with direct and quantifiable measures of peptide and protein identification accuracy as opposed to vague assessments such as "high confidence."

Databases, Protein↗

Identification and functional analysis of 'hypothetical' genes expressed in Haemophilus influenzae.

The progress in genome sequencing has led to a rapid accumulation in GenBank submissions of uncharacterized 'hypothetical' genes. These genes, which have not been experimentally characterized and whose functions cannot be deduced from simple sequence comparisons alone, now comprise a significant fraction of the public databases. Expression analyses of Haemophilus influenzae cells using a combination of transcriptomic and proteomic approaches resulted in confident identification of 54 'hypothetical' genes that were expressed in cells under normal growth conditions. In an attempt to understand the functions of these proteins, we used a variety of publicly available analysis tools. Close homologs in other species were detected for each of the 54 'hypothetical' genes. For 16 of them, exact functional assignments could be found in one or more public databases. Additionally, we were able to suggest general functional characterization for 27 more genes (comprising approximately 80% total). Findings from this analysis include the identification of a pyruvate-formate lyase-like operon, likely to be expressed not only in H.influenzae but also in several other bacteria. Further, we also observed three genes that are likely to participate in the transport and/or metabolism of sialic acid, an important component of the H.influenzae lipo-oligosaccharide. Accurate functional annotation of uncharacterized genes calls for an integrative approach, combining expression studies with extensive computational analysis and curation, followed by eventual experimental verification of the computational predictions.

Amino Acid Sequence↗

Statistical analysis of global gene expression data: some practical considerations.

Applying appropriate error models and conservative estimates to microarray data helps to reduce the number of false predictions and allows one to focus on biologically relevant observations. Several key conclusions have been drawn from the statistical analysis of global gene expression data: it is worth keeping core information for each experiment, including raw and processed data; biological and technical replicates are needed; careful experimental design makes the analysis simpler and more powerful; the choice of the similarity measure is nontrivial and depends on the goal of an experiment; array information must be complemented with other data; and gene expression studies are 'hypothesis generators'.

Algorithms↗

In Silico Metabolic Model and Protein Expression of Haemophilus influenzae Strain Rd KW20 in Rich Medium.

The intermediary metabolism of Haemophilus influenzae strain Rd KW20 was studied by a combination of protein expression analysis using a recently developed direct proteomics approach, mutational analysis, and mathematical modeling. Special emphasis was placed on carbon utilization, sugar fermentation, TCA cycle, and electron transport of H. influenzae cells grown microaerobically and anaerobically in a rich medium. The data indicate that several H. influenzae metabolic proteins similar to Escherichia coli proteins, known to be regulated by low concentrations of oxygen, were well expressed in both growth conditions in H. influenzae. An in silico model of the H. influenzae metabolic network was used to study the effects of selective deletion of certain enzymatic steps. This allowed us to define proteins predicted to be essential or non-essential for cell growth and to address numerous unresolved questions about intermediary metabolism of H. influenzae. Comparison of data from in vivo protein expression with the protein list associated with a genome-scale metabolic model showed significant coverage of the known metabolic proteome. This study demonstrates the significance of an integrated approach to the characterization of H. influenzae metabolism.

Biochemistry↗

Standard mixtures for proteome studies.

Mixtures of moderate complexity were formed from 23 peptides and 12 proteins digested with trypsin, all individually characterized. These mixtures were analyzed with replicates in full and windowed m/z ranges using online high-performance reverse phase liquid chromatography coupled via electrospray ionization to an ion trap mass spectrometer. The resulting spectra were searched using SEQUEST against databases of different sizes and contents and confidences of the observed identifications were evaluated by our earlier statistical model. These data were then combined with biologically derived spectral data, searched, and further evaluated. All peptides but one and all proteins were identified with high confidence. Additionally, the presence and behavior of quadruply charged peptides was analyzed. The properties of the proposed peptide and protein mixtures as well as the performance of the statistical model were carefully investigated. These mixtures mimic the complexity seen in large-scale proteomics experiments, and are proposed to serve as quality assessment standards for future proteome studies.

Animals↗

Spectral quality assessment for high-throughput tandem mass spectrometry proteomics.

Current techniques in tandem mass spectrometric analyses of cellular protein contents often produce thousands to tens of thousands of spectra per experiment. This study introduces a new algorithm, named SPEQUAL, which is aimed at automated tandem mass spectral quality assessment. The quality of a given spectrum can be evaluated from three basic components: (i) charge state differentiation, (ii) total signal intensity, and (iii) signal-to-noise estimates. The differentiation between single and multiple precursor charge states (i) provides a binary score for a given spectrum. Components (ii) and (iii) provide partial scores which are subsequently summarized and multiplied by the first score. SPEQUAL was applied to over 10,000 data files derived from almost 3,000 tandem mass spectra, and the results (final cumulative scores) were manually verified. SPEQUAL's performance was determined to have high sensitivity and specificity and low error rates for both spectral quality estimates in general and precursor charge state differentiation in particular. Each of the partial scores is controlled by adjustable thresholds to fine-tune SPEQUAL's performance for different analysis pipelines and instrumentation. This spectral quality assessment tool is intended to act in an advisory role to the researcher, assisting in filtration of thousands of spectra typically produced by high throughput tandem mass spectrometric proteome analyses. Lastly, SPEQUAL was implemented as Java GUI-based and command-line-based interfaces freely available for both academic and industrial researchers.

Mass Spectrometry↗

LIP index for peptide classification using MS/MS and SEQUEST search via logistic regression.

This study addresses the issue of peptide identification resulting from tandem mass spectrometry proteomics analysis followed by database search. This work shows that the Logistic Identification of Peptides (LIP) Index achieves high sensitivity and specificity for peptide classification relative to a manually verified "gold" standard and also accurately estimates the probability of a correct peptide match. The LIP Index is a weighted average of SEQUEST output variables based on logistic regression models and is a transparent, easy to use, inclusive, extendable, and statistically sound approach to classify correct peptide identifications. Modifications, such as normalizing cross-correlations (Xcorr) for peptide length, adjusting for charge state, and the number of tryptic termini, significantly improve the fit the logistic regression models, as well as increase sensitivity and specificity. The LIP Index also incorporates earlier developed statistical models on spectral quality assessment and peptide identification, which further improves sensitivity and specificity.

Algorithms↗

A statistical model for identifying proteins by tandem mass spectrometry.

A statistical model is presented for computing probabilities that proteins are present in a sample on the basis of peptides assigned to tandem mass (MS/MS) spectra acquired from a proteolytic digest of the sample. Peptides that correspond to more than a single protein in the sequence database are apportioned among all corresponding proteins, and a minimal protein list sufficient to account for the observed peptide assignments is derived using the expectation-maximization algorithm. Using peptide assignments to spectra generated from a sample of 18 purified proteins, as well as complex H. influenzae and Halobacterium samples, the model is shown to produce probabilities that are accurate and have high power to discriminate correct from incorrect protein identifications. This method allows filtering of large-scale proteomics data sets with predictable sensitivity and false positive identification error rates. Fast, consistent, and transparent, it provides a standard for publishing large-scale protein identification data sets in the literature and for comparing the results obtained from different experiments.

Amino Acid Sequence↗

Initial proteome analysis of model microorganism Haemophilus influenzae strain Rd KW20.

The proteome of Haemophilus influenzae strain Rd KW20 was analyzed by liquid chromatography (LC) coupled with ion trap tandem mass spectrometry (MS/MS). This approach does not require a gel electrophoresis step and provides a rapidly developed snapshot of the proteome. In order to gain insight into the central metabolism of H. influenzae, cells were grown microaerobically and anaerobically in a rich medium and soluble and membrane proteins of strain Rd KW20 were proteolyzed with trypsin and directly examined by LC-MS/MS. Several different experimental and computational approaches were utilized to optimize the proteome coverage and to ensure statistically valid protein identification. Approximately 25% of all predicted proteins (open reading frames) of H. influenzae strain Rd KW20 were identified with high confidence, as their component peptides were unambiguously assigned to tandem mass spectra. Approximately 80% of the predicted ribosomal proteins were identified with high confidence, compared to the 33% of the predicted ribosomal proteins detected by previous two-dimensional gel electrophoresis studies. The results obtained in this study are generally consistent with those obtained from computational genome analysis, two-dimensional gel electrophoresis, and whole-genome transposon mutagenesis studies. At least 15 genes originally annotated as conserved hypothetical were found to encode expressed proteins. Two more proteins, previously annotated as predicted coding regions, were detected with high confidence; these proteins also have close homologs in related bacteria. The direct proteomics approach to studying protein expression in vivo reported here is a powerful method that is applicable to proteome analysis of any (micro)organism.

Aerobiosis↗

Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.

We present a statistical model to estimate the accuracy of peptide assignments to tandem mass (MS/MS) spectra made by database search applications such as SEQUEST. Employing the expectation maximization algorithm, the analysis learns to distinguish correct from incorrect database search results, computing probabilities that peptide assignments to spectra are correct based upon database search scores and the number of tryptic termini of peptides. Using SEQUEST search results for spectra generated from a sample of known protein components, we demonstrate that the computed probabilities are accurate and have high power to discriminate between correctly and incorrectly assigned peptides. This analysis makes it possible to filter large volumes of MS/MS database search results with predictable false identification error rates and can serve as a common standard by which the results of different research groups are compared.

Algorithms↗

Transcriptome analysis of Escherichia coli using high-density oligonucleotide probe arrays.

Microarrays traditionally have been used to analyze the expression behavior of large numbers of coding transcripts. Here we present a comprehensive approach for high-throughput transcript discovery in Escherichia coli focused mainly on intergenic regions which, together with analysis of coding transcripts, provides us with a more complete insight into the organism's transcriptome. Using a whole genome array, we detected expression for 4052 coding transcripts and identified 1102 additional transcripts in the intergenic regions of the E.coli genome. Further classification reveals 317 novel transcripts with unknown function. Our results show that, despite sophisticated approaches to genome annotation, many cellular transcripts remain unidentified. Through the experimental identification of all RNAs expressed under a specific condition, we gain a more thorough understanding of all cellular processes.

3' Untranslated Regions↗

H. influenzae Consortium: integrative study of H. influenzae-human interactions.

Developments in high-throughput analysis tools coupled with integrative computational techniques have enabled biological studies to reach new levels. The ability to correlate large volumes of diverse data types into cohesive models of organism function has spawned a new systematic approach to biological investigation. The creation of a new consortium has been proposed to investigate a single organism utilizing these comprehensive approaches. The Haemophilus influenzae Consortium (HIC) would be comprised of five laboratories, each providing separate and complementary areas of expertise in the study of Haemophilus influenzae (HI). The 5-year study proposes to develop coherent models of HI, both as a stand-alone organism, and more importantly, as a human pathogen. Studies in growth condition specificity followed by genomic, metabolic, and proteomic experimentation will be combined and integrated through computational and experimental analyses to form dynamic and predictive models of HI and its responses. Data from the HIC will allow greater understanding of cellular behavior, pathogen-host interactions, bacterial infection, and provide future scientific endeavors with a template for studies of other pathogens.

Bacterial Adhesion↗