PubMed Health⌕ Search

Biomedical subjects

David Fenyö

Publications and source records attributed to David Fenyö.

13 recordsLinked to original sources

Reproducibility of LC-MS-based protein identification.

Traditional analysis of liquid chromatography-mass spectrometry (LC-MS) data, typically performed by reviewing chromatograms and the corresponding mass spectra, is both time-consuming and difficult. Detailed data analysis is therefore often omitted in proteomics applications. When analysing multiple proteomics samples, it is usually only the final list of identified proteins that is reviewed. This may lead to unnecessarily complex or even contradictory results because the content of the list of identified proteins depends heavily on the conditions for triggering the collection of tandem mass spectra. Small changes in the signal intensity of a peptide in different LC-MS experiments can lead to the collection of a tandem mass spectrum in one experiment but not in another. Also, the quality of the tandem mass spectrometry experiments can vary, leading to successful identification in some cases but not in others. Using a novel image analysis approach, it is possible to achieve repeat analysis with a very high reproducibility by matching peptides across different LC-MS experiments using the retention time and parent mass over charge (m/z). It is also easy to confirm the final result visually. This approach has been investigated by using tryptic digests of integral membrane proteins from organelle-enriched fractions from Arabidopsis thaliana and it has been demonstrated that very highly reproducible, consistent, and reliable LC-MS data interpretation can be made.

Arabidopsis↗

SwePep, a database designed for endogenous peptides and mass spectrometry.

A new database, SwePep, specifically designed for endogenous peptides, has been constructed to significantly speed up the identification process from complex tissue samples utilizing mass spectrometry. In the identification process the experimental peptide masses are compared with the peptide masses stored in the database both with and without possible post-translational modifications. This intermediate identification step is fast and singles out peptides that are potential endogenous peptides and can later be confirmed with tandem mass spectrometry data. Successful applications of this methodology are presented. The SwePep database is a relational database developed using MySql and Java. The database contains 4180 annotated endogenous peptides from different tissues originating from 394 different species as well as 50 novel peptides from brain tissue identified in our laboratory. Information about the peptides, including mass, isoelectric point, sequence, and precursor protein, is also stored in the database. This new approach holds great potential for removing the bottleneck that occurs during the identification process in the field of peptidomics. The SwePep database is available to the public.

Animals↗

Optimizing search conditions for the mass fingerprint-based identification of proteins.

The two central problems in protein identification by searching a protein sequence collection with MS data are the optimal use of experimental information to allow for identification of low abundance proteins and the accurate assignment of the probability that a result is false. For comprehensive MS-based protein identification, it is necessary to choose an appropriate algorithm and optimal search conditions. We report a systematic study of the quality of PMF-based protein identifications under different sequence collection search conditions using the Probability algorithm, which assigns the statistical significance to each result. We employed 2244 PMFs from 2-DE-separated human blood plasma proteins, and performed identification under various search constraints: mass accuracy (0.01-0.3 Da), maximum number of missed cleavage sites (0-2), and size of the sequence collection searched (5.6 x 10(4)-1.8 x 10(5)). By counting the number of significant results (significance levels 0.05, 0.01, and 0.001) for each condition, we demonstrate the search condition impact on the successful outcome of proteome analysis experiments. A mass correction procedure utilizing mass deviations of albumin matching peptides was tested in an attempt to improve the statistical significance of identifications and iterative searching was employed for identification of multiple proteins from each PMF.

Blood Proteins↗

A proteomics approach to the study of absorption, distribution, metabolism, excretion, and toxicity.

A proteomics approach was used to identify liver proteins that displayed altered levels in mice following treatment with a candidate drug. Samples from livers of mice treated with candidate drug or untreated were prepared, quantified, labeled with CyDye DIGE Fluors, and subjected to two-dimensional electrophoresis. Following scanning and imaging of gels from three different isoelectric focusing intervals (3-10, 7-11, 6.2-7.5), automated spot handling was performed on a large number of gel spots including those found to differ more than 20% between the treated and untreated condition. Subsequently, differentially regulated proteins were subjected to a three-step approach of mass spectrometry using (a) matrix-assisted laser desorption/ionization time-of-flight mass spectrometry peptide mass fingerprinting, (b) post-source decay utilizing chemically assisted fragmentation, and (c) liquid chromatography-tandem mass spectrometry. Using this approach we have so far resolved 121 differentially regulated proteins following treatment of mice with the candidate drug and identified 110 of these using mass spectrometry. Such data can potentially give improved molecular insight into the metabolism of drugs as well as the proteins involved in potential toxicity following the treatment. The differentially regulated proteins could be used as targets for metabolic studies or as markers for toxicity.

Acrylamides↗

A modular cross-linking approach for exploring protein interactions.

A method is described for the elucidation of protein-protein interactions using novel cross-linking reagents and mass spectrometry. The method incorporates (1) a modular solid-phase synthetic strategy for generating the cross-linking reagents, (2) enrichment and digestion of cross-linked proteins using microconcentrators, (3) mass spectrometric analysis of cross-linked peptides, and (4) comprehensive computational analysis of the cross-linking data. This integrated approach has been applied to the study of cross-linking between the components of the heterodimeric protein complex negative cofactor 2.

Acrylic Resins↗

A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes.

This paper investigates the use of survival functions and expectation values to evaluate the results of protein identification experiments. These functions are standard statistical measures that can be used to reduce various protein identification scoring schemes to a common, easily interpretably representation. The relative merits of scoring systems were explored using this approach, as well as the effects of altering primary identification parameters. We would advocate the widespread use of these simple statistical measures to simplify and standardize the reporting of the confidence of protein identification results, allowing the users of different identification algorithms to compare their results in a straightforward and statistically significant manner. A method is described for measuring these distributions using information that is being discarded by most protein identification search engines, resulting in accurate survival functions that are specific to any combination of scoring algorithms, sequence databases, and mass spectra.

Mass Spectrometry↗

A model of random mass-matching and its use for automated significance testing in mass spectrometric proteome analysis.

A rapid and accurate method for testing the significance of protein identities determined by mass spectrometric analysis of protein digests and genome database searching is presented. The method is based on direct computation using a statistical model of the random matching of measured and theoretical proteolytic peptide masses. Protein identification algorithms typically rank the proteins of a genome database according to a score based on the number of matches between the masses obtained by mass spectrometry analysis and the theoretical proteolytic peptide masses of a database protein. The random matching of experimental and theoretical masses can cause false results. A result is significant only if the score characterizing the result deviates significantly from the score expected from a false result. A distribution of the score (number of matches) for random (false) results is computed directly from our model of the random matching, which allows significance testing under any experimental and database search constraints. In order to mimic protein identification data quality in large-scale proteome projects, low-to-high quality proteolytic peptide mass data were generated in silico and subsequently submitted to a database search program designed to include significance testing based on direct computation. This simulation procedure demonstrates the usefulness of direct significance testing for automatically screening for samples that must be subjected to peptide sequence analysis by e.g. tandem mass spectrometry in order to determine the protein identity.

Algorithms↗

Informatics and data management in proteomics.

Proteomics has become dominated by large amounts of experimental data and interpreted results. This experimental data cannot be effectively used without understanding the fundamental structure of its information content and representing that information in such a way that knowledge can be extracted from it. This review explores the structure of this information with regard to three fundamental issues: the extraction of relevant information from raw data, the scale of the projects involved and the statistical significance of protein identification results.

Algorithms↗

RADARS, a bioinformatics solution that automates proteome mass spectral analysis, optimises protein identification, and archives data in a relational database.

RADARS, a rapid, automated, data archiving and retrieval software system for high-throughput proteomic mass spectral data processing and storage, is described. The majority of mass spectrometer data files are compatible with RADARS, for consistent processing. The system automatically takes unprocessed data files, identifies proteins via in silico database searching, then stores the processed data and search results in a relational database suitable for customized reporting. The system is robust, used in 24/7 operation, accessible to multiple users of an intranet through a web browser, may be monitored by Virtual Private Network, and is secure. RADARS is scalable for use on one or many computers, and is suited to multiple processor systems. It can incorporate any local database in FASTA format, and can search protein and DNA databases online. A key feature is a suite of visualisation tools (many available gratis), allowing facile manipulation of spectra, by hand annotation, reanalysis, and access to all procedures. We also described the use of Sonar MS/MS, a novel, rapid search engine requiring 40 MB RAM per process for searches against a genomic or EST database translated in all six reading frames. RADARS reduces the cost of analysis by its efficient algorithms: Sonar MS/MS can identifiy proteins without accurate knowledge of the parent ion mass and without protein tags. Statistical scoring methods provide close-to-expert accuracy and brings robust data analysis to the non-expert user.

Amino Acid Sequence↗

Probity: a protein identification algorithm with accurate assignment of the statistical significance of the results.

An algorithm for protein identification based on mass spectrometric proteolytic peptide mapping and genome database searching is presented. The algorithm ranks database proteins based on direct calculation of the probability of random matching and assigns the statistical significance to each result. We investigate the performance of the algorithm by simulation and show that the algorithm responds to random data in the desired manner and that the statistical significance computed indicates the risk that a particular identification result is false.

Algorithms↗

Protein identification in complex mixtures.

This paper investigates the prospects of successful mass spectrometric protein identification based on mass data from proteolytic digests of complex protein mixtures. Sets of proteolytic peptide masses representing various numbers of digested proteins in a mixture were generated in silico. In each set, different proteins were selected from a protein sequence collection and for each protein the sequence coverage was randomly selected within a particular regime (15-30% or 30-60%). We demonstrate that the Probity algorithm, which is characterized by an optimal tolerance for random interference, employed in an iterative procedure can correctly identify >95% of proteins at a desired significance level in mixtures composed of hundreds of yeast proteins under realistic mass spectrometric experimental constraints. By using a model of the distribution of protein abundance, we demonstrate that the very high efficiency of identification of protein mixtures that can be achieved by appropriate choices of informatics procedures is hampered by limitations of the mass spectrometric dynamic range. The results stress the desire to choose carefully experimental protocols for comprehensive proteome analysis, focusing on truly critical issues such as the dynamic range, which potentially limits the possibilities of identifying low abundance proteins.

Fungal Proteins↗

The statistical significance of protein identification results as a function of the number of protein sequences searched.

The potential for obtaining a true mass spectrometric protein identification result depends on the choice of algorithm as well as on experimental factors that influence the information content in the mass spectrometric data. Current methods can never prove definitively that a result is true, but an appropriate choice of algorithm can provide a measure of the statistical risk that a result is false, i.e., the statistical significance. We recently demonstrated an algorithm, Probity, which assigns the statistical significance to each result. For any choice of algorithm, the difficulty of obtaining statistically significant results depends on the number of protein sequences in the sequence collection searched. By simulations of random protein identifications and using the Probity algorithm, we here demonstrate explicitly how the statistical significance depends on the number of sequences searched. We also provide an example on how the practitioner's choice of taxonomic constraints influences the statistical significance.

Algorithms↗