PubMed Health⌕ Search

Biomedical subjects

Daniel Crowther

Publications and source records attributed to Daniel Crowther.

5 recordsLinked to original sources

GAPP: a fully automated software for the confident identification of human peptides from tandem mass spectra.

This paper introduces the genome annotating proteomic pipeline (GAPP), a totally automated publicly available software pipeline for the identification of peptides and proteins from human proteomic tandem mass spectrometry data. The pipeline takes as its input a series of MS/MS peak lists from a given experimental sample and produces a series of database entries corresponding to the peptides observed within the sample, along with related confidence scores. The pipeline is capable of finding any peptides expected, including those that cross intron-exon boundaries, and those due to single nucleotide polymorphisms (SNPs), alternate splicing, and post-translational modifications (PTMs). GAPP can therefore be used to re-annotate genomes, and this is supported through the inclusion of a Distributed Annotation System (DAS) server, which allows the peptides identified by the pipeline to be displayed in their genomic context within the Ensembl genome browser. GAPP is freely available via the web, at www. gapp.info.

Alternative Splicing↗

DING proteins are from Pseudomonas.

DING proteins have been described as animal and plant proteins with potential biomineralisation, receptor or signalling roles that have been characterised by an N-terminal DINGGG-sequence. However, these sequences have only ever been identified as either N-terminal peptides or partial cDNA sequences, and have yet to be detected in any of the many genomic animal and plant genomes now available. Microbial relatives of the DING proteins have been described, which appear to be periplasmic phosphate-binding proteins. Recently, full-length Pseudomonas aeruginosa UCBPP-PA14 and Hypericum perforatum genes have been sequenced that show high homology to the published DING protein N-terminal sequences, and small peptides previously identified in conjunction with the peptide sequencing of DING proteins can also be mapped to regions across these full-length sequences. Searching with these sequences identifies other plant and animal cDNA fragments in the public nucleotide databases, and, additionally, an unordered rat genomic contig that contains a DING-like sequence on a small fragment. Analysing the codon usage of these DNA sequences identifies all of these sequences as of Pseudomonas origin, suggesting that DING proteins do not exist in eukaryotes, but instead are potentially due to microbial contamination or infection.

Amino Acid Sequence↗

Determination of partial amino acid composition from tandem mass spectra for use in peptide identification strategies.

We demonstrate a new approach to the determination of amino acid composition from tandem mass spectrometrically fragmented peptides using both experimental and simulated data. The approach has been developed to be used as a search-space filter in a protein identification pipeline with the aim of increased performance above that which could be attained by using immonium ion information. Three automated methods have been developed and tested: one based upon a simple peak traversal, in which all intense ion peaks are treated as being either a b- or y-ion using a wide mass tolerance; a second which uses a much narrower tolerance and does not perform transformations of ion peaks to the complementary type; and the unique fragments method which allows for b- or y-ion type to be inferred and corroborated using a scan of the other ions present in each peptide spectrum. The combination of these methods is shown to provide a high-accuracy set of amino acid predictions using both experimental and simulated data sets. These high quality predictions, with an accuracy of over 85%, may be used to identify peptide fragments that are hard to identify using other methods. The data simulation algorithm is also shown post priori to be a good model of noiseless tandem mass spectrometric peptide data.

Algorithms↗

Protein and peptide identification algorithms using MS for use in high-throughput, automated pipelines.

Current proteomics experiments can generate vast quantities of data very quickly, but this has not been matched by data analysis capabilities. Although there have been a number of recent reviews covering various aspects of peptide and protein identification methods using MS, comparisons of which methods are either the most appropriate for, or the most effective at, their proposed tasks are not readily available. As the need for high-throughput, automated peptide and protein identification systems increases, the creators of such pipelines need to be able to choose algorithms that are going to perform well both in terms of accuracy and computational efficiency. This article therefore provides a review of the currently available core algorithms for PMF, database searching using MS/MS, sequence tag searches and de novo sequencing. We also assess the relative performances of a number of these algorithms. As there is limited reporting of such information in the literature, we conclude that there is a need for the adoption of a system of standardised reporting on the performance of new peptide and protein identification algorithms, based upon freely available datasets. We go on to present our initial suggestions for the format and content of these datasets.

Algorithms↗

Confident protein identification using the average peptide score method coupled with search-specific, ab initio thresholds.

Perhaps the greatest difficulty in interpreting large sets of protein identifications derived from mass spectrometric methods is whether or not to trust the results. For such experiments, the level of confidence in each protein identification made needs to be far greater than the often used 95% significance threshold to avoid the identification of many false-positives. To provide higher confidence results, we have developed an innovative scoring strategy coupling the recently published Average Peptide Score (APS) method with pre-filtering of peptide identifications, using a simple peptide quality filter. Iterative generation of these filters in conjunction with reversed database searching is used to determine the correct levels at which the APS and peptide quality thresholds should be set to return virtually zero false-positive reports. This proceeds without the need to reference a known dataset.

Computational Biology↗