PubMed Health⌕ Search

Biomedical subjects

Michael L Nielsen

Publications and source records attributed to Michael L Nielsen.

16 recordsLinked to original sources

De novo peptide sequencing and identification with precision mass spectrometry.

The recent proliferation of novel mass spectrometers such as Fourier transform, QTOF, and OrbiTrap marks a transition into the era of precision mass spectrometry, providing a 2 orders of magnitude boost to the mass resolution, as compared to low-precision ion-trap detectors. We investigate peptide de novo sequencing by precision mass spectrometry and explore some of the differences when compared to analysis of low-precision data. We demonstrate how the dramatically improved performance of de novo sequencing with precision mass spectrometry paves the way for novel approaches to peptide identification that are based on direct sequence lookups, rather than comparisons of spectra to a database. With the direct sequence lookup, it is not only possible to search a database very efficiently, but also to use the database in novel ways, such as searching for products of alternative splicing or products of fusion proteins in cancer. Our de novo sequencing software is available for download at http://peptide.ucsd.edu/.

Algorithms↗

Liquid chromatography at critical conditions: comprehensive approach to sequence-dependent retention time prediction.

An approach to sequence-dependent retention time prediction of peptides based on the concept of liquid chromatography at critical conditions (LCCC) is presented. Within the LCCC approach applied to biopolymers (BioLCCC), the specific retention time corresponds to a particular sequence. In combination with mass spectrometry, this approach provides an efficient tool to solve problems wherein the protein sequencing is essential. In this work, we present a theoretical background of the BioLCCC concept and demonstrate experimentally its feasibility for sequence-dependent LC retention time prediction for peptides. BioLCCC model is based on three notions: (a) a random walk model for a macromolecule chain; (b) an entropy and energy compensation for the macromolecules within the adsorbent pore; and (c) a set of phenomenological parameters for the effective interaction energies of interactions between the amino acid residues and the adsorbent surface. In this work, the phenomenological parameters have been obtained for C18 reversed-phase HPLC. Note, that contrary to alternative additive models for retention time prediction based on summation of the so-called "retention coefficients", the BioLCCC approach takes into account the location of amino acids within the primary structure of a peptide and, thus, allows the identification of the peptides having the same composition of amino acids but differing by their arrangement. As a result, this new approach allows prediction of retention time for any possible amino acid sequence in particular HPLC experiments. In addition, the BioLCCC model lacks of main drawbacks of additive approaches that predict retention time for sequences of limited chain lengths and provide information about amino acid composition only. The proposed BioLCCC approach was characterized experimentally using LTQ FT LC-MS and LC-MS/MS data obtained earlier for Escherichia coli. The HPLC system calibration was performed using peptide retention standards. The results received show a linear correlation between predicted and experimental retention times, with a correlation coefficient, R2, of 0.97 for a peptide standard mixture and 0.9 for E. coli data, respectively, with the standard error below 1 min. The work presents the first description of a BioLCCC approach for high-throughput peptide characterization and preliminary results of its feasibility tests.

Adsorption↗

Hydrogen rearrangement to and from radical z fragments in electron capture dissociation of peptides.

Hydrogen rearrangement is an important process in radical chemistry. A high degree of H. rearrangement to and from z. ionic fragments (combined occurrence frequency 47% compared with that of z.) is confirmed in analysis of 15,000 tandem mass spectra of tryptic peptides obtained with electron capture dissociation (ECD), including previously unreported double H. losses. Consistent with the radical character of H. abstraction, the residue determining the formation rate of z' = z. + H. species is found to be the N-terminal residue in z. species. The size of the complementary c(m)' fragment turned out to be another important factor, with z' species dominating over z. ions for m < or = 6. The H. atom was found to be abstracted from the side chains as well as from alpha-carbon groups of residues composing the c' species, with Gln and His in the c' fragment promoting H. donation and Asp and Ala opposing it. Ab initio calculations of formation energies of .A radicals (A is an amino acid) confirmed that the main driving force for H. abstraction by z. is the process exothermicity. No valid correlation was found between the NC(alpha) bond strength and the frequency of this bond cleavage, indicating that other factors than thermochemistry are responsible for directing the site of ECD cleavage. Understanding hydrogen attachment to and loss from ECD fragments should facilitate automatic interpretation ECD mass spectra in protein identification and characterization, including de novo sequencing.

Cell Line↗

Extent of modifications in human proteome samples and their effect on dynamic range of analysis in shotgun proteomics.

The complexity of the human proteome, already enormous at the organism level, increases further in the course of the proteome analysis due to in vitro sample evolution. Most of in vitro alterations can also occur in vivo as post-translational modifications. These two types of modifications can only be distinguished a posteriori but not in the process of analysis, thus rendering necessary the analysis of every molecule in the sample. With the new software tool ModifiComb applied to MS/MS data, the extent of modifications was measured in tryptic mixtures representing the full proteome of human cells. The estimated level of 8-12 modified peptides per each unmodified tryptic peptide present at >or=1% level is approaching one modification per amino acid on average. This is a higher modification rate than was previously thought, posing an additional challenge to analytical techniques. The solution to the problem is seen in improving sample preparation routines, introducing dynamic range-adjusted thresholds for database searches, using more specific MS/MS analysis using high mass accuracy and complementary fragmentation techniques, and revealing peptide families with identification of additional proteins only by unfamiliar peptides. Extensive protein separation prior to analysis reduces the requirements on speed and dynamic range of a tandem mass spectrometer and can be a viable alternative to the shotgun approach.

Amino Acid Sequence↗

ModifiComb, a new proteomic tool for mapping substoichiometric post-translational modifications, finding novel types of modifications, and fingerprinting complex protein mixtures.

A major challenge in proteomics is to fully identify and characterize the post-translational modification (PTM) patterns present at any given time in cells, tissues, and organisms. Here we present a fast and reliable method ("ModifiComb") for mapping hundreds types of PTMs at a time, including novel and unexpected PTMs. The high mass accuracy of Fourier transform mass spectrometry provides in many cases unique elemental composition of the PTM through the difference DeltaM between the molecular masses of the modified and unmodified peptides, whereas the retention time difference DeltaRT between their elution in reversed-phase liquid chromatography provides an additional dimension for PTM identification. Abundant sequence information obtained with complementary fragmentation techniques using ion-neutral collisions and electron capture often locates the modification to a single residue. The (DeltaM, DeltaRT) maps are representative of the proteome and its overall modification state and may be used for database-independent organism identification, comparative proteomic studies, and biomarker discovery. Examples of newly found modifications include +12.000 Da (+C atom) incorporation into proline residues of peptides from proline-rich proteins found in human saliva. This modification is hypothesized to increase the known activity of the peptide.

Adult↗

PhosTShunter: a fast and reliable tool to detect phosphorylated peptides in liquid chromatography Fourier transform tandem mass spectrometry data sets.

A database independent search algorithm for the detection of phosphopeptides is described. The program interrogates the tandem mass spectra of LC-MS/MS data sets regarding the presence of phosphorylation specific signatures. To achieve maximum informational content, the complementary fragmentation techniques electron capture dissociation (ECD) and collisionally activated dissociation (CAD) are used independently for peptide fragmentation. Several criteria characteristic for peptides phosphorylated on either serine or threonine residues were evaluated. The final algorithm searches for product ions generated by either the neutral loss of phosphoric acid or the combined neutral loss of phosphoric acid and water. Various peptide mixtures were used to evaluate the program. False positive results were not observed because the program utilizes the parts-per-million mass accuracy of Fourier transform ion cyclotron resonance mass spectrometry. Additionally, false negative results were not generated owing to the high sensitivity of the chosen criteria. The limitations of database dependent data interpretation tools are discussed and the potential of the novel algorithm to overcome these limitations is illustrated.

Algorithms↗

Efficient PCR-based gene targeting with a recyclable marker for Aspergillus nidulans.

The rapid accumulation of genomic sequences from a large number of eukaryotes, including numerous filamentous fungi, has created a tremendous scientific potential, which can only be realized if precise site-directed genome modifications, like gene deletions, promoter replacements, in-frame GFP fusions and specific point mutations can be made rapidly and reliably. The development of gene-targeting techniques in filamentous fungi and other higher eukaryotes has been hampered because foreign DNA is predominantly integrated randomly into the genome. For Aspergillus nidulans, we have developed a flexible method for gene-targeting employing a bipartite gene-targeting substrate. This substrate is made solely by PCR, which obviates the need for bacterial subcloning steps. The method reduces the number of false positives and can be used to produce virtually any genome alteration. A major advance of the method is that it allows multiple subsequent genome manipulations to be performed as the selectable marker is recycled.

Aspergillus nidulans↗

New data base-independent, sequence tag-based scoring of peptide MS/MS data validates Mowse scores, recovers below threshold data, singles out modified peptides, and assesses the quality of MS/MS techniques.

The Mascot score (M-score) is one of the conventional validity measures in data base identification of peptides and proteins by MS/MS data. Although tremendously useful, M-score has a number of limitations. For the same MS/MS data, M-score may change if the protein data base is expanded. A low M-value may not necessarily mean poor match but rather poor MS/MS quality. In addition M-score does not fully utilize the advantage of combined use of complementary fragmentation techniques collisionally activated dissociation (CAD) and electron capture dissociation (ECD). To address these issues, a new data base-independent scoring method (S-score) was designed that is based on the maximum length of the peptide sequence tag provided by the combined CAD and ECD data. The quality of MS/MS spectra assessed by S-score allows poor data (39% of all MS/MS spectra) to be filtered out before the data base search, speeding up the data analysis and eliminating a major source of false positive identifications. Spectra with below threshold M-scores (poor matches) but high S-scores are validated. Spectra with zero M-score (no data base match) but high S-score are classified as belonging to modified sequences. As an extension of S-score, an extremely reliable sequence tag was developed based on complementary fragments simultaneously appearing in CAD and ECD spectra. Comparison of this tag with the data base-derived sequence gives the most reliable peptide identification validation to date. The combined use of M- and S-scoring provides positive sequence identification from >25% of all MS/MS data, a 40% improvement over traditional M-scoring performed on the same Fourier transform MS instrumentation. The number of proteins reliably identified from Escherichia coli cell lysate hereby increased by 29% compared with the traditional M-score approach. Finally S-scoring provides a quantitative measure of the quality of fragmentation techniques such as the minimum abundance of the precursor ion, the MS/MS of which gives the threshold S-score value of 2.

Bacterial Proteins↗

Improving protein identification using complementary fragmentation techniques in fourier transform mass spectrometry.

Identification of proteins by MS/MS is performed by matching experimental mass spectra against calculated spectra of all possible peptides in a protein data base. The search engine assigns each spectrum a score indicating how well the experimental data complies with the expected one; a higher score means increased confidence in the identification. One problem is the false-positive identifications, which arise from incomplete data as well as from the presence of misleading ions in experimental mass spectra due to gas-phase reactions, stray ions, contaminants, and electronic noise. We employed a novel technique of reduction of false positives that is based on a combined use of orthogonal fragmentation techniques electron capture dissociation (ECD) and collisionally activated dissociation (CAD). Since ECD and CAD exhibit many complementary properties, their combined use greatly increased the analysis specificity, which was further strengthened by the high mass accuracy (approximately 1 ppm) afforded by Fourier transform mass spectrometry. The utility of this approach is demonstrated on a whole cell lysate from Escherichia coli. Analysis was made using the data-dependent acquisition mode. Extraction of complementary sequence information was performed prior to data base search using in-house written software. Only masses involved in complementary pairs in the MS/MS spectrum from the same or orthogonal fragmentation techniques were submitted to the data base search. ECD/CAD identified twice as many proteins at a fixed statistically significant confidence level with on average a 64% higher Mascot score. The confidence in protein identification was hereby increased by more than 1 order of magnitude. The combined ECD/CAD searches were on average 20% faster than CAD-only searches. A specially developed test with scrambled MS/MS data revealed that the amount of false-positive identifications was dramatically reduced by the combined use of CAD and ECD.

Cyclotrons↗

Physicochemical properties determining the detection probability of tryptic peptides in Fourier transform mass spectrometry. A correlation study.

Sequence verification and mapping of posttranslational modifications require nearly 100% sequence coverage in the "bottom-up" protein analysis. Even in favorable cases, routine liquid chromatography-mass spectrometry detects from protein digests peptides covering 50-90% of the sequence. Here we investigated the reasons for limited peptide detection, considering various physicochemical aspects of peptide behavior in liquid chromatography-Fourier transform mass spectrometry (LC-FTMS). No overall correlation was found between the detection probability and peptide mass. In agreement with literature data, the signal increased with peptide hydrophobicity. Surprisingly, the pI values exhibited an opposite trend, with more acidic tryptic peptides detected with higher probability. A mixture of synthesized peptides of similar masses confirmed the hydrophobicity dependence but showed strong positive correlation between pI and signal response. An explanation of this paradoxal behavior was found through the observation that more acidic tryptic peptide lengths tend to be longer. Longer peptides tend to acquire higher average charge state in positive mode electrospray ionization than more basic but shorter counterparts. The induced-current detection in FTMS favors ions in higher charge states, thus providing the observed pI-FTMS signal anticorrelation.

Animals↗

Shifted-basis technique improves accuracy of peak position determination in Fourier transform mass spectrometry.

The present paper suggests a new algorithm for estimation of peak positions in FTMS spectra. It is shown theoretically and experimentally that the new technique yields superior results compared to the currently applied techniques, when the noise level is high and/or the peaks are located close to each other. Cases are presented where the deviation from the true mass could be mistaken for space charge effect, while the shift is in fact solely due to the shortcomings of the current techniques and can be corrected by applying the shifted-basis technique. In two out of three cases, this technique gave more accurate (>5 times) result compared to the conventional analysis. In the third case, where the signal was high compared to the noise, the results were comparable. The new technique can be used to achieve better mass accuracy for noisy and not well resolved spectra, and to further investigate the features of the space charge effect.

Algorithms↗

HysTag--a novel proteomic quantification tool applied to differential display analysis of membrane proteins from distinct areas of mouse brain.

A novel isotopically labeled cysteine-tagging and complexity-reducing reagent, called HysTag, has been synthesized and used for quantitative proteomics of proteins from enriched plasma membrane preparations from mouse fore- and hindbrain. The reagent is a 10-mer derivatized peptide, H(2)N-(His)(6)-Ala-Arg-Ala-Cys(2-thiopyridyl disulfide)-CO(2)H, which consists of four functional elements: i) an affinity ligand (His(6)-tag), ii) a tryptic cleavage site (-Arg-Ala-), iii) Ala-9 residue that contains four (d(4)) or no (d(0)) deuterium atoms, and iv) a thiol-reactive group (2-thiopyridyl disulfide). For differential analysis cysteine residues in the compared samples are modified using either (d(4)) or (d(0)) reagent. The HysTag peptide is preserved in Lys-C digestion of proteins and allows charge-based selection of cysteine-containing peptides, whereas subsequent tryptic digestion reduces the labeling group to a di-peptide, which does not hinder effective fragmentation. Furthermore, we found that tagged peptides containing Ala-d(4) co-elute with their d(0)-labeled counterparts. To demonstrate effectiveness of the reagent, a differential analysis of mouse forebrain versus hindbrain plasma membranes was performed. Enriched plasma membrane fractions were partially denatured, reduced, and reacted with the reagent. Digestion with endoproteinase Lys-C was carried out on nonsolubilized membranes. The membranes were sedimented by ultra centrifugation, and the tagged peptides were isolated by Ni(2+) affinity or cation-exchange chromatography. Finally, the tagged peptides were cleaved with trypsin to release the histidine tag (residues 1-8 of the reagent) followed by liquid chromatography tandem mass spectroscopy for relative protein quantification and identification. A total of 355 unique proteins were identified, among which 281 could be quantified. Among a large majority of proteins with ratios close to one, a few proteins with significant quantitative changes were retrieved. The HysTag offers advantages compared with the isotope-coded affinity tag reagent, because the HysTag reagent is easy to synthesize, economical due to use of deuterium instead of (13)C isotope label, and allows robust purification and flexibility through the affinity tag, which can be extended to different peptide functionalities.

Amino Acid Sequence↗

Peptide end sequencing by orthogonal MALDI tandem mass spectrometry.

Highly sensitive peptide fragmentation and identification in sequence databases is a cornerstone of proteomics. Previously, a two-layered strategy consisting of MALDI peptide mass fingerprinting followed by electrospray tandem mass spectrometry of the unidentified proteins has been successfully employed. Here, we describe a high-sensitivity/high-throughput system based on orthogonal MALDI tandem mass spectrometry (o-MALDI) and the automated recognition of fragments corresponding to the N- and C-terminal amino acid residues. Robotic deposition of samples onto hydrophobic anchor substrates is employed, and peptide spectra are acquired automatically. The pulsing feature of the QSTAR o-MALDI mass spectrometer enhances the low mass region of the spectra by approximately 1 order of magnitude. Software has been developed to automatically recognize characteristic features in the low mass region (such as the y1 ion of tryptic peptides), maintaining high mass accuracy even with very low count events. Typically, the sum of the N-terminal two ions (b2 ion), the third N-terminal ion (b3 ion), and the two C-terminal fragments of the peptide (y1 and y2) can be determined. Given mass accuracy in the low ppm range, peptide end sequencing on one or two tryptic peptides is sufficient to uniquely identify a protein from gel samples in the low silver-stained range.

Calibration↗

Proteomics-grade de novo sequencing approach.

The conventional approach in modern proteomics to identify proteins from limited information provided by molecular and fragment masses of their enzymatic degradation products carries an inherent risk of both false positive and false negative identifications. For reliable identification of even known proteins, complete de novo sequencing of their peptides is desired. The main problems of conventional sequencing based on tandem mass spectrometry are incomplete backbone fragmentation and the frequent overlap of fragment masses. In this work, the first proteomics-grade de novo approach is presented, where the above problems are alleviated by the use of complementary fragmentation techniques CAD and ECD. Implementation of a high-current, large-area dispenser cathode as a source of low-energy electrons provided efficient ECD of doubly charged peptides, the most abundant species (65-80%), in a typical trypsin-based proteomics experiment. A new linear de novo algorithm is developed combining efficiency and speed, processing on a conventional 3 GHz PC, 1000 MS/MS data sets in 60 s. More than 6% of all MS/MS data for doubly charged peptides yielded complete sequences, and another 13% gave nearly complete sequences with a maximum gap of two amino acid residues. These figures are comparable with the typical success rates (5-15%) of database identification. For peptides reliably found in the database (Mowse score > or = 34), the agreement with de novo-derived full sequences was >95%. Full sequences were derived in 67% of the cases when full sequence information was present in MS/MS spectra. Thus the new de novo sequencing approach reached the same level of efficiency and reliability as conventional database-identification strategies.

Algorithms↗