PubMed Health⌕ Search

Biomedical subjects

Parag Mallick

Publications and source records attributed to Parag Mallick.

17 recordsLinked to original sources

Computational prediction of proteotypic peptides for quantitative proteomics.

Mass spectrometry-based quantitative proteomics has become an important component of biological and clinical research. Although such analyses typically assume that a protein's peptide fragments are observed with equal likelihood, only a few so-called 'proteotypic' peptides are repeatedly and consistently identified for any given protein present in a mixture. Using >600,000 peptide identifications generated by four proteomic platforms, we empirically identified >16,000 proteotypic peptides for 4,030 distinct yeast proteins. Characteristic physicochemical properties of these peptides were used to develop a computational tool that can predict proteotypic peptides for any protein from any organism, for a given platform, with >85% cumulative accuracy. Possible applications of proteotypic peptides include validation of protein identifications, absolute quantification of proteins, annotation of coding sequences in genomes, and characterization of the physical principles governing key elements of mass spectrometric workflows (e.g., digestion, chromatography, ionization and fragmentation).

Algorithms↗

The PeptideAtlas project.

The completion of the sequencing of the human genome and the concurrent, rapid development of high-throughput proteomic methods have resulted in an increasing need for automated approaches to archive proteomic data in a repository that enables the exchange of data among researchers and also accurate integration with genomic data. PeptideAtlas (http://www.peptideatlas.org/) addresses these needs by identifying peptides by tandem mass spectrometry (MS/MS), statistically validating those identifications and then mapping identified sequences to the genomes of eukaryotic organisms. A meaningful comparison of data across different experiments generated by different groups using different types of instruments is enabled by the implementation of a uniform analytic process. This uniform statistical validation ensures a consistent and high-quality set of peptide and protein identifications. The raw data from many diverse proteomic experiments are made available in the associated PeptideAtlas repository in several formats. Here we present a summary of our process and details about the Human, Drosophila and Yeast PeptideAtlas builds.

Animals↗

Quantitative proteomic analysis of the budding yeast cell cycle using acid-cleavable isotope-coded affinity tag reagents.

Quantitative profiling of proteins, the direct effectors of nearly all biological functions, will undoubtedly complement technologies for the measurement of mRNA. Systematic proteomic measurement of the cell cycle is now possible by using stable isotopic labeling with isotope-coded affinity tag reagents and software tools for high-throughput analysis of LC-MS/MS data. We provide here the first such study achieving quantitative, global proteomic measurement of a time-course gene expression experiment in a model eukaryote, the budding yeast Saccharomyces cerevisiae, during the cell cycle. We sampled 48% of all predicted ORFs, and provide the data, including identifications, quantitations, and statistical measures of certainty, to the community in a sortable matrix. We do not detect significant concordance in the dynamics of the system over the time-course tested between our proteomic measurements and microarray measures collected from similarly treated yeast cultures. Our proteomic dataset therefore provides a necessary and complementary measure of eukaryotic gene expression, establishes a rich database for the functional analysis of S. cerevisiae proteins, and will enable further development of technologies for global proteomic analysis of higher eukaryotes.

Carbon Isotopes↗

Mutagenesis of putative serine-threonine phosphorylation sites proximal to Arg255 of human cytochrome P450c17 does not selectively promote its 17,20-lyase activity.

OBJECTIVE: To investigate the role of serine-threonine phosphorylation on the activity of human P450c17. DESIGN: In vitro study. SETTING: Academic basic research laboratory. PATIENT(S): None. INTERVENTION(S): P450c17 expression constructs with a FLAG-tag on either the C-terminus or N-terminus of the protein were generated. Human C-terminal FLAG-tagged P450c17 chromosomal DNA was subjected to site-directed mutagenesis. Serine 258 and threonine 260 each were mutated to alanine and aspartic acid. The mutant P450c17s were expressed in COS-7 cells, and the enzymatic activities were measured. MAIN OUTCOME MEASURE(S): 17alpha-Hydroxylase and C(17-20) lyase activities of human P450c17. RESULT(S): C-terminal FLAG-tagged P450c17 functioned indistinguishably from the wild-type P450c17. Mutants S258A, S258D, and T260D had significantly less 17alpha-hydroxylase and C(17-20) lyase activities than the wild type. CONCLUSION(S): Adding an epitope tag to the C-terminus of the P450c17 protein does not interfere with its activities and will be a useful tool to isolate human P450c17 protein from cultured cells. Phosphorylation of serine 258 but not threonine 260 may act as a physiologic regulator of both enzymatic activities through interaction with obligatory redox partners.

Amino Acids↗

Protein cross-linking analysis using mass spectrometry, isotope-coded cross-linkers, and integrated computational data processing.

Distance constraints in proteins and protein complexes provide invaluable information for calculation of 3D structures, identification of protein binding partners and localization of protein-protein contact sites. We have developed an integrative approach to identify and characterize such sites through the analysis of proteolytic products derived from proteins chemically cross-linked by isotopically coded cross-linkers using LC-MALDI tandem mass spectrometry and computer software. This method is specifically tailored toward the rapid analysis of low microgram amounts of proteins or multimeric protein complexes cross-linked with nonlabeled and deuterium-labeled bis-NHS ester cross-linking reagents (both commercially available and readily synthesized). Through labeling with [18O]water solvent and LC-MALDI analysis, the method further allows the possible distinction between Type 0 and Type 1 or Type 2 modified peptides (monolinks and looplinks or cross-links), although such a distinction is more readily made from analysis of tandem mass spectrometry data. When applied to the bacterial Colicin E7 DNAse/Im7 heterodimeric protein complex, 23 cross-links were identified including six intersubunit cross-links, all between residues that are close in space when examined in the context of the X-ray structure of the heterodimer. In addition, cross-links were successfully identified in five single subunit proteins, beta-lactoglobulin, cytochrome c, lysozyme, myoglobin, and ribonuclease A, establishing the generality of the approach.

Chromatography, Liquid↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

Signal maps for mass spectrometry-based comparative proteomics.

Mass spectrometry-based proteomic experiments, in combination with liquid chromatography-based separation, can be used to compare complex biological samples across multiple conditions. These comparisons are usually performed on the level of protein lists generated from individual experiments. Unfortunately given the current technologies, these lists typically cover only a small fraction of the total protein content, making global comparisons extremely limited. Recently approaches have been suggested that are built on the comparison of computationally built feature lists instead of protein identifications. Although these approaches promise to capture a bigger spectrum of the proteins present in a complex mixture, their success is strongly dependent on the correctness of the identified features and the aligned retention times of these features across multiple experiments. In this experimental-computational study, we went one step further and performed the comparisons directly on the signal level. First signal maps were constructed that associate the experimental signals across multiple experiments. Then a feature detection algorithm used this integrated information to identify those features that are discriminating or common across multiple experiments. At the core of our approach is a score function that faithfully recognizes mass spectra from similar peptide mixtures and an algorithm that produces an optimal alignment (time warping) of the liquid chromatography experiments on the basis of raw MS signal, making minimal assumptions on the underlying data. We provide experimental evidence that suggests uniqueness and correctness of the resulting signal maps even on low accuracy mass spectrometers. These maps can be used for a variety of proteomic analyses. Here we illustrate the use of signal maps for the discovery of diagnostic biomarkers. An imple-mentation of our algorithm is available on our Web server.

Algorithms↗

Scoring proteomes with proteotypic peptide probes.

Technologies for genome-wide analyses typically undergo a transition from a discovery phase to a scoring phase. In the discovery phase, the genomic universe is explored and all pertinent data are noted. In the scoring phase, relevant entities are screened to reveal groups of genes that are associated with specific biological processes or conditions. In this article, we propose that the transition from a discovery to a scoring phase is also essential, feasible and imminent for proteomics.

Amino Acid Sequence↗

High throughput quantitative analysis of serum proteins using glycopeptide capture and liquid chromatography mass spectrometry.

It is expected that the composition of the serum proteome can provide valuable information about the state of the human body in health and disease and that this information can be extracted via quantitative proteomic measurements. Suitable proteomic techniques need to be sensitive, reproducible, and robust to detect potential biomarkers below the level of highly expressed proteins, generate data sets that are comparable between experiments and laboratories, and have high throughput to support statistical studies. Here we report a method for high throughput quantitative analysis of serum proteins. It consists of the selective isolation of peptides that are N-linked glycosylated in the intact protein, the analysis of these now deglycosylated peptides by liquid chromatography electrospray ionization mass spectrometry, and the comparative analysis of the resulting patterns. By focusing selectively on a few formerly N-linked glycopeptides per serum protein, the complexity of the analyte sample is significantly reduced and the sensitivity and throughput of serum proteome analysis are increased compared with the analysis of total tryptic peptides from unfractionated samples. We provide data that document the performance of the method and show that sera from untreated normal mice and genetically identical mice with carcinogen-induced skin cancer can be unambiguously discriminated using unsupervised clustering of the resulting peptide patterns. We further identify, by tandem mass spectrometry, some of the peptides that were consistently elevated in cancer mice compared with their control littermates.

Animals↗

Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.

A crucial aim upon the completion of the human genome is the verification and functional annotation of all predicted genes and their protein products. Here we describe the mapping of peptides derived from accurate interpretations of protein tandem mass spectrometry (MS) data to eukaryotic genomes and the generation of an expandable resource for integration of data from many diverse proteomics experiments. Furthermore, we demonstrate that peptide identifications obtained from high-throughput proteomics can be integrated on a large scale with the human genome. This resource could serve as an expandable repository for MS-derived proteome information.

Amino Acid Sequence↗

PFIT and PFRIT: bioinformatic algorithms for detecting glycosidase function from structure and sequence.

The identification of the enzymes involved in the metabolism of simple and complex carbohydrates presents one bioinformatic challenge in the post-genomic era. Here, we present the PFIT and PFRIT algorithms for identifying those proteins adopting the alpha/beta barrel fold that function as glycosidases. These algorithms are based on the observation that proteins adopting the alpha/beta barrel fold share positions in their tertiary structures having equivalent sets of atomic interactions. These are conserved tertiary interaction positions, which have been implicated in both structure and function. Glycosidases adopting the alpha/beta barrel fold share more conserved tertiary interactions than alpha/beta barrel proteins having other functions. The enrichment pattern of conserved tertiary interactions in the glycosidases is the information that PFIT and PFRIT use to predict whether any given alpha/beta barrel will function as a glycosidase or not. Using as a test set a database of 19 glycosidase and 45 nonglycosidase alpha/beta barrel proteins with low sequence similarity, PFIT and PFRIT can correctly predict glycosidase function for 84% of the proteins known to function as glycosidases. PFIT and PFRIT incorrectly predict glycosidase function for 25% of the nonglycosidases. The program PSI-BLAST can also correctly identify 84% of the 19 glycosidases, however, it incorrectly predicts glycosidase function for 50% of the nonglycosidases (twofold greater than PFIT and PFRIT). Overall, we demonstrate that the structure-based PFIT and PFRIT algorithms are both more selective and sensitive for predicting glycosidase function than the sequence-based PSI-BLAST algorithm.

Algorithms↗

Inference of protein function and protein linkages in Mycobacterium tuberculosis based on prokaryotic genome organization: a combined computational approach.

The genome of Mycobacterium tuberculosis was analyzed using recently developed computational approaches to infer protein function and protein linkages. We evaluated and employed a method to infer genes likely to belong to the same operon, as judged by the nucleotide distance between genes in the same genomic orientation, and combined this method with those of the Rosetta Stone, Phylogenetic Profile and conserved Gene Neighbor computational methods for the inference of protein function.

Bacterial Proteins↗

The directional atomic solvation energy: an atom-based potential for the assignment of protein sequences to known folds.

The Directional Atomic Solvation EnergY (DASEY) is an atom-based description of the environment of an amino acid position within a known 3D protein structure. The DASEY has been developed to align and score a probe amino acid sequence to a library of template protein structures for fold assignment. DASEY is computed by summing the atomic solvation parameters of atoms falling within a tetrahedral sector, or petal, extending 16 A along each of the four bond axes of each alpha-carbon atom of the protein. The DASEY discriminates between pairs of structurally equivalent positions and random pairs in protein structures sharing a fold but belonging to different superfamilies, unlike some previous descriptors of protein environments, such as buried area. Furthermore, the DASEY values have characteristic patterns of residue replacement, an essential feature of a successful fold assignment method. Benchmarking fold assignment with DASEY achieves coverage of 56% of sequences with 90% accuracy when probe sequences are matched to protein structural templates belonging to the same fold but to a different superfamily, an improvement of greater than 200% over a previous method.

Algorithms↗

Genomic evidence that the intracellular proteins of archaeal microbes contain disulfide bonds.

Disulfide bonds have only rarely been found in intracellular proteins. That pattern is consistent with the chemically reducing environment inside the cells of well-studied organisms. However, recent experiments and new calculations based on genomic data of archaea provide striking contradictions to this pattern. Our results indicate that the intracellular proteins of certain hyperthermophilic archaea, especially the crenarchaea Pyrobaculum aerophilum and Aeropyrum pernix, are rich in disulfide bonds. This finding implicates disulfide bonding in stabilizing many thermostable proteins and points to novel chemical environments inside these microbes. These unexpected results illustrate the wealth of biochemical insights available from the growing reservoir of genomic data.

Archaeal Proteins↗

A modeled hydrophobic domain on the TCL1 oncoprotein mediates association with AKT at the cytoplasmic membrane.

AKT has a critical role in relaying cell survival and proliferation signals initiated by ligand binding to surface receptors in mammalian cells. Induction of AKT serine/threonine kinase activity is augmented by the T-cell leukemia-1 (TCL1) oncoprotein through a physical association requiring the AKT pleckstrin homology domain. Here, we used molecular modeling and identified an exposed hydrophobic patch composed of two discontinuous amino acid stretches near one end of the TCL1 beta-barrel that was required for a TCL1-AKT association. Site-directed mutations of this region did not affect TCL1 secondary structure, yet they disrupted interactions with AKT. This region was found in other members of the TCL1 oncoprotein family, such as TCL1b and MTCP1, and suggested a conserved, novel AKT binding domain. Interestingly, TCL1 and AKT co-localize in multiple cell compartments, but only extracts from the plasma membrane stimulate optimal complex formation in vitro. Identification of an AKT binding domain on TCL1 is an important step in deciphering the complex interactions that regulate AKT kinase activity in lymphocyte development and neoplasia within the immune system.

Animals↗

GXXXG and AXXXA: common alpha-helical interaction motifs in proteins, particularly in extremophiles.

The GXXXG motif is a frequently occurring sequence of residues that is known to favor helix-helix interactions in membrane proteins. Here we show that the GXXXG motif is also prevalent in soluble proteins whose structures have been determined. Some 152 proteins from a non-redundant PDB set contain at least one alpha-helix with the GXXXG motif, 41 +/- 9% more than expected if glycine residues were uniformly distributed in those alpha-helices. More than 50% of the GXXXG-containing alpha-helices participate in helix-helix interactions. In fact, 26 of those helix-helix interactions are structurally similar to the helix-helix interaction of the glycophorin A dimer, where two transmembrane helices associate to form a dimer stabilized by the GXXXG motif. As for the glycophorin A structure, we find backbone-to-backbone atomic contacts of the C alpha-H...O type in each of these 26 helix-helix interactions that display the stereochemical hallmarks of hydrogen bond formation. These glycophorin A-like helix-helix interactions are enriched in the general set of helix-helix interactions containing the GXXXG motif, suggesting that the inferred C alpha-H...O hydrogen bonds stabilize the helix-helix interactions. In addition to the GXXXG motif, some 808 proteins from the non-redundant PDB set contain at least one alpha-helix with the AXXXA motif (30 +/- 3% greater than expected). Both the GXXXG and AXXXA motifs occur frequently in predicted alpha-helices from 24 fully sequenced genomes. Occurrence of the AXXXA motif is enhanced to a greater extent in thermophiles than in mesophiles, suggesting that helical interaction based on the AXXXA motif may be a common mechanism of thermostability in protein structures. We conclude that the GXXXG sequence motif stabilizes helix-helix interactions in proteins, and that the AXXXA sequence motif also stabilizes the folded state of proteins.

Amino Acid Motifs↗