PubMed HealthSearch

SEARCH · PubMed Health

Results for “peptide identification”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

On-filter fractionation by empFASP improves identification of membrane peptides in proteomic experiments.

Membrane proteins remain among the most analytically challenging targets in bottom-up proteomics due to their limited solubility and low abundance of protease-accessible sites within transmembrane domains. In addition, hydrophobic peptides are frequently lost during detergent removal and the on-filter processing steps. Here, we present empFASP, a straightforward on-filter-fractionation-based modification of the enhanced filter-aided sample preparation (eFASP) workflow that enhances recovery of membrane-embedded peptides otherwise lost during digestion and cleanup. The method combines controlled on-filter inversion with sequential ethyl acetate extraction at defined pH values, enabling recovery of peptide material retained on the filter and redistributed into detergent micelles. Compared with SP3 and SP4 in HEK293T lysates, empFASP increased unique hydrophobic peptide identifications by up to 48% and increased the proportion of detected transmembrane peptides. Application to mouse mitochondrial membranes and phosphatidylethanolamine-deficient and PE-containing Escherichia coli membranes showed that the additional fractions of empFASP contribute complementary recovery of hydrophobic and membrane-associated peptides, with the strongest gains observed at the peptide level. Because empFASP requires no specialized reagents or instrumentation, it can be readily implemented in standard proteomics workflows to improve coverage of membrane-embedded regions. SIGNIFICANCE: The empFASP (enhanced membrane peptide) workflow offers a practical solution to one of the persistent limitations in membrane proteomics-the underrepresentation of hydrophobic and transmembrane peptides in standard digests. By integrating simple pH-controlled extractions into an on-filter format, empFASP recovers peptides otherwise lost through adsorption or detergent micelle retention, substantially improving coverage of the membrane proteome. This method expands the analytical reach of bottom-up proteomics without requiring specialized instrumentation, making it immediately applicable for studies of membrane topology, protein-lipid interactions, and the structural consequences of altered membrane composition.

Proteomics

Intestinal surface peptide hydrolases: identification and characterization of three enzymes from rat brush border.

Peptide hydrolases were solubilized from rat small intestinal brush border by papain and separated by Sephadex G-200 chromatography, velocity gradient ultracentrifugation and polyacrylamide disc electrophoresis and designated according to approximate molecular size from sedimentation studies. Peptidases I (apparent Mr 230 000) and II (apparent Mr 160 000) are oligopeptidases with maximum specificity for tripeptides with identical pH optima (7.5) and similar apparent Km with L-Leu-Gly (I, 0.60 MM; II, 0.76 mM). L-Leucyl-beta-naphthylamide is a competitive inhibitor of both enzymes. Concentration of peptidase II produced partial conversion to peptidase I on polyacrylamide disc electrophoresis. The third peptide hydrolase (III, Mr 120 000) is a dipeptidase with pH optimum 8.5 and apparent Km for L-Leu-Gly of 0.65 mM. These peptide hydrolases were inhibited appreciably (37-59%) by 0.2 M glycine/NaOH, Tris - HCl or Tris - glycine buffers. EDTA (5 mM) completely inhibited these enzymes but all activity was restored by dialysis against buffer without divalent ions. Subsequent addition of Mg2+, Mn2+, Co2+ or Zn2+ (1-2 mM) inhibited peptidases I and II variably (4-81%) depending upon the substrate and buffer used. In contrast peptidase III was activated slightly by metal ions (5-20%). These peptide hydrolases are strategically located at the intestinal lumen-cell interface and possess biochemical characteristics making them ideally suited to play a pivotal role in the final stage of protein digestion.

Aminopeptidases

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics

Large-scale discovery platform enables identification of peptides targeting drug-resistant candidiasis.

Natural products have an unparalleled track record as sources of clinical drugs. Among them, nonribosomal peptides (NRPs) stand as one of the most therapeutically significant classes, encompassing numerous approved anti-infective and anticancer agents. Yet, discovering bioactive NRPs remains profoundly challenging due to their complex biosynthesis and chemical architecture. Here, we present NPDiscover, a pathogen-oriented, scalable bioinformatics platform that integrates genome mining, metabolomics, and machine learning to identify NRPs active against drug-resistant pathogens. Applying NPDiscover to Actinobacteria datasets, we discovered edaphochelin A, a previously unreported NRP that kills multi-drug-resistant Candida auris and Candida glabrata by disrupting respiratory chain proteins. Structural elucidation via nuclear magnetic resonance and mass spectrometry, alongside in vitro and in vivo validation, confirmed its efficacy, safety, and a mode of action distinct from existing antifungals-establishing edaphochelin A as a compelling drug candidate and NPDiscover as a powerful engine for scalable natural product discovery.

CP: biotechnology

[Studies on cytochrome c oxidase, I. Purification and characterization of bovine myocardial enzyme and identification of peptide chains in the complex].

As part of the preliminary work for the structural elucidation of cytochrome c oxidase, the enzyme complex was isolated from bovine heart muscle and characterised chemically. The enzyme contains 10-11 nmol haem a, and 12-13 nmol copper per mg protein. The solubilised active enzyme also contains 5% phospholipid, comprising about 2 mol each of cardiolipin and phosphatidylethanolamine per mol haem a. In addition, the preparation contains a small number of detergent molecules (Tween-80). Eight polypeptide components were isolated by preparative dodecylsulphate gel electrophoresis, gel filtration on Biogel P-60, and counter current distribution. The apparent molecular weights of these components were I - 36 000, II - 28 000 (21 000), III - 19 000, IV - 14 000, V - 12 500, VI - 11 000, VII - 10 000 and VIII - 6000. At least seven intact polypeptide chains contribute to the structure of the enzyme complex of the terminal oxidase. On the basis of amino acid analysis and end group determination, they can be divided into two groups. The high molecular weight peptides I -III are hydrophobic and their amino acid compositions differ markedly from those of known enzyme proteins, especially with respect to their contents of leucine and methionine. Components I and II have formyl methionine at their N-termini. They are therefore possibly mitochondrial membrane components from complex 4 of the respiratory chain. Polypeptides IV - VII resemble functional enzyme subunits in their amino acid composition. Some of them possess free N-termini (alanine). The low molecular weight component VIII is heterogeneous and contains the N-terminal amino acids isoleucine, serine and phenylelanine in non-stoichiometric amounts. Analysis gives a minimal protein molecular weight of 130 000 (65 000 per haem a) for the two haem and two copper-containing "monomers". The molecular weight of the moiety preliminarily defined as enzymatic is about 48 000. The chemical characterisation provides data for the strategy of the subsequent sequence analysis of the polypeptides.

Amino Acids

Identification by peptide analysis of the spectrin-binding protein in human erythrocytes.

One-dimensional and two-dimensional peptide-mapping techniques are used to identify the protein which gives rise to the 72,000 dalton alpha-chymotryptic fragment previously shown to be the membrane attachment site for spectrin. Peptide maps of the 72,000 dalton fragment are very different from maps of Bands 1, 2, 2.9, 3, 3.1, 4.1, and 4.2 and very similar to maps of the apparently closely homologous polypeptides, Bands 2.1, 2.2, 2.3, and 2.6. Limited proteolysis of erythrocyte membranes is shown to generate Band 3', another polypeptide which has been associated with spectrin-binding activity. Peptide maps of Band 3' are very similar to maps of Band 2.1, suggesting that Band 3' is also a proteolytic fragment of Band 2.1. It is concluded that Band 2.1 and possibly some or all of the other, related polypeptides which electrophorese in the 2 region is (are) the spectrin-binding protein(s) of the human erythrocyte.

Carrier Proteins

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Strategy for Simultaneous Multiomic Survey of N-Glycomic and Extracellular Matrix Proteome by Mass Spectrometry Imaging.

Recent advances in spatially resolved molecular profiling have positioned matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI) as a powerful platform for multiomic tissue analyses. However, conventional workflows that sequentially target distinct molecular classes are time- and resource-intensive, requiring repeated sequential sample preparation, imaging, and data integration. Here, we evaluate streamlined strategies for simultaneous or combined acquisition of N-glycan and collagen-derived peptide information using PNGase F and collagenase. In-solution studies demonstrate that simultaneous enzymatic digestion yields comparable peptide identifications and glycan profiles relative to traditional sequential workflows, with minimal impact on enzymatic specificity. On the basis of these findings, we developed and optimized MALDI-MSI protocols enabling either simultaneous enzyme application or sequential enzyme treatment with unified matrix deposition and single-pass imaging. While direct coapplication reduced image uniformity, a hybrid approach that used sequential enzyme deposition with combined imaging preserved spatial fidelity and spectral quality while significantly reducing processing and computational demands. Application to human tissues, including vertebral bone and ocular samples, highlights the utility of this workflow for fragile specimens and exploratory multiomic surveys. Collectively, these results establish a framework for integrated glycomic and proteomic imaging targeting the extracellular microenvironment, expanding multiomic MALDI-MSI analyses.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Immunoglobulin D myeloma and amyloidosis: immunochemical and structural studies of Bence Jones and amyloid fibrillar proteins.

Urinary Bence Jones protein and amyloid fibril protein isolated from the subcutaneous tissue of a patient with IgD myeloma and associated amyloidosis were subjected to physicochemical and immunochemical identification. Peptide maps and amino-terminal tetrapeptide composition obtained from the two proteins were comparable. Immunochemical cross-reactivity between the two proteins, with other lambda-type amyloid and Bence Jones proteins, and with a serum component was demonstrated. The results suggest that the source of the amyloid fibril protein is an intact circulating light polypeptide chain as well as smaller amino-terminal fragments.

Aged

Mass Spectrometry-Based Proteomics for the Masses: Peptide and Protein Identification in the Hunt Laboratory During the 2000's.

There has been a rapid increase in the number of individuals utilizing mass spectrometry-based proteomics to study complex biological systems and questions since the start of the 2000's. Building off the advancements in ionization and liquid chromatography scientists continued to push towards technology that would enable in-depth analysis of biological specimen. Donald F Hunt and the Hunt laboratory were major contributors to this effort with their work on improving upon existing Fourier Transform MS, development of electron transfer dissociation, and continued work on ion-ion reactions to improve intact protein analysis. Collaboration with other instrumentation laboratories and instrument companies led to the sharing of technology and eventual commercialization providing greater access. Additionally, the Hunt laboratory spread the gospel of MS-based proteomics through collaborations that lasted decades with other scientists who were experts in immunology, cellular signaling, epigenetics, and other fascinating fields. This article attempts to highlight the many contributions of Don and the Hunt laboratory to peptide and protein identification since the year 2000.

Humans

Isolation and thin-layer chromatographic identification of several peptides with NH2-terminal tryptophan and their histochemical demonstration in the ACTH cells of the rat hypophysis.

Methods for the isolation and thin-layer chromatographic identification of amino-terminal tryptophyl-peptides presumably responsible for histochemical tryptophyl-peptide reactions in the ACTH cells of the rat hypophysis are described. In the hypophyseal extract several tryptophylpeptide bands--depending on the homogenization solution--were demonstrated on thin-layer chromatograms. Tryptophyl-peptides were demonstrated from their fluorescence induced 1) with glyoxylic acid (glyoxylic acid introduced into the homogenization solution), 2) by exposure of the chromatographic plates to combined formaldehyde and chloral vapour or 3) by exposure to combined formaldehyde and acetyl chloride vapour. A positive PAS reaction was demonstrated in some tryptophyl-peptide bands. Thus, some tryptophylpeptides seem to contribute to the observed PAS positivity of the ACTH cells.

Animals

Characterization of a common precursor to corticotropin and beta-lipotropin: cell-free synthesis of the precursor and identification of corticotropin peptides in the molecule.

mRNA was isolated from cultures of AtT-20/D-16v tumor cells and translated in a mRNA-dependent reticulocyte cell-free system. The corticotropin (ACTH) product was purified by a double-antibody immunoprecipitation procedure using antisera specific for the alpha(1-24) sequence of ACTH. The product is shown by sodium dodecyl sulfate/gel electrophoresis and gel filtration on guanidine-HCl columns to be homogeneous with an apparent molecular weight (Mr) of 28,500. A product with the same molecular weight is synthesized when membrane-bound polysomes from D-16v cells are allowed to complete their nascent chains in a reticulocyte cell-free system. Mr 31,000 ACTH isolated from tumor cells has been separated into three proteins of different apparent Mr:29,000, 32,000, and 34,000. The cell-free product contains the same lysine-, methionine-, and phenylalanine-labeled tryptic peptides as the Mr 29,000 ACTH synthesized in the tumor cells. Tryptic peptide analysis also reveals the presence of the alpha(1-39) sequence in the Mr 28,500 cell-free product and suggests that there is only one copy of this sequence in the Mr 28,500 molecule.

Adrenocorticotropic Hormone

Coexistence of desmin and the fibroblastic intermediate filament subunit in muscle and nonmuscle cells: identification and comparative peptide analysis.

Extraction of chicken embryo fibroblasts (CEF) or baby hamster kidney (BHK) cells with 1% Triton X-100 and 0.6 M KCl leaves an insoluble cytoskeletal residue composed primarily of the 52,000 Mr subunit of intermediate filaments (F-IFP). In addition, CEF cytoskeletons exhibit a minor component with Mr of 50,000, identified as alpha-desmin, one of the two major isoelectric variants of the intermediate filament subunit from smooth muscle. BHK cytoskeletons contain the 50,000 Mr mammalian desmin variant. Cytoskeletons prepared from chicken embryonic myotubes contain F-IFP and both alpha- and beta-desmin. These data suggest that two distinct 10-nm filament subunits coexist in a single cell. One-dimensional peptide analysis of F-IFP and desmin from avian and mammalian cells reveals significant interspecies homology, as well as homology between F-IFP and desmin from the same species. Peptide analyses of 32P-labeled intermediate filament subunits suggest that there is considerable similarity in the phosphorylation sites of these proteins. These results indicate that F-IFP and desmin might be evolutionally related.

Amino Acid Sequence

Characterization of a common precursor to corticotropin and beta-lipotropin: identification of beta-lipotropin peptides and their arrangement relative to corticotropin in the precursor synthesized in a cell-free system.

Radioactive proteins synthesized in an mRNA-dependent reticulocyte cell-free system under the direction of mRNA from AtT-20/D-16v mouse cells were isolated by specific immunoprecipitation using antiserum to either alpha(1-24) corticotropin or beta-endorphin [beta(61-91) lipotropin]. Each immunoprecipitate was fractionated by sodium dodecyl sulfate/polyacrylamide gel electrophoresis and shown to contain only one labeled protein with an apparent molecular weight of 28,500. Tryptic peptide analysis of the Mr 28,500 corticotropin and beta-lipotropin molecules isolated from the gels demonstrated that the two proteins had the same lysine, methionine, and tryptophan peptides. Four tryptic peptides from the cell-free product exhibited the same electrophoretic and chromatographic mobilities as marker tryptic peptides from bovine beta-melanotropin and porcine beta-endorphin. The identification of these peptides was confirmed by amino acid composition studies with a variety of labeled amino acids. The beta-lipotropin tryptic peptides were also shown to be located carboxy terminal to the corticotropin tryptic peptides.

Adrenocorticotropic Hormone