PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “proteoform characterization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

4 recordsLinked to original sources

Mass spectrometry-based top-down proteomics for proteoform profiling of protein coronas.

The protein corona is a layer of biomolecules-primarily proteins-that adsorbs to nanoparticle (NP) surfaces in biological fluids. If the purpose of the NP is therapeutic, this can have a profound effect on its biological activity and function in vivo. Protein corona formation can also be exploited for diagnostic purposes and to differentially enrich proteins for biomarker discovery. For all of these applications, it is useful to determine which proteins, and which specific proteoforms, bind to different types of NP. The traditional mass spectrometry (MS)-based bottom-up proteomics does not accurately identify specific proteoforms within the protein corona. This limitation impedes the nanomedicine field's ability to precisely predict the biological fate and pharmacokinetics of nanomedicines and their effectiveness in early-stage biomarker discovery and disease detection because many different proteoforms of the same gene could exist in the corona, and they have divergent biological functions. Here, we describe how to use capillary zone electrophoresis (CZE)-MS-based top-down proteomics to characterize the proteoform landscape of the protein corona. Our procedures detail the recovery of intact proteoforms from NP surfaces by using detergent-assisted proteoform elution and the measurement of these proteoforms by using CZE-tandem MS (MS/MS) and CZE-high-field asymmetric waveform ion mobility spectrometry (FAIMS)-MS/MS. The entire workflow is completed within 3-4 d. Using this protocol, hundreds of proteoforms from the protein corona of polystyrene NPs can be identified. Distinct protein corona proteoform profiles were observed from NPs with different physicochemical properties. The addition of FAIMS is beneficial for more in-depth proteoform characterization.

Proteomics↗

A Draft Map of E. coli Proteoforms.

Top-down proteomics (TDP) enables direct characterization of intact proteoforms, providing protein-level insights into molecular diversity arising from post-translational modifications and sequence variations. Despite this advantage, proteome coverage in TDP remains limited relative to bottom-up proteomics (BUP). To expand coverage, we developed an integrated multidimensional approach combining sequential protein extraction, size-exclusion chromatography (SEC) fractionation, and capillary zone electrophoresis (CZE)-tandem mass spectrometry (MS/MS) and reversed-phase liquid chromatography (RPLC)-MS/MS. This approach identified 743 proteoform families and 10,613 proteoforms from E. coli cells through hundreds of MS runs. By incorporating previous E. coli TDP data sets from our group, we identified 14,932 proteoforms from 985 proteoform families, covering 43% of the E. coli proteome. The data represent the highest proteome coverage of cells by MS-based TDP, creating a draft map of E. coli proteoforms. The results offer strong evidence that MS-based TDP can reach high proteome coverage.

Escherichia coli↗

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans↗

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics↗