PubMed HealthSearch

SEARCH · PubMed Health

Results for “proteomics database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multiple urinary peptides are associated with hypertension: a link to molecular pathophysiology.

OBJECTIVES: Hypertension is a common condition worldwide; however, its underlying mechanisms remain largely unknown. This study aimed to identify urinary peptides associated with hypertension to further explore the relevant molecular pathophysiology. METHODS: Peptidome data from 2876 individuals without end-organ damage were retrieved from the Human Urinary Proteome Database, belonging to general population (discovery) or type 2 diabetic (validation) cohorts. Participants were divided based on systolic blood pressure (SBP) and diastolic BP (DBP) into hypertensive (SBP &#x2265;140&#x200a;mmHg and/or DBP &#x2265;90&#x200a;mmHg) and normotensive (SBP <120&#x200a;mmHg and DBP <80&#x200a;mmHg, without antihypertensive treatment) groups. Differences in peptide abundance between the two groups were confirmed using an external cohort ( n &#x200a;=&#x200a;420) of participants without end-organ damage, matched for age, BMI, eGFR, sex, and the presence of diabetes. Furthermore, the association of the peptides with BP as a continuous variable was investigated. The findings were compared with peptide biomarkers of chronic diseases and bioinformatic analyses were conducted to highlight the underlying molecular mechanisms. RESULTS: Between hypertensive and normotensive individuals, 96 (mostly COL1A1 and COL3A1) peptides were found to be significantly different in both the discovery (adjusted) and validation (nominal significance) cohorts, with consistent regulation. Of these, 83 were consistently regulated in the matched cohort. A weak, yet significant, association between their abundance and standardized BP was also observed. CONCLUSION: Hypertension is associated with an altered urinary peptide profile with evident differential regulation of collagen-derived peptides. Peptides related to vascular calcification and sodium regulation were also affected. Whether these modifications reflect the pathophysiology of hypertension and/or early subclinical organ damage requires further investigation.

Humans

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49&#x2009;h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases

Advancing proteomic discovery through optimized multi-stage scoring and deep learning-enhanced open search.

MOTIVATION: Protein search engines are essential for interpreting mass spectrometry data into biological insight. Current tools often face limitations in sensitivity when analyzing complex modern datasets, and lack a unified framework that effectively integrates deep learning features for both restricted and open searches, especially for scenarios aimed at discovering unknown modifications. RESULTS: We present pFind+, a high-performance search engine for data-dependent acquisition (DDA) proteomics, extending pFind. It introduces an enhanced raw scoring that delivers substantially improved pre-filtering ability, while recovering most of the computational overhead through a tailored acceleration strategy. Coupled with an enhanced rescoring framework that effectively integrates deep learning features, pFind+ uniquely supports high-sensitivity, DL-enhanced open search, enabling comprehensive PTM discovery while incorporating hardware-aware inference optimizations for practical deployment. Evaluations across diverse datasets demonstrate its superior sensitivity, with gains of 12.7%-29.3% (average 17.9%) in restricted search and 8.0%-38.4% (average 25.8%) in open search over the best existing tools.

Deep Learning

The Proteomic Landscape of CTNNB1 Mutated Low-Grade Early-Stage Endometrial Carcinomas.

Endometrial carcinoma is the most frequent gynecologic malignancy in western countries. In recent years, mutations in CTNNB1 have been associated with worse prognosis in low-risk carcinomas. However, there is a lack of understanding of the proteomic implications of CTNNB1 mutations in this type of tumor. In this study, we performed shotgun proteomics using Formalin-Fixed Paraffin-Embedded (FFPE) tissue samples of CTNNB1 mutated and wild-type low-risk endometrial carcinomas. A publicly available proteomic and transcriptomic database was used to validate results. Differential protein expression and Gene Set Enrichment Analysis revealed dysregulation of pathways associated with cell keratinization, immune response modulation, and intracellular calcium regulation. CTNNB1 mutated tumors showed immune dysregulation at multiple levels including cytokine secretion, cell adhesion, and lymphocyte activation. These results were supported by tissue multiplex immunofluorescence analysis, demonstrating reduced CD8 tumor-infiltrating lymphocytes and different immune spatial interaction patterns. Intracellular calcium dysfunction was associated with key transcript dysregulation. We found an increased expression of CAMK2A and ROR2, suggesting a potential role for non-canonical Wnt pathway activation in CTNNB1 mutated tumors.

Humans

Heme oxygenase 1 (HO-1) is a drug target for reversing cisplatin resistance in non-small cell lung cancer.

INTRODUCTION: Platinum-based drugs, the most widely used chemotherapeutic drugs in clinical oncology, have long faced the problem of drug resistance, which is urgently in need of resolution. Identifying biomarkers of drug resistance may help reduce platinum resistance and improve therapeutic efficacy. OBJECTIVES: This study aims to identify potential biomarkers associated with the development of cisplatin resistance in non-small cell lung cancer (NSCLC) and explore mechanisms to overcome chemoresistance. METHODS: NSCLC cisplatin resistance cell lines were constructed, and transcriptome sequencing was performed. Results were validated using Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA) databases. Molecular docking, proteomics sequencing, and in vitro and in vivo experiments were conducted to evaluate the role of Heme Oxygenase 1 (HO-1) in cisplatin resistance. RESULTS: NSCLC cisplatin resistance cell lines, GEO and TCGA data identified HMOX1, downstream of Nrf2, as a key drug resistance gene induced by cisplatin. Activation of the Nrf2/HO-1 pathway was found to induce ferroptosis resistance, a critical mechanism of cisplatin resistance. Candidate compounds SB 202190 and Nordihydroguaiaretic acid (NDGA) effectively reactivated ferroptosis by inhibiting HO-1, thereby increasing cisplatin sensitivity. CONCLUSION: The Nrf2/HO-1 pathway is a significant contributor to cisplatin resistance in NSCLC. Targeting HO-1 with SB 202190 and NDGA presents a promising strategy to overcome resistance and improve chemotherapy outcomes.

Cisplatin

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Comprehensive Assessment of the Intrinsic Pancreatic Microbiome.

OBJECTIVE: To sought comprehensively profile tissue and cyst fluid in patients with benign, precancerous, and cancerous conditions of the pancreas to characterize the intrinsic pancreatic microbiome. BACKGROUND: Small studies in pancreatic ductal adenocarcinoma (PDAC) and intraductal papillary mucinous neoplasm (IPMN) have suggested that intrapancreatic microbial dysbiosis may drive malignant transformation. METHODS: Pancreatic samples were collected at the time of resection from 109 patients. Samples included tumor tissue (control, n = 20; IPMN, n = 20; PDAC, n = 19) and pancreatic cyst fluid (IPMN, n = 30; serous cystadenomas, n = 10; mucinous cystic neoplasm, n = 10). Assessment of bacterial DNA by quantitative polymerase chain reaction and 16S ribosomal RNA gene sequencing was performed. Downstream analyses determined the relative abundances of individual taxa between groups and compared intergroup diversity. Whole-genome sequencing data from 140 patients with PDAC in the National Cancer Institute's Clinical Proteomic Tumor Analysis Consortium were analyzed to validate findings. RESULTS: Sequencing of pancreatic tissue yielded few microbial reads regardless of diagnosis, and analysis of pancreatic tissue showed no difference in the abundance and composition of bacterial taxa between normal pancreas, IPMN, or PDAC groups. Low-grade and high-grade dysplasia IPMN were characterized by low bacterial abundances with no difference in tissue composition and a slight increase in Pseudomonas and Sediminibacterium in high-grade dysplasia cyst fluid. Decontamination analysis using the Clinical Proteomic Tumor Analysis Consortium database confirmed a low-biomass, low-diversity intrinsic pancreatic microbiome that did not differ by pathology. CONCLUSIONS: Our analysis of the pancreatic microbiome demonstrated very low intrinsic biomass that is relatively conserved across diverse neoplastic conditions and thus unlikely to drive malignant transformation.

Humans

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics

DORSSAA: Drug-Target interactOmics Resource Based on Stability/Solubility Alteration Assay.

Advancements in high-throughput techniques such as Thermal Proteome Profiling and the high-throughput Proteome Integral Solubility Alteration assay have revolutionized our understanding of drug-protein interactions. Despite these innovations, the absence of an integrative platform for cross-study analysis of stability and solubility alteration data represents a significant bottleneck. To address this gap, we introduce Drug-target interactOmics Resource based on Stability/Solubility Alteration Assay (DORSSAA), an interactive and expandable web-based platform for the systematic analysis and visualization of proteome stability and solubility alteration assay datasets. Currently, DORSSAA features 1,135,985 records spanning 38 cell lines and organisms, 135 compounds, and 40,742 protein targets. Through its user-friendly interface, the resource supports comparative drug-protein interaction analysis and facilitates the discovery of actionable therapeutic targets. Through two case studies, methotrexate target profiling in A549 cells and combinatorial-therapy drug-target interactions in leukemia cell lines, we demonstrate DORSSAA's utility for identifying protein-drug interactions across diverse experimental contexts. This resource empowers researchers to accelerate drug discovery and enhance our understanding of protein behavior. Compared with data repositories and interaction databases, DORSSAA provides direct protein-level evidence of mechanisms of action with strict statistical control for each study. This enables more reliable identification of drug targets, off-target effects, and potential drug combinations.

Humans

Cilia.Pro database of ciliary proteins from vertebrates, Chlamydomonas, and Caenorhabditis.

Cilia and flagella are microtubule-based organelles that generate force and sense the extracellular environment. In humans, these structures are essential for development, homeostasis, and reproduction, with defects contributing to a wide array of congenital and degenerative disorders. As cilia were present on the last common ancestor of all eukaryotes, research on cilia across model organisms holds significant relevance for understanding human disease. The green alga Chlamydomonas, which diverged from the human lineage with the animal-plant split, shares striking similarities in ciliary structure and function with humans. Two decades ago, our group published the proteome of the Chlamydomonas cilium, identifying hundreds of new ciliary proteins that were organized in an online database. Since then, advances have brought us a more comprehensive understanding of both Chlamydomonas and mammalian cilia. Our database, www.Cilia.Pro, has been continually updated to integrate proteomic, transcriptomic, and genomic data from Chlamydomonas and Caenorhabditis along with humans, and other vertebrates providing a valuable tool for the ciliary research community.

Cilia

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics