PubMed HealthSearch

SEARCH · PubMed Health

Results for “proteomics database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

NovoBoard: A Comprehensive Framework for Evaluating the False Discovery Rate and Accuracy of De Novo Peptide Sequencing.

De novo peptide sequencing is one of the most fundamental research areas in mass spectrometry-based proteomics. Many methods have often been evaluated using a couple of simple metrics that do not fully reflect their overall performance. Moreover, there has not been an established method to estimate the false discovery rate (FDR) of de novo peptide-spectrum matches. Here we propose NovoBoard, a comprehensive framework to evaluate the performance of de novo peptide-sequencing methods. The framework consists of diverse benchmark datasets (including tryptic, nontryptic, immunopeptidomics, and different species) and a standard set of accuracy metrics to evaluate the fragment ions, amino acids, and peptides of the de novo results. More importantly, a new approach is designed to evaluate de novo peptide-sequencing methods on target-decoy spectra and to estimate and validate their FDRs. Our FDR estimation provides valuable information to assess the reliability of new peptides identified by de novo sequencing tools, especially when no ground-truth information is available to evaluate their accuracy. The FDR estimation can also be used to evaluate the capability of de novo peptide sequencing tools to distinguish between de novo peptide-spectrum matches and random matches. Our results thoroughly reveal the strengths and weaknesses of different de novo peptide-sequencing methods and how their performances depend on specific applications and the types of data.

Peptides

Omics in hereditary optic neuropathies: A systematic review of clinical studies with an integrated point of view.

Hereditary optic neuropathies are characterized by bilateral visual loss due to the degeneration of retinal ganglion cells, resulting in optic nerve degeneration and atrophy. Although the genetic origin of the main isolated and syndromic hereditary optic neuropathies has been characterized, the clinical phenotypes exhibit significant and poorly understood variability in both penetrance and expressivity. Additionally, the genetic and environmental factors that influence the onset of these optic neuropathies remain poorly understood, with limited biomarkers to predict disease progression or as readouts for therapeutic trials. Data-driven omics strategies allow deep phenotyping to improve our understanding of pathophysiological mechanisms and to search for new biomarkers and therapeutic targets. We explore whether the omics strategies applied to patients with hereditary optic neuropathies have provided such new insights. MEDLINE, Web of Science and EMBASE databases were screened for studies with terms relating to hereditary optic neuropathies, transcriptomics, epigenomics, proteomics, metabolomics and lipidomics in clinical studies exploring patients' samples. Out of 1244 references identified, 22 articles were included after double-masked data curation. These articles focused only on the 3 main forms of hereditary optic neuropathies, namely, OPA1-related dominant optic atrophy (n = 4), Leber hereditary optic neuropathy (n = 13), and Wolfram syndrome (n = 5). While the methodological designs and results of these studies were highly heterogeneous, they revealed molecular alterations that we have attempted to discuss at the integrated multi-omics level. This data integration highlighted several common pathophysiological mechanisms such as energetic impairment, endoplasmic reticulum stress, proteotoxic and oxidative stresses, lipid remodeling and altered amino acid and purine metabolisms, while suggesting potential new biomarkers and therapeutic targets. These findings underscore the potential of integrated multi-omics approaches to deepen our understanding of the phenotypic complexity of hereditary optic neuropathies and to support the development of innovative diagnostic and therapeutic strategies.

Humans

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

Influence of nicotine on protein expression around hydrophilic osseointegrated implants: A proteomic study in male rats.

OBJECTIVE: To ensure the success of dental implant treatment, various factors must be considered, including osseointegration and systemic conditions. There is evidence in the literature that smokers may exhibit alterations in tissue healing, which can compromise the success of implant rehabilitation. Therefore, this study aimed to investigate the influence of nicotine on the protein profile of bone tissue around hydrophilic implants during the osseointegration process in rats. DESIGN: Bone tissue samples from the control and nicotine groups (n = 3 per group) were subjected to protein extraction, mass spectrometry, and bioinformatic analyses. Protein identification was performed using Proteome Discoverer 2.1 software and the SEQUEST algorithm, and the protein data were compared with those of a protein database of Rattus norvegicus obtained from UniProt. RESULTS: A total of 740 proteins were detected in both the control group and the nicotine-exposed group. Among them, the proteins biglycan, periostin and histone H4 were highlighted because of their higher abundance in the healthy implant group, while they were reduced in the nicotine-exposed group. CONCLUSIONS: Nicotine has the potential to alter the protein profile of bone tissue around hydrophilic implants during osseointegration, which may impair tissue remodeling and healing.

Animals

Risk of Cardiovascular Disease Mortality in Patients With Diagnosed Cancer and Associated Genetic and Proteomic Mechanisms: A UK Biobank-Based Cohort Study.

BACKGROUND: Previous studies have identified a link between cancer and cardiovascular disease; however, the underlying genetic and proteomic mechanisms remain unclear. Therefore, this study aimed to investigate the association between cancer diagnosis and cardiovascular mortality and to explore the potential mechanisms involved. METHODS: A total of 379 944 participants without cardiovascular disease at baseline, including 65 047 individuals with cancer, were recruited from the UK Biobank database. The primary end point was cardiovascular death. Multivariate Cox regression was performed to evaluate the risk of cardiovascular death in populations with and without cancer. Genome-wide association studies, phenome-wide association studies, and proteomic analyses were applied to investigate the underlying genetic and proteomic mechanisms. RESULTS: Multivariate Cox regression analysis showed an increased risk of cardiovascular death in the group with cancer (hazard ratio, 1.50 [95% CI, 1.40-1.61]) after multivariable adjustment. Proteomic analysis confirmed a strong association between cancer and cardiovascular disease, primarily involving pathways related to complement and coagulation cascades, and various inflammatory processes. In contrast, genome-wide association studies and phenome-wide association studies revealed only a limited number of shared genetic variations between cancer and cardiovascular conditions, such as hypertension and cardiac dysrhythmias. CONCLUSIONS: Cardiovascular risk is increased in patients with cancer and may be related to altered expression of inflammation- and coagulation-related proteins. In clinical practice, it is recommended to emphasize the management of endocrine, kidney, and inflammation-related risk factors in the population with cancer.

Humans

Deciphering the ghost proteome in ovarian cancer cells by deep proteogenomic characterization.

Proteogenomics is becoming a powerful tool in personalized medicine by linking genomics, transcriptomics and mass spectrometry (MS)-based proteomics. Due to increasing evidence of alternative open reading frame-encoded proteins (AltProts), proteogenomics has a high potential to unravel the characteristics, variants, expression levels of the alternative proteome, in addition to already annotated proteins (RefProts). To obtain a broader view of the proteome of ovarian cancer cells compared to ovarian epithelial cells, cell-specific total RNA-sequencing profiles and customized protein databases were generated. In total, 128 RefProts and 30 AltProts were identified exclusively in SKOV-3 and PEO-4 cells. Among them, an AltProt variant of IP_715944, translated from DHX8, was found mutated (p.Leu44Pro). We show high variation in protein expression levels of RefProts and AltProts in different subcellular compartments. The presence of 117 RefProt and two AltProt variants was described, along with their possible implications in the different physiological/pathological characteristics. To identify the possible involvement of AltProts in cellular processes, cross-linking-MS (XL-MS) was performed in each cell line to identify AltProt-RefProt interactions. This approach revealed an interaction between POLD3 and the AltProt IP_183088, which after molecular docking, was placed between POLD3-POLD2 binding sites, highlighting its possibility of the involvement in DNA replication and repair.

Humans

BRIDGE: an interactive application for multi-omics data analysis, visualization and integration.

SUMMARY: BRIDGE is a Shiny-based application that provides an accessible, modular platform for individual and integrative multi-omics analysis. Using an independent SQLite database backend, it offers a local, private, and user-friendly environment that requires no prior computational expertise. The application supports proteomics, phospho-proteomics, and RNA-seq analyses through a comprehensive suite of visualization and analytical modules, together with an integrated multi-omics analysis pipeline. Built-in caching and asynchronous processing improve responsiveness, enabling efficient exploration, analysis, and visualization of multi-omics datasets on moderate hardware. AVAILABILITY AND IMPLEMENTATION: BRIDGE is implemented in R using Shiny and is freely available as a Docker container at https://ghcr.io/paulilab/bridge. A public demonstration server with example datasets is available at https://bridge.imp.ac.at. Code and datasets are also available at https://github.com/paulilab/BRIDGE and under DOI: https://doi.org/10.5281/zenodo.20215824.

Multiomics

Chromosomal toxin-antitoxin systems in Pseudomonas putida are rather selfish than beneficial.

Chromosomal toxin-antitoxin (TA) systems are widespread genetic elements among bacteria, yet, despite extensive studies in the last decade, their biological importance remains ambivalent. The ability of TA-encoded toxins to affect stress tolerance when overexpressed supports the hypothesis of TA systems being associated with stress adaptation. However, the deletion of TA genes has usually no effects on stress tolerance, supporting the selfish elements hypothesis. Here, we aimed to evaluate the cost and benefits of chromosomal TA systems to Pseudomonas putida. We show that multiple TA systems do not confer fitness benefits to this bacterium as deletion of 13 TA loci does not influence stress tolerance, persistence or biofilm formation. Our results instead show that TA loci are costly and decrease the competitive fitness of P. putida. Still, the cost of multiple TA systems is low and detectable in certain conditions only. Construction of antitoxin deletion strains showed that only five TA systems code for toxic proteins, while other TA loci have evolved towards reduced toxicity and encode non-toxic or moderately potent proteins. Analysis of P. putida TA systems' homologs among fully sequenced Pseudomonads suggests that the TA loci have been subjected to purifying selection and that TA systems spread among bacteria by horizontal gene transfer.

Anti-Bacterial Agents

Identification of Glioblastoma Cell Surface Proteins and Assessment of Their Expression Across Patient-Derived Stem-Like Cell Cultures.

Glioblastoma (GBM) is the most common primary brain cancer in adults and remains fatal, with a median survival of a few months. There is an urgent need to develop novel therapeutic strategies against this aggressive malignancy. Modern cancer research increasingly focuses on personalized therapies tailored toward unique molecular features of each tumor or patient. In this context, cell surface proteins (CSPs) represent an attractive class of therapeutic targets due to their accessibility and central roles in physiological and pathological processes, making them among the most targeted proteins in current drug development. In this study, promising CSPs were identified through an untargeted proteomics approach using high-resolution mass spectrometry on patient-derived GBM stem-like cell (GSC) cultures, complemented by RNA-seq data and computational database analyses. From this primary discovery, five CSPs, namely PTK7, PTPRZ1, OSMR, CSPG4, and IGDCC4, were selected for detailed investigation. A targeted UHPLC-multiple reaction monitoring (MRM) method was developed and optimized to assess their expression and evaluate their abundance variations across different GSC cultures and cell passage levels. Beyond confirming these CSPs as potential therapeutic targets in GBM, our study demonstrates the value of three-dimensional GSC cultures as robust models for biomarker research and target assessment.

Humans

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals

Integrative pan-cancer analysis of transferrin reveals context-dependent prognostic associations and links to immune and metabolic disease-related programs.

BACKGROUND: Iron metabolism is closely linked to tumor biology, yet the pan-cancer significance of transferrin (TF), the major circulating iron-transport protein, remains insufficiently defined. Although TF has been implicated in cancer-related processes, its prognostic relevance, immune associations, and broader disease-related transcriptional context have not been systematically characterized across tumor types. OBJECTIVE: This study aimed to perform an integrative pan-cancer analysis of TF to characterize its expression patterns, clinical associations, immune context, pathway features, and pharmacogenomic correlations, and to explore whether TF-related signals extend to selected metabolic and chronic organ injury settings. METHODS: We used multiple public databases, including The Cancer Genome Atlas (TCGA), Human Protein Atlas (HPA), Gene Expression Omnibus (GEO), and Cancer Cell Line Encyclopedia (CCLE), to integrate transcriptomic, proteomic, and clinical data across 33 tumor types and selected non-malignant conditions. TF expression was evaluated across normal tissues, tumors, and cell lines, followed by survival analysis, immune infiltration analysis, TMB/MSI and methylation assessment, pathway enrichment, and drug-response correlation. Independent GEO cohorts of non-alcoholic steatohepatitis (NASH), heart failure (HF), and liver cirrhosis (LC) were used for cross-disease extension. Selected findings were further explored in OA/PA-treated hepatocytes, 786-O renal carcinoma cells, and AC16 cardiomyocytes. RESULTS: TF showed pronounced tissue specificity and cancer-type-dependent dysregulation. Across pan-cancer cohorts, the most consistent adverse survival associations were observed in kidney renal clear cell carcinoma (KIRC) and stomach adenocarcinoma (STAD), where TF remained associated with overall survival (OS) in multivariable analyses. TF expression was also correlated with cancer-type-specific immune infiltration patterns and selected drug-response profiles. Across independent NASH, HF, and LC datasets, TF expression was elevated and TF-associated pathways partially overlapped with those observed in cancer. In vitro experiments provided preliminary support that TF modulation is associated with proliferative phenotypes in KIRC cells and stress- and metabolism-related phenotypes in hepatocyte and cardiomyocyte models. CONCLUSION: These findings support TF as a context-dependent biomarker candidate in cancer, with the most consistent prognostic relevance observed in KIRC and STAD. Rather than establishing a unified mechanism across diseases, this study provides an integrative framework suggesting that TF is associated with malignant behavior, immune context, and selected metabolic stress-related programs, and warrants further mechanistic investigation.

Iron metabolism

Verification of biological markers of subacute cutaneous lupus erythematosus via TMT labelling proteomics combined with transcriptome data.

OBJECTIVE: This study aimed to investigate biological markers in subacute cutaneous lupus erythematosus (SCLE). METHODS: The tandem mass tag (TMT)-labelling proteomics method was used to explore differentially expressed proteins between SCLE lesions and normal skin tissues. The differences in transcriptomic data between SCLE tissues and normal skin tissues were analysed from the GEO database (GSE81071, GSE109248 and GSE112943). The differences in transcriptomic data from peripheral blood mononuclear cells (PBMCs) of patients with systemic lupus erythematosus (SLE) and normal controls were analysed (GSE81622 and GSE154851). The 35 healthy controls, 30 SCLE patients, 35 SLE patients and 30 lupus nephritis (LN) patients were diagnosed and enrolled. The serum expression levels of IFI44 and EPSTI1 were detected. Data were presented as the mean&#xa0;&#xb1;&#xa0;standard deviation or frequency and were analysed using Student's t-test, Chi-square test and one-way ANOVA between the groups. Receiver operating characteristic (ROC) curves were used to analyse the clinical efficacy of IFI44 and EPSTI1 in distinguishing SCLE from SLE. RESULTS: In a comparative analysis of SCLE lesions and normal skin tissues, proteomics studies identified 376 proteins that exhibited significant differential expression. In GO and KEGG analyses, the enriched terms mainly included the interferon-gamma-mediated signalling pathway (p&#xa0;<&#xa0;.001), immune receptor activity (p&#xa0;<&#xa0;.001) and cell adhesion molecules (p&#xa0;<&#xa0;.001). The top 10 hub genes were screened in SCLE as follows: CD8A, CXCL10, IFI44, CD7, CCL5, TLR4, EPSTI1, ISG15, KLRD1 and SELL using Cytoscape (3.10.1) software. The 15 common proteins/genes between proteomics and three datasets results were found, including CXCL10, OAS1, DDX60L, CFB, IFI6, HERC6, IFI44L, GBP1, EPSTI1, OAS2, CXCL11, TYMP, IFI44, ISG15 and IFIT3. The 61 differentially expressed genes in GSE81622 and the top 100 differentially expressed genes in GSE154851, alongside the 15 identified genes described above through Venn diagram analysis. Four common genes, IFI44L, IFI44, EPSTI1 and OAS1, were identified. Two common genes, IFI44 and EPSTI1, were found in hub genes from the proteomics results. The serum levels of IFI44 and EPSTI1 in LN were significantly higher than those in SLE patients (p&#xa0;<&#xa0;.05). ROC curve analysis demonstrated that serum levels of IFI44 and EPSTI1 could differentiate SCLE from SLE with an area under the curve (AUC) of 0.898 and 0.847, respectively. CONCLUSIONS: The IFI44 and EPSTI1 proved to be closely involved in the progression from SCLE to SLE, and can represent new candidate diagnostic molecular markers of occurrence and progression of SCLE.

Humans

Harnessing deep learning for proteome-scale detection of amyloid signaling motifs.

MOTIVATION: Amyloid signaling sequences adopt the cross-&#x3b2; fold that is capable of self-replication in the templating process. Propagation of the amyloid fold from the receptor to the effector protein is used for signal transduction in the immune response pathways in animals, fungi, and bacteria. So far, a dozen of families of amyloid signaling motifs (ASMs) have been classified. Unfortunately, due to the wide variety of ASMs it is difficult to identify them in large protein databases available, which limits the possibility of conducting experimental studies. To date, various deep learning (DL) models have been applied across a range of protein-related tasks, including domain family classification and the prediction of protein structure and protein-protein interactions. RESULTS: In this study, we develop tailor-made bidirectional LSTM and BERT-based architectures to model ASM, and compare their performance against a state-of-the-art machine learning grammatical model. Our research is focused on developing a discriminative model of generalized ASMs, capable of detecting ASMs in large datasets. The DL-based models are trained on a diverse set of motif families and a global negative set, and used to identify ASMs from remotely related families. We analyze how both models represent the data and demonstrate that the DL-based approaches effectively detect ASMs, including novel motifs, even at the genome scale. AVAILABILITY AND IMPLEMENTATION: The models are provided as a Python package, asmscan-bilstm, and a Docker image at https://github.com/chrispysz/asmscan-proteinbert-run. The source code can be accessed at https://github.com/jakub-galazka/asmscan-bilstm and https://github.com/chrispysz/asmscan-proteinbert. Data and results are at https://github.com/wdyrka-pwr/ASMscan.

Deep Learning

Proteomic insights into Helicobacter pylori infection in stomach cells, revealing host response and host-targeted therapeutics repurposing.

BACKGROUND: Helicobacter pylori (H. pylori) is a globally prevalent gastric pathogen strongly associated with chronic gastritis, peptic ulcers, and gastric cancer. While bacterial factors have been extensively studied, host proteomic responses and their therapeutic potential remain largely underexplored. RESEARCH DESIGN AND METHODS: Current analyses employed a systematic proteomics-based data integration and harmonization approach (retrospective qualitative cohort study) to identify important differentially regulated host proteins. Proteomic datasets were curated from in vitro studies and analyzed for functional enrichment, protein-protein interaction networks, and hub protein identification. To explore therapeutic repurposing, drug repositioning was performed using the DrugBank database. RESULTS: Data summation describing protein differential regulation in human gastric cells as a result of the infection revealed 1672 perturbed host proteins. Bioinformatics analysis revealed 11 proteins including CSK, MET, RELA, MARK2, GRB2, FTO, PLCG1, CRKL, RPS5, RPS9, and RPS27A to be ideal host targets for therapeutic repurposing. Clinically approved drugs such as Dasatinib (targeting CSK) and Crizotinib (targeting MET) emerged as promising candidates due to favorable pharmacokinetics and known bioactivity. CONCLUSIONS: Host-directed therapeutics could offer alternative strategies to conventional antibiotic therapy, addressing challenges such as resistance and infection recurrence, providing a foundation for future experimental validation and development of host-targeted interventions for infection control.

Humans