PubMed HealthSearch

PubMed · 42695892

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Abstract

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Diya Mathai, Stefan Schulze. 2026-09-04. ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.. https://doi.org/10.1021/acs.jproteome.6c00074

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The application of the CRISPR-Cas system in Pseudomonas aeruginosa infections.

Due to the extensive drug resistance of Pseudomonas aeruginosa (P. aeruginosa), it is still a great clinical challenge. The clustered regularly interspaced short palindromic repeats and associated proteins (CRISPR-Cas) system has become a promising strategy against this pathogen. This review critically evaluates the multifaceted applications of CRISPR-Cas technology in P. aeruginosa, including its role in antimicrobial resistance, diagnostics, genome editing, and emerging therapeutic and vaccine strategies. In addition to conducting a comprehensive analysis of various studies, we also compared the performance and limitations of various CRISPR platforms, and discussed the main technologies and transformation obstacles in this field. Finally, we look forward to the direction of applying these experimental tools to clinical research in the future.

Pseudomonas aeruginosa

Comparative genomics reveals genotype-phenotype concordance and cryptic resistomes in clinical Pseudomonas aeruginosa.

BACKGROUND: Pseudomonas aeruginosa (P. aeruginosa) is a major pathogen because of its adaptability. It shows rapid evolution of multidrug resistance (MDR). Phenotype-based diagnostics often fail to detect silent resistance determinants and early adaptive changes. This study integrates phenotypic profiling with whole-genome sequencing (WGS) to examine resistance architecture in clinical isolates from eastern India. METHODS: From 1295 culture-positive P. aeruginosa specimens collected at a tertiary care hospital in eastern India. Using predefined criteria, representative MDR and non-MDR isolates were selected, including distinct resistance phenotypes, specimen-source diversity, and hospital and community-acquired settings; multivariate analysis of resistance profiles illustrated phenotypic diversity. Antimicrobial susceptibility assessed using VITEK-2 and Kirby-Bauer disk diffusion, species identity confirmed by 16 S rRNA sequencing, and genomic analysis processed through a reference-guided workflow. Antimicrobial Resistance (AMR) determinants were identified through CARD, and phylogenetic tree constructed from 454 publicly available P. aeruginosa genomes. RESULTS: MDR exhibited greater sequence divergence relative to PA14 (~ 69,000 variants) than the non-MDR isolate (~ 58,700 variants), with > 92% coverage at ≥ 30X depth. Strong genotype-phenotype concordance observed in MDR isolates across five antibiotic classes, associated with β-lactamase variants (PDC-67, OXA-396) and regulatory adaptations (ArmR, cprS). The non-MDR isolate harboured gyrA (T83I) resistance-associated mutations, PDC-1, and OXA-847 without phenotypic expression, indicating silent resistome. Phylogenetically, MDR isolates clustered tightly within the phylogeny, while the non-MDR isolate formed a distinct lineage. CONCLUSION: Observed genomic differences align with adaptation under antimicrobial selection, though confirmation requires larger collections. The non-MDR isolate retained a silent resistome. Findings highlight limitations of phenotype-only diagnostics, support genomic data integration, and emphasize transcriptomics for hidden resistance expression and regulatory dynamics.

Pseudomonas aeruginosa

Transposon insertion sequencing of Pseudomonas aeruginosa identifies multiple intersecting pathways essential for extreme colistin resistance.

Colistin is used to treat antibiotic resistant gram-negative infections, including those caused by Pseudomonas aeruginosa (Pa). Using a diverse collection of clinical isolates, we identified BWH047, a colistin-resistant isolate with an extremely high minimum inhibitory concentration (MIC, 1280 µg/mL). To characterize the genes conditionally essential for colistin resistance in BWH047, we employed transposon insertion sequencing and identified 20 gene candidates. In-frame deletion validated 75% of the candidates and identified genes in several new pathways that contribute to colistin resistance in Pa, including algU and wapH. We also identified several candidate genes from previously reported colistin resistance pathways (e.g., arn, pmrAB). We further investigated the impact of a colistin resistance-associated inner membrane DedA-family undecaprenyl phosphate flippase, which we named DpcA (DedA of Pseudomonas necessary for colistin resistance A). Deletion of dpcA in BWH047 restored sensitivity to colistin (MIC = 0.5 µg/mL) and resulted in several unique changes to the structure of lipopolysaccharide (LPS), including production of decreased amounts of the colistin resistance-conferring 4-amino-4-deoxy-L-arabinose (L-Ara4N) modification on lipid A. This work represents a robust analysis of colistin resistance in Pa and identifies intersecting pathways that contribute to extreme phenotypic resistance.

Pseudomonas aeruginosa