PubMed HealthSearch

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Genomic and Structural Analysis of Gamete Recognition Proteins in a Broadcast Spawning Echinoderm Mesocentrotus franciscanus.

Gamete recognition proteins are expressed on the surfaces of sperm and eggs, where they mediate interactions between gametes. The genetic basis for gamete recognition proteins, as well as their structure and interactions, have yet to be fully resolved. Using a new high-quality de novo genome assembly for the sea urchin Mesocentrotus franciscanus, we investigated the genomic structure, expression, and protein forms of several gamete recognition proteins: sperm bindin, egg receptor for sperm (HSP110), and egg bindin receptor (EBR1), as well as the receptor for egg jelly (REJ) and its paralogs. To inform future population genetic and evolutionary studies, we resolve the genomic structure of the large EBR1 protein, identifying fewer tandem CUB-TSP1 repeats in EBR1 compared to the initial characterization of this protein. As expected for an egg receptor for sperm, EBR1 is highly expressed in female reproductive tissues (eggs and female gonad), compared to other tissues. In contrast, HSP110 shows similar levels of expression across male and female reproductive tissues, as well as across non-reproductive tissues and development stages. HSP110 might be a pleiotropic gene that in part influences fertilization. Using protein structural modeling and functional domain predictions, we propose hypotheses about potential interactions among EBR1, bindin, and HSP110 proteins that may provide insight into sperm-egg interactions in sea urchins. Resolving the genomic structure of genes encoding gamete recognition proteins, in combination with functional annotations and protein structural modeling, enables deeper investigation into the consequences of variation in gamete recognition proteins and the evolution of reproductive isolation.

Mesocentrotus franciscanus

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability. We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system. These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida

Metaproteomic Analysis to Assess the Impact of Storage Media on Human Gut Microbiome in Fecal Samples.

The human gut microbiome is a diverse community of microorganisms residing in the gastrointestinal tract. The storage condition of fecal samples may impact the taxonomic and protein compositions of microbiomes in these samples. Here, we performed a mass spectrometry-based metaproteomic study to assess the impact of storage media on human gut microbiome in fecal samples. We evaluated FDA-authorized OMNIgene·GUT (OG), phosphate-buffered saline (PBS), and RNALater (RNAL) buffers and identified 38,185 microbial peptides corresponding to 7348 microbial proteins, which matched 16 phyla, 20 classes, 50 orders, 104 families, 332 genera, and 453 species. We found a high similarity among the fecal microbiomes preserved in OG, PBS, and RNAL in terms of the identification of proteins, taxa, and functional annotations. Both alpha and beta diversity suggested the high similarity among samples stored in the three media. Nonetheless, we also found some notable differences among buffers regarding the abundances of a few taxon groups. A partial human proteome (over 400 proteins) was identified in the fecal samples, with most of these proteins associated with the membrane and extracellular regions. The findings indicate the similarity among microbiomes in the fecal samples stored in OG, PBS, and RNAL regarding proteome profile, taxa, and functional capacity. SUMMARY: This study thoroughly analyzed and compared the metaproteomes of fecal samples preserved at -80°C in PBS, RNALater, and OMNIgene·GUT Dx buffers, offering novel insights into the effectiveness of these buffers in maintaining the stability and composition of the human gut microbiome. We found a high similarity in the identification and quantification of proteins, taxa, and functional annotations across the three buffers, with notable quantitative differences highlighting subtle yet important variations in preservation efficacy. The unique datasets and findings could offer valuable revelations into the impact of fecal sample preservation on translational and clinical analyses of the human gut microbiome.

Humans

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

Proteomics-Based Identification of the Pyroptosis-Related Biomarker PCSK9 and Its Association With the Pathogenesis of Rheumatoid Arthritis.

Rheumatoid arthritis (RA) is a common autoimmune disease, and early diagnosis is critical for effective treatment. This study aims to identify potential biomarkers related to pyroptosis through serum proteomics analysis, offering new insights for the early diagnosis of RA. We enrolled 100 participants, including 50 patients with RA and 50 healthy controls. Serum samples were collected and analyzed using high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) for proteomics profiling. Differential protein expression analysis and functional annotation revealed significant upregulation of pyroptosis-related proteins in the serum of patients with RA. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, along with protein-protein interaction (PPI) network analysis, showed that these proteins are involved in inflammation and immune pathways, particularly the activation of the NOD-like receptor protein 3 (NLRP3) inflammasome. Enzyme-linked immunosorbent assay (ELISA) validation confirmed a significant increase in PCSK9 levels in patients with RA, suggesting that PCSK9 may play a key role in the pathogenesis of RA. This study provides new directions for biomarker research in RA, particularly regarding the potential involvement of the pyroptosis pathway, with significant clinical application prospects.

Humans

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals

Whole-Genome Analysis of Bacillus Licheniformis Ali5 and Synthesis of Lichenysin via Genome Shuffling.

Whole-genome sequencing of Bacillus licheniformis Ali5 was performed via MGI-seq PE150 and Nanopore single-molecule real-time sequencing. The strain has a 4,114,664 bp circular genome encoding 4030 protein-coding genes. Functional annotation across NR, COG, GO, KEGG, CARD, BacMet, and CAZy databases identified 4025, 2812, 988, 1242, 72, 69, and 94 corresponding genes, respectively, and antiSMASH 6.0 revealed multiple antimicrobial biosynthetic gene clusters, including intact lichenysin and lichenicidin VK21 A1/A2 gene clusters. Three rounds of recursive protoplast fusion-based genome shuffling, paired with a dual-index screening system, significantly improved strain growth and lichenysin biosynthesis. Recombinants exhibited shortened lag phase, enhanced proliferation, improved stationary-phase stability, and higher diauxic peak biomass. PP3-176 and PP3-186 showed 4.6%-8.1% higher 12-h shake-flask titer and 3.1%-4.0% higher maximum titer than the parental average, with excellent fermentation stability. 1-L bioreactor validation confirmed strong scale-up potential. PP3-186 achieved 27.2% and 31.6% titer increases at 12 h and 20 h, while PP3-176 yielded 20.4% and 14.6% improvements with robust metabolic performance. This study validates genome shuffling as an effective strategy for enhancing lichenysin production, providing candidate strains and technical support for industrial application.

Bacillus licheniformis

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

Genomic and transcriptomic characterization of genes expressed at 20 MPa by the marine actinobacterium Kocuria flava.

A marine hydrocarbonoclastic actinobacterium Kocuria flava IOS11 was isolated from 3500 m deep-sea water of the Indian Ocean. The isolate efficiently degraded phenanthrene (250 mg/L) achieving 82 and 98% of degradation at 0.1 MPa and 20 MPa, respectively within a period of 5 days. Whole genome, transcriptomee and metabolomic analysis elucidated its phenanthrene biodegradation efficiency under in situ deep-sea conditions. The genome sequence comprises 3.47 Mb distributed across 88 scaffolds with a high GC content of 74.30%. The genome analysis encoded 3126 genes including 3052 protein coding sequences with functional annotation identifying a broad array of genes associated with PAHs degradation, environmental stress adaptation, biosurfactant and siderophore synthesis. Transcriptome profiling under 0.1 and 20 MPa conditions with phenanthrene as a sole carbon source revealed enhanced expression of hydrocarbon degrading genes, transporters, biosurfactant associated enzymes and stress responsive genes including integrases, DNA repair protein Rad, alanine ligase, heat and cold shock proteins under high pressure conditions underscoring the deep-sea adaptation capabilities of the strain. The degradation pathway of phenanthrene was proposed through integrated genome, transcriptome and metabolomic analysis. These studies provided K. flava IOS11 as a metabolically versatile and pressure adapted bacterium with promising potential for bioremediation application in extreme marine environment.

Transcriptome