PubMed HealthSearch

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC.

AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Humans

Reevaluating human gene annotation: a second-generation analysis of chromosome 22.

We report a second-generation gene annotation of human chromosome 22. Using expressed sequence databases, comparative sequence analysis, and experimental verification, we have extended genes, fused previously fragmented structures, and identified new genes. The total length in exons of annotation was increased by 74% over our previously published annotation and includes 546 protein-coding genes and 234 pseudogenes. Thirty-two potential protein-coding annotations are partial copies of other genes, and may represent duplications on an evolutionary path to change or loss of function. We also identified 31 non-protein-coding transcripts, including 16 possible antisense RNAs. By extrapolation, we estimate the human genome contains 29,000-36,000 protein-coding genes, 21,300 pseudogenes, and 1500 antisense RNAs. We suggest that our revised annotation criteria provide a paradigm for future annotation of the human genome.

Animals

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals

Serine-threonine phosphoregulation by PknB and Stp contributes to quiescence and antibiotic tolerance in Staphylococcus aureus.

Staphylococcus aureus can cause infections that are often chronic and difficult to treat, even when the bacteria are not antibiotic resistant because most antibiotics act only on metabolically active cells. Subpopulations of persister cells are metabolically quiescent, a state associated with delayed growth, reduced protein synthesis, and increased tolerance to antibiotics. Serine-threonine kinases and phosphatases similar to those found in eukaryotes can fine-tune essential bacterial cellular processes, such as metabolism and stress signaling. We found that acid stress-mimicking conditions that S. aureus experiences in host tissues delayed growth, globally altered the serine and threonine phosphoproteome, and increased threonine phosphorylation of the activation loop of the serine-threonine protein kinase B (PknB). The deletion of stp, which encodes the only annotated functional serine-threonine phosphatase in S. aureus, increased the growth delay and phenotypic heterogeneity under different stress challenges, including growth in acidic conditions, the intracellular milieu of human cells, and abscesses in mice. This growth delay was associated with reduced protein translation and intracellular ATP concentrations and increased antibiotic tolerance. Using phosphopeptide enrichment and mass spectrometry-based proteomics, we identified targets of serine-threonine phosphorylation that may regulate bacterial growth and metabolism. Together, our findings highlight the importance of phosphoregulation in mediating bacterial quiescence and antibiotic tolerance and suggest that targeting PknB or Stp might offer a future therapeutic strategy to prevent persister formation during S. aureus infections.

Animals

Receptor-defined targeting of a genomically unique melanoma-enriched noncanonical antigen.

Effective T cell-based immunotherapies require functional receptors that can be engineered and redeployed to recognize tumor-restricted antigens. Noncanonical peptides arising from transcription outside annotated protein-coding regions expand the antigenic landscape of cancer; however, systematic strategies to biologically prioritize and functionally validate such targets remain underdeveloped. Here, we integrated de novo transcript analysis, exon-resolved quantification, RNA in situ hybridization, and immunopeptidomics to identify melanoma-associated noncanonical transcripts and advance candidates through receptor-level validation. Among three recurrent melanoma-associated transcripts, EVA003 emerged as a lead target based on its distinct repeat-enriched genomic architecture, consistent tumor-enriched exon-level expression across independent datasets, and a genomically unique immunogenic core sequence. We demonstrate endogenous presentation of EVA003-derived peptides on HLA-A*03:01 and detect specific reactivity in patient-derived tumor-infiltrating lymphocytes. Single-cell transcriptomic profiling identified a dominant peptide-reactive clonotype, enabling isolation of a naturally occurring T cell receptor. Transfer of this receptor into healthy donor T cells conferred antigen-dependent activation and cytotoxicity against both peptide-pulsed targets and melanoma cells expressing EVA003 endogenously. Together, these findings establish a biologically informed strategy for prioritizing noncanonical tumor antigens and demonstrate that genomically unique, tumor-enriched noncanonical peptides can be presented to molecularly defined receptors capable of mediating cancer cell killing. These findings support the integration of prioritized noncanonical antigens into engineered T cell therapeutic strategies.

Humans

Chromosome-level genome assembly of bivalve mollusk, Xishishe Coelomactra antiquata.

Coelomactra antiquata, a significant marine economic shellfish in China, is experiencing a natural population decline due to habitat destruction and overfishing, making the restoration and conservation of its natural resources an urgent priority. This study provides a high - quality chromosome - level genome assembly for C. antiquata, created by PacBio and Hi - C sequencing and resulting in a 19 - chromosome map. The assembly encompasses a genome size of 807.31 Mb, with a contig N50 of 17.35 Mb and a scaffold N50 of 42.90 Mb. A total of 28,070 protein - coding genes were identified, 25,959 of which were functionally annotated. Overall, this study offers a chromosome - level genome for C. antiquata that is highly continuous and complete, providing an indispensable resource for subsequent molecular and genetic studies of this species.

Animals

Systematic discovery of pathogen effector functions across human pathogens and pathways.

Pathogens deploy effector proteins to exploit host cell biology, and most effector open reading frames (ORFs) are rapidly evolving and lack functional annotation. We developed the effector ORFeome (eORFeome), a scalable functional genomics platform encompassing 3,835 effector ORFs from diverse viruses, bacteria, and parasites. High-throughput barcoded screens across nuclear factor κB (NF-κB), apoptosis, p53, cGAS-STING, and major histocompatibility complex class I (MHC class I) pathways revealed novel pathway-modulating functions for hundreds of uncharacterized eORFs, unexpected activities of known effectors, and distinct pathway-specific functions encoded by single ORFs. Illustrating the power of this approach, we identified HHV6A U14 as a p53 antagonist, HHV7 U21 as a dual-function STING antagonist and MHC-I antigen display inhibitor, and adenoviral 13.6K/i-leader protein as a de novo-evolved TAP inhibitor that suppresses MHC-I display. These results establish a general framework for systematic effector annotation, uncover new mechanisms of host-pathogen interaction across kingdoms, and highlight pathogen effectors as a versatile toolkit for rewiring and probing human cellular pathways.

Humans

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14 Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16 Mb and scafold N50 length of 21.24 Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40 Mb, with a scaffold N50 of 125.01 Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22 Mb in size with a scaffold N50 of 7.81 Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes

Genome-wide screening and functional analysis of protein glycosylation-related genes involved in tomato fruit ripening.

Protein glycosylation, an essential co- and post-translational modification, plays critical roles in plant growth, development, and stress responses. However, its functional role in tomato fruit ripening has not been extensively investigated. Here, key protein glycosylation-related genes involved in tomato fruit ripening were identified by genome-wide screen and subsequently functional characterization. First, a dataset comprising 242 glycosylation-related proteins was established based on Gene Ontology annotations in tomato, combined with sequence homology to protein glycosylation-related proteins from Arabidopsis thaliana and Homo sapiens. Then, Subsequently, 28 genes encoding highly expressed glycosylation-related proteins (RPKM > 30) at the breaker (BR) stage were selected for functional screening, and subsequently 6 genes were identified as regulators of fruit ripening by method of virus-induced gene silencing (VIGS). Among them, Solyc03g098600 (STT3B), Solyc01g109410 (OST48), Solyc04g082670 (RPN1), and Solyc08g076460 (DAD1) functioned as positive regulators of tomato fruit ripening, whereas Solyc04g005340 (UAM2) and Solyc08g075340 (XEG113), acted as negative regulators. The expression of these genes responded dynamically to multiple ripening-related cues, including temperature, light, ethylene, and transcription factors. Furthermore, silencing of these genes individually affected the expression of genes involved in fruit ripening, including ethylene biosynthesis genes (ACS2, ACS4, ACO1, and ACO3), ripening-associated transcription factors (RIN, NOR, NOR-LIKE1, FUL1, and FUL2), and the key gene (PSY1) of lycopene biosynthesis pathway. Collectively, these findings demonstrate that protein glycosylation plays an important role in tomato fruit ripening by modulating ethylene signaling, ripening-associated transcriptional regulation, and lycopene biosynthesis.

Fruit ripening

Genomes of the ex-type strains of Elsinoë mangiferae and E. perseae, the causal agents of scab on mango and avocado.

Elsinoë species are slow-growing, hemibiotrophic to necrotrophic fungi that cause scab diseases on economically important fruit crops. Genome resources for many host-specific species remain limited. We report high-quality draft genome assemblies for the ex-type strains of Elsinoë mangiferae (CBS 226.50) and E. perseae (CBS 406.34), causal agents of mango and avocado scab, respectively. Among 5 approaches tested, a Nanopore-only NextDenovo assembly produced the most contiguous genomes, yielding 24.5 Mb (E. mangiferae) and 25.1 Mb (E. perseae) assemblies with 13 and 18 contigs, respectively, BUSCO completeness scores of ∼94%, and multiple putative telomere-to-telomere chromosomes. Gene prediction identified 9,134 and 9,243 genes, respectively. Functional annotation revealed enrichment of metabolic and regulatory pathways, including those involved in posttranslational modification, protein transport, and secondary metabolism. Carbohydrate-active enzyme repertoires were small but conserved, consistent with stealth pathogenicity strategies and low plant cell wall degradation. Both genomes encoded large secretomes (>850 proteins), diverse protease repertoires (>300 proteins), Ecp2-like effector proteins, and multiple biosynthetic gene clusters, including clusters with similarity to those associated with elsinochrome and ACT-toxin II biosynthesis, some of which may contribute to host-pathogen interactions and disease development. A large fraction of genes lacked functional characterization, suggesting incomplete databases and/or the presence of lineage-specific genes potentially involved in virulence or host adaptation. These genome resources fill critical gaps for underrepresented Elsinoë species and provide taxonomically anchored references essential for diagnostics, comparative genomics, and research into the molecular basis of host specificity and pathogenicity in scab-causing fungi.

Persea