PubMed HealthSearch

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

Systematic Proteome Profiling of Maternal Plasma for Development of Preeclampsia Biomarkers.

Preeclampsia (PE) is a hypertensive disorder of pregnancy with various clinical symptoms. However, traditional markers for the disease including high blood pressure and proteinuria are poor indicators of the related adverse outcomes. Here, we performed systematic proteome profiling of plasma samples obtained from pregnant women with PE to identify clinically effective diagnostic biomarkers. Proteome profiling was performed using TMT-based liquid chromatography-mass spectrometry (LC-MS/MS) followed by subsequent verification by multiple reaction monitoring (MRM) analysis on normal and PE maternal plasma samples. Functional annotations of differentially expressed proteins (DEPs) in PE were predicted using bioinformatic tools. The diagnostic accuracies of the biomarkers for PE were estimated according to the area under the receiver-operating characteristics curve (AUC). A total of 1307 proteins were identified, and 870 proteins of them were quantified from plasma samples. Significant differences were evident in 138 DEPs, including 71 upregulated DEPs and 67 downregulated DEPs in the PE group, compared with those in the control group. Upregulated proteins were significantly associated with biological processes including platelet degranulation, proteolysis, lipoprotein metabolism, and cholesterol efflux. Biological processes including blood coagulation and acute-phase response were enriched for down-regulated proteins. Of these, 40 proteins were subsequently validated in an independent cohort of 26 PE patients and 29 healthy controls. APOM, LCN2, and QSOX1 showed high diagnostic accuracies for PE detection (AUC >0.9 and p&#xa0;<&#xa0;0.001, for all) as validated by MRM and ELISA. Our data demonstrate that three plasma biomarkers, identified by systematic proteomic profiling, present a possibility for the assessment of PE, independent of the clinical characteristics of pregnant women.

Humans

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (&#xb1;5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

Serum Proteomic Profiling Reveals Renin-Associated Immune and Cytoskeletal Dysregulation in Post-COVID-19 Condition Patients with Secondary Adrenal Insufficiency.

Post-COVID-19 condition (PCC) with secondary adrenal insufficiency (SAI) involves multiorgan dysfunction, potentially linked to renin-angiotensin-aldosterone system dysregulation. The molecular basis of renin-associated pathology remains unclear. Here, PCC+SAI patients were stratified by upright renin into low- (<38.8&#x202f;pg/mL) and high-renin (&#x2265;38.8&#x202f;pg/mL) groups. Clinical, endocrine, and proteomic analyses were performed. We found that high-renin patients showed increased BMI, lipids, renin, and aldosterone, but reduced aldosterone-to-renin ratio. Proteomic annalysis identified 20 differentially expressed proteins (DEPs), including 17 upregulated and 3 downregulated proteins in Ren-H patients. Functional annotation revealed that 15 DEPs were immune-related (e.g., APOC4, APOE, C4BPA, CFAH, CFHR3, PF4V, PLF4), while FLNA and COF1 represented cytoskeletal proteins. These DEPs were primarily involved in immune response, complement and coagulation cascades, and MAPK signaling pathways. Correlation analyses indicated that upright renin was positively correlated with complement-related proteins and platelet-derived immune factors, while cytoskeletal proteins (FLNA, COF1) showed positive associations with serum Na+ levels. Additionally, white blood cell and platelet counts were positively correlated with the majority of DEPs. In conclusion, exploratory proteomic analyses suggest that elevated upright renin in PCC+SAI may be associated with immune dysregulation, complement activation, and cytoskeletal remodeling, offering novel insights into the endocrine-immune interactions driving postviral sequelae.

Humans

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency &#x2264;&#x2009;0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans

Shared genetic architecture and therapeutic targets across paediatric immune-mediated diseases.

OBJECTIVES: Paediatric-onset immune-mediated inflammatory diseases (IMIDs), including juvenile idiopathic arthritis and related rheumatic diseases, remain genetically undercharacterised. We aimed to define shared and category-specific genetic architecture across paediatric IMIDs, compare signals with adult IMIDs, and identify therapeutic opportunities. METHODS: We analysed 24 paediatric IMIDs classified as autoimmune, polygenic-autoinflammatory, mixed-pattern, or allergic. Genome-wide association analyses included 18,086 cases and 131,019 controls of European ancestry. We estimated single nucleotide polymorphism (SNP)-based heritability, genetic correlations, and polygenic overlap; performed subset-based meta-analysis; and conducted functional annotation, gene prioritisation, pathway and protein network analyses, adult-IMID comparison, and drug-target prioritisation. RESULTS: SNP-based heritability ranged from 28.9% for allergic IMIDs to 61.9% for autoimmune IMIDs. Genetic correlation and polygenic modelling supported partial sharing across categories with category-specific components. Meta-analysis identified 39 genome-wide significant loci outside the Major Histocompatibility Complex (MHC) region, including 15 previously unreported loci; 19 loci were shared between categories. Gene-prioritisation and protein interaction analyses identified a core MHC-centred antigen-presentation network, with category-enriched modules involving complement, innate/barrier pathways, epithelial biology, and type 2 immunity. Enriched pathways included nuclear factor &#x3ba;B signalling, T helper 17 related pathways, Janus kinase-signal transducer and activator of transcription signalling, programmed cell death protein 1/programmed death&#x2011;ligand 1, cytotoxic T&#x2011;lymphocyte associated protein 4 regulation, and osteoclast differentiation, several of which are relevant to rheumatic diseases. Paediatric IMIDs shared broad polygenic architecture with adult IMIDs, whereas top-ranked genes converged strongly with adult rheumatic diseases. Priority Index analysis identified 178 high-scoring genes, including 43 approved or investigational IMID drug targets. CONCLUSIONS: Paediatric-onset IMIDs share core pathways with adult forms but exhibit distinct genetic architecture shaped by age-specific immune and neurodevelopmental biology. These findings provide a genomic framework for paediatric precision medicine, guiding classification, risk prediction, and therapeutic development.

Humans

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Draft genome assembly of the green-bronze dung beetle, Onthophagus orpheus.

Dung beetles (Coleoptera: Scarabaeinae) are ecologically important insects, yet genomic resources for this diverse lineage remain limited. Here, we present a high-quality genome assembly for Onthophagus orpheus, an understudied species that is abundant in urban forests in the eastern United States. The assembled genome is a scaffold-level assembly, with a high degree of genic completeness as assessed by Benchmarking Universal Single-Copy Ortholog (BUSCO) analyses, indicating robust representation of conserved protein-coding genes. Structural and functional annotation recovered a comprehensive gene set consistent with expectations for coleopteran genomes. This genome assembly provides an important resource for future work on the behavioral ecology and population genetics of Onthophagus orpheus, specifically, and Scarabaeidae more broadly.

Onthophagus

A Network Pharmacology and Molecular Docking Study of TongBi Formula for Osteoarthritis.

This study applied network pharmacology combined with molecular docking to predict the potential therapeutic targets and molecular mechanisms of TongBi Formula (TBF) in osteoarthritis (OA). Active components and corresponding targets of TBF were retrieved from the traditional Chinese medicine Systems Pharmacology Database and Analysis Platform, while OA-related targets were collected from Online Mendelian Inheritance in Man, GeneCards, DrugBank, and Therapeutic Target Database. A network visualization and analysis software was used to construct compound-target and protein-protein interaction (PPI) networks. Gene Ontology functional annotation and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed using the Database for Annotation, Visualization and Integrated Discovery platform. Molecular docking analysis was conducted using a molecular docking software to evaluate the predicted binding affinity between key active compounds and core target proteins. A total of 47 overlapping targets between TBF and OA were identified. PPI network analysis highlighted JUN, RELA, IL6, MAPK1, and IL10 as potential hub targets. Enrichment analysis suggested that TBF may regulate inflammation, lipid metabolism, and multiple intracellular signaling pathways associated with OA progression. Molecular docking results demonstrated favorable predicted binding affinities between core active compounds and key OA-related protein targets. These findings provide a computational framework for understanding the potential mechanisms of TBF against OA and support further experimental validation.

Molecular Docking Simulation

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3&#x2006;460&#x2006;657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120&#x2006;985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-&#xc5; resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between &#x3b2;H and &#x3b2;I of the core &#x3b2;-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15&#x2009;Gb (contig N50&#x2009;=&#x2009;43.57&#x2009;Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03&#x2009;Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17&#x2009;Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

RNAi screening of uncharacterized genes identifies promising druggable targets in Schistosoma japonicum.

Schistosomiasis affects more than 250 million people worldwide and is one of the neglected tropical diseases. Currently, the treatment of schistosomiasis relies on a single drug-praziquantel-which has led to increasing pressure from drug resistance. Therefore, there is an urgent need to find new treatments. The development of genome sequencing has provided valuable information for understanding the biology of schistosomes. In the genome of Schistosoma japonicum, approximately 11% of the protein-coding sequences are uncharacterized genes (UGs) annotated as "hypothetical protein" or "protein of unknown function." These poorly understood genes have been unjustifiably neglected, although some may be essential for the survival of the parasites and serve as potential drug targets. In this study, we systematically mined the highly expressed UGs in both genders of this parasite throughout key developmental stages in their mammalian host, using our previously published S. japonicum genome and RNA-seq data. By employing in vitro RNA interference (RNAi), we screened 126 UGs that lack homologs in Homo sapiens and identified 8 that are essential for the parasite vitality. We further investigated two UGs, Sjc_0002003 and Sjc_0009272, which resulted in the most severe phenotypes. Fluorescence in situ hybridization demonstrated that both genes were expressed throughout the body without sex bias. Silencing either Sjc_0002003 or Sjc_0009272 reduced the cell proliferation in the body. Furthermore, in vivo RNAi indicated both genes are required for the growth and survival of the parasites in the mammalian host. For Sjc_0002003, we further characterize the underlying molecular cause of the observed phenotype. Through RNA-seq analysis and functional studies, we revealed that silencing Sjc_0002003 reduces the expression of a series of intestinal genes, including Sjc_0007312 (hypothetical protein), Sjc_0008276 (vha-17), Sjc_0002942 (PLA2G15), and Sjc_0003646 (SJCHGC09134 protein), leading to gut dilation. Our work highlights the importance of UGs in schistosomes as promising targets for drug development in the treatment of the schistosomiasis.

Schistosoma japonicum