PubMed HealthSearch

SEARCH · PubMed Health

Results for “Protein Domains”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The Arabidopsis TIRome informs the design of artificial TIR (Toll/interleukin-1 receptor) domain proteins.

The TIR (Toll/interleukin-1 receptor) domain is an ancient protein module that functions in immune and cell death responses across the Tree of Life. TIR domains encoded by plants and prokaryotes function as enzymes to produce diverse small molecule immune signals. Plant genomes can encode hundreds of TIR-domain containing proteins-many of which confer important agricultural disease resistance as TIR-NLR (nucleotide-binding, leucine-rich repeat) immune receptors. Despite their importance, how natural variation influences TIR enzymatic output and immunity-associated cell death is largely unexplored. We assayed a complete collection of the TIR domains of Arabidopsis thaliana Col-0 (the "AtTIRome") to explore variation in TIR metabolite production and cell death signaling. Roughly half of the AtTIRome triggered cell death in transient assays. Artificial TIR proteins designed based on consensus sequences of the AtTIRome's cell death phenotypic classes revealed polymorphisms controlling variation in TIR cell death elicitation and metabolite production. Structure-function analyses of artificial TIRs revealed that natural variation in the "BB-loop", a flexible region overlying the catalytic pocket, determines differences in function across Arabidopsis TIR-containing proteins. We further demonstrate that artificial TIRs are functional on an NLR chassis and that BB-loop variation can tune the activity of a natural TIR-NLR protein. These findings shed light on the diversity of TIR outputs and reveal methods to design and engineer TIR-based immune receptors.

Arabidopsis

SUPFAM: a database of sequence superfamilies of protein domains.

BACKGROUND: SUPFAM database is a compilation of superfamily relationships between protein domain families of either known or unknown 3-D structure. In SUPFAM, sequence families from Pfam and structural families from SCOP are associated, using profile matching, to result in sequence superfamilies of known structure. Subsequently all-against-all family profile matches are made to deduce a list of new potential superfamilies of yet unknown structure. DESCRIPTION: The current version of SUPFAM (release 1.4) corresponds to significant enhancements and major developments compared to the earlier and basic version. In the present version we have used RPS-BLAST, which is robust and sensitive, for profile matching. The reliability of connections between protein families is ensured better than before by use of benchmarked criteria involving strict e-value cut-off and a minimal alignment length condition. An e-value based indication of reliability of connections is now presented in the database. Web access to a RPS-BLAST-based tool to associate a query sequence to one of the family profiles in SUPFAM is available with the current release. In terms of the scientific content the present release of SUPFAM is entirely reorganized with the use of 6190 Pfam families and 2317 structural families derived from SCOP. Due to a steep increase in the number of sequence and structural families used in SUPFAM the details of scientific content in the present release are almost entirely complementary to previous basic version. Of the 2286 families, we could relate 245 Pfam families with apparently no structural information to families of known 3-D structures, thus resulting in the identification of new families in the existing superfamilies. Using the profiles of 3904 Pfam families of yet unknown structure, an all-against-all comparison involving sequence-profile match resulted in clustering of 96 Pfam families into 39 new potential superfamilies. CONCLUSION: SUPFAM presents many non-trivial superfamily relationships of sequence families involved in a variety of functions and hence the information content is of interest to a wide scientific community. The grouping of related proteins without a known structure in SUPFAM is useful in identifying priority targets for structural genomics initiatives and in the assignment of putative functions. Database URL: http://pauling.mbu.iisc.ernet.in/~supfam.

Amino Acid Sequence

Nuclear and cytosolic J-domain proteins provide synergistic control of Hsf1 at distinct phases of the heat shock response.

The heat shock response (HSR) is the major defense mechanism against proteotoxic stress in the cytosol and nucleus of eukaryotic cells. Initiation and attenuation of the response are mediated by stress-dependent regulation of heat shock transcription factors (HSFs). Saccharomyces cerevisiae encodes a single HSF (Hsf1), facilitating the analysis of HSR regulation. Hsf1 is repressed by Hsp70 chaperones under non-stress conditions and becomes activated under proteotoxic stress, directly linking protein damage and its repair to the HSR. J-domain proteins (JDPs) are essential for targeting of Hsp70s to their substrates, yet the specific JDP(s) regulating Hsf1 and connecting protein damage to HSR activation remain unclear. Here, we show that the yeast nuclear JDP Apj1 primarily controls the attenuation phase of the HSR by promoting Hsf1's displacement from heat shock elements in target DNA. In apj1Δ cells, HSR attenuation is significantly impaired. Additionally, yeast cells lacking both Apj1 and the major JDP Ydj1 exhibit increased HSR activation even in non-stress conditions, indicating their distinct regulatory roles. Apj1's role in both nuclear protein quality control and Hsf1 regulation underscores its role in directly linking nuclear proteostasis to HSR regulation. Together, these findings establish the nucleus as key stress-sensing signaling hub.

Saccharomyces cerevisiae Proteins

A novel transformer model of protein domains for viral taxonomy classification.

MOTIVATION: Viruses with carefully curated taxonomic assignments (such as those in the ICTV taxonomy) still represent only a small fraction of viruses identified through sequencing data from virome or microbiome projects. It is therefore critical to develop methods that can assign viruses at multiple taxonomic ranks, so that a virus deemed novel at a given rank may still be placed into a higher-level taxon. Sequence-similarity-based approaches can classify viruses that share substantial genomic similarity with known viruses (e.g. those belonging to the same species or genus); however, their performance drops significantly when applied to more divergent viruses. Recent deep learning models, such as ViTax, which utilize DNA language models, aim to address these limitations, but their performance also degrades when applied to novel viruses lacking genus-level similarity to known references. Proteins are more conserved than genomic sequences, and the multiple proteins encoded by a virus can be leveraged to reveal evolutionary relationships among viruses. RESULTS: We propose a new tool, D2T (Domain-to-Taxonomy), that leverages recent advances in protein language models to improve viral taxonomic assignment. D2T represents a virus as a sequence of protein domain tokens and learns a transformer-based model for taxonomic classification. Experiments on multiple closed-set and open-set datasets show that D2T excels at assigning higher-level taxonomic labels (family and above). Furthermore, by combining D2T with Kraken2, which performs well at the genus level, the hybrid method (K+D2T) achieves accurate viral taxonomic classification across multiple taxonomic ranks. AVAILABILITY AND IMPLEMENTATION: D2T is available as a GitHub repository at https://github.com/mgtools/D2T.

Viruses

Drug target ontology to classify and integrate drug discovery data.

BACKGROUND: One of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome. RESULTS: As part of that effort, we have developed a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while supporting manual knowledge acquisition. A modular, hierarchical implementation facilitate ontology development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships. CONCLUSIONS: DTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/ , Github ( http://github.com/DrugTargetOntology/DTO ), and the NCBO Bioportal ( http://bioportal.bioontology.org/ontologies/DTO ). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

Biological Ontologies

Golgi_traff phylogeny reveals ancient eukaryotic genes with recent surprises: replication and diversification of HID1 domain-containing protein unique to Schizosaccharomyces.

Golgi_traff is a Pfam clan containing two members, Dymeclin (DYM) and HID1 domain-containing protein (HID). Interrogation of over 900 eukaryotic genomes with sequence models showed that both are ancient eukaryotic genes, which have exhibited different paths of gene loss, including from major taxonomic groups. For example, the Metazoa have both genes, whereas the Viridiplantae and Dikarya have lost HID and DYM, respectively. A unique replication event occurred within the genus Schizosaccharomyces in that all sequenced species possess three HID-encoding paralogs, whereas its nearest fungal relatives and other eukaryotes are almost exclusively monogenic. A phylogenetic analysis of yeasts revealed that the Golgi-resident paralog Human ortholog 3 (SPAC17A5.16) is more similar to the HID of other yeasts than to its paralogs. Transmission electron microscopy revealed that the SPAC17A5.16 mutant lacks a stacked Golgi apparatus (GA) form, suggesting a role in maintaining GA structure. Altered proliferation of the SPAC17A5.16 mutant in response to GA disrupting chemical agents indicated a perturbation of GA-related functions. Structural models suggest SPAC17A5.16 has a long, disordered N-terminal region that may facilitate anchoring to GA membranes. A modification to Schizosaccharomyces HID nomenclature is proposed to reflect their evolutionary and functional characteristics. The potential of the Golgi_traff clan to serve as a model for the diversification of protein function according to the concepts of sub/neofunctionalization is discussed.

Schizosaccharomyces

The Key Trichoderma-Induced Gene Encoding a DUF568 Domain-Containing Protein Mediates Defense Responses in Wheat.

Genes encoding DUF568 domain-containing proteins participate in plant stress adaptation. To elucidate the functional role of DUF568 domain-containing genes in Trichoderma-induced wheat defense responses against wheat Fusarium crown rot, we performed a genome-wide identification and characterization of the TaDUF568 gene family in hexaploid wheat (Triticum aestivum L.). In this study, a total of 33 TaDUF568 family genes were systematically identified and characterized at the genome-wide level, exhibiting uneven chromosomal distribution and diverse physicochemical properties. Phylogenetic, structural, and collinearity analyses revealed conserved family characteristics among monocot species. Segmental duplication was verified as the primary driver of gene family expansion. Expression profiling revealed divergent tissue-specific expression patterns among TaDUF568 family members, among which TaDUF568.18 was strongly induced by Trichoderma M2. Subcellular localization assays confirmed that TaDUF568.18 is a plasma membrane-localized protein. Functional validation via stable transgenes demonstrated that overexpression of TaDUF568.18 restricted lesion expansion, improved agronomic traits, and enhanced disease resistance. This study is the first to characterize the wheat DUF568 family and confirm that TaDUF568.18 (annotated as TaAIR12) acts as a positive regulator of Trichoderma-mediated wheat defense, providing a valuable gene resource for wheat disease-resistance breeding.

DUF568

CoDIAC: A comprehensive approach for interaction analysis reveals novel insights into SH2 domain function and regulation.

Protein domains are conserved structural and functional units that serve as building blocks of proteins. Through evolutionary expansion, domain families are represented by multiple members in diverse configurations with other domains, evolving new specificities for their interacting partners. Here, we develop a structure-based interface analysis to comprehensively map domain interfaces from experimental and predicted structures, including interfaces with macromolecules and intraprotein interfaces. We hypothesized that comprehensive contact mapping of domains could yield new insights into domain selectivity, conservation of domain-domain interfaces across proteins, and identify conserved post-translational modifications (PTMs), relative to interaction interfaces, allowing for the inference of specific effects due to PTMs or mutations. We applied this approach to the human SH2 domain family, a modular unit central to phosphotyrosine-mediated signaling, identifying a novel approach to understanding binding selectivity and evidence of coordinated regulation of SH2 domain binding interfaces by tyrosine and serine/threonine phosphorylation and acetylation. These findings suggest multiple signaling systems can regulate protein activity and SH2 domain interactions in a coordinated manner. We provide the extensive features of the human SH2 domain family and this modular approach as an open source Python package for COmprehensive Domain Interface Analysis of Contacts (CoDIAC).

SH2 domains

Transposable Elements Drive Regulatory and Functional Innovation of F-box Genes.

Protein domains of transposable elements (TEs) and viruses increase the protein diversity of host genomes by recombining with other protein domains. By screening 10 million eukaryotic proteins, we identified several domains that define multicopy gene families and frequently co-occur with TE/viral domains. Among these, a Tc1/Mariner transposase helix-turn-helix (HTH) domain was captured by F-box genes in the Caenorhabditis genus, creating a new class of F-box genes. For specific members of this class, like fbxa-215, we found that the HTH domain is required for diverse processes including germ granule localization, fertility, and thermotolerance. Furthermore, we provide evidence that Heat Shock Factor 1 (HSF-1) mediates the transcriptional integration of fbxa-215 into the heat shock response by binding to Helitron TEs directly upstream of the fbxa-215 locus. The interactome of HTH-bearing F-box factors suggests roles in post-translational regulation and proteostasis, consistent with established functions of F-box proteins. Based on AlphaFold2 multimer proteome-wide screens, we propose that the HTH domain may diversify the repertoire of protein substrates that F-box factors regulate post-translationally. We also describe an independent capture of a TE domain by F-box genes in zebrafish. In conclusion, we identify two independent TE domain captures by F-box genes in eukaryotes and provide insights into how these novel proteins are integrated within host gene regulatory networks.

Animals

Genetic Testing Unveils a Novel Thrombospondin-1 Domain Containing Protein 1 Gene Variant as the Cause of Chronic Edema in a 79-Year-Old Woman.

Edema requires management tailored to its underlying etiology; however, in some cases the cause remains elusive. We describe a 79-year-old woman with lifelong unexplained peripheral edema. Comprehensive evaluation excluded common etiologies such as heart failure, renal dysfunction, and venous thrombosis. Whole-genome sequencing identified a novel homozygous splice-site variant (NM_018676.4:c.58+2T>G) in the thrombospondin-1 domain-containing protein 1 (THSD1) gene, which has previously been associated with non-immune hydrops fetalis (NIHF). This report describes, to our knowledge, the first elderly patient with chronic peripheral edema harboring a likely pathogenic THSD1 variant, suggesting that THSD1-related disease may, in rare instances, persist beyond the perinatal period and present into late adulthood. Although causality cannot be established definitively from a single case, the findings highlight the potential utility of genetic testing in adults with chronic unexplained edema.

chronic edema

Spatio-genetically coordinated TPR domain-containing proteins modulate c-di-GMP signaling in Vibrio vulnificus.

Vibrio species, which include several pathogens, are autochthonous to estuarine and warm coastal marine environments, where biofilm formation bolsters their ecological persistence and transmission. Here, we identify a bicistronic operon, rcbAB, whose products synergistically inhibit motility and promote biofilm maturation post-attachment by modulating intracellular c-di-GMP levels in the human and animal pathogen V. vulnificus. RcbA contains an N-terminal tetratricopeptide repeat (TPR) domain and a structured C-terminal region of unknown function, while RcbB possesses an N-terminal TPR domain and a C-terminal GGDEF domain characteristic of diguanylate cyclases. The TPR domain of RcbB represses its diguanylate cyclase activity, while RcbA's TPR domain and C-terminal region co-operatively de-repress it. Localization of both proteins to the flagellar pole is TPR-dependent but not co-dependent, although RcbA anchors RcbB to the pole in the absence of polar landmarks such as HubP and flagella. The conservation of rcbAB across diverse bacterial taxa substantiates its fundamental importance in bacterial biology. This work demonstrates how spatio-genetically coordinated TPR domain-containing proteins modulate c-di-GMP signaling, contributing to our understanding of biofilm formation in Vibrio species and potentially other bacteria. It also reveals the first evidence of inter-protein interaction via the TPR domains of both partners, challenging the conventional paradigm in which only one bears the domain.

Vibrio vulnificus

Targeted Epigenetic Silencing of Jumonji Domain-Containing Protein 3 Alleviates Nuclear Factor-Kappa B-Mediated Inflammation in Familial Mediterranean Fever.

BACKGROUND: Familial Mediterranean fever (FMF) is an inherited autoinflammatory condition caused by variants in the MEFV gene encoding pyrin, the essential component of the NLRP3/NF-κB complex of inflammasomes. Deregulation of nuclear factor-kappa B (NF-κB), a key proinflammatory mediator, leads to chronic inflammation in autoinflammatory/autoimmune diseases. Epigenetic modulation offers a new approach to regulate inflammasome activity, with Jumonji domain-containing protein 3 (JMJD3) being a promising target for managing inflammatory illnesses. GSK-J4 is a selective inhibitor of JMJD3, restricting pro-inflammatory cytokines and inflammation. AIM: Our research aimed to elucidate the role of JMJD3 and the NF-κB-JMJD3 signaling pathways in regulating inflammation in an in vitro model, and to investigate GSK-J4's effect in inhibiting inflammasome activation in primed peripheral blood mononuclear cells (PBMCs) isolated from FMF cases. METHODS: PBMCs were cultured and primed with LPS, and then treated with GSK-J4. JMJD3 knockdown was achieved using siRNA interference. Cellular inflammatory dynamics were assessed by Western blotting (WB) and ELISA. The qRT-PCR was used for gene expression quantification. Untreated cells served as a negative control. RESULTS: Our results showed significantly downregulated gene expression of NF-κB, NLRP3, and inflammatory cytokines in GSK-J4-treated cells compared to untreated cells, as confirmed by ELISA. WB reported a reduction of NF-κB in induced cells following GSK-J4 treatment. Knocking down JMJD3 also showed decreased levels of JMJD3, NF-κB, and inflammatory cytokines, indicating its proinflammatory role. CONCLUSION: The study showed that selective inhibition or silencing of JMJD3 significantly suppressed the inflammasome in FMF cases, suggesting its role as a therapeutic target for alleviating inflammation in various autoinflammatory diseases.

Humans

Characterizing the Discordance between AT-Rich Interacting Domain 1A Protein and Genotype in Endometrioid-Type Endometrial Tumors.

ARID1A is one of the most frequently mutated genes in endometrial cancer, with approximately 40% of patients harboring an ARID1A mutation. However, relatively little is known about how AT-rich interacting domain 1A (ARID1A) protein loss shapes endometrial cancer pathogenesis. Mounting evidence from other malignancies suggests that ARID1A protein can be regulated post-translationally, independent of genotype. However, most studies in endometrial cancer evaluate genotype alone, overlooking the potential for alternative mechanisms of ARID1A loss. To address this gap, ARID1A protein expression and genotype were examined in endometrioid tumors, and associated transcriptional changes were characterized. Evaluation of ARID1A protein in 71 human endometrioid tumors demonstrates that protein loss can occur regardless of ARID1A genotype. Retention or deficiency of ARID1A protein was not significantly related to variant allele frequency or location of mutation in human tumors with mutant ARID1A. A human endometrial cancer cell model suggests that ARID1A protein loss can occur through proteasomal degradation. Furthermore, ARID1A protein expression was found to be a predictor of worse overall survival in The Cancer Genome Atlas cohort of ARID1A wild-type endometrioid tumors. Spatial transcriptomics of 16 human endometrioid tumors revealed that both genotype and protein expression of ARID1A play a role in shaping unique transcriptional signatures in endometrial cancer and can be used to predict patient prognosis. Suggesting evaluation of ARID1A should not be done solely by sequencing techniques.

Humans

The Jumonji C domain-containing proteins GmJMJ19 and GmJMJ20 link florigen signaling with epigenetic regulation of photoperiodic flowering and post-flowering plant height in soybean.

Soybean (Glycine max) is a photoperiod-sensitive legume whose latitudinal adaptation depends on the precise control of flowering time and plant height. Histone demethylases of the JmjC domain-containing (JMJ) protein family have been implicated in these processes across plant species, but their specific roles in soybean remain largely unexplored. Here, we identify soybean GmJMJ19 and GmJMJ20, two closely related JMJD5/KDM8 orthologs, as master epigenetic regulators that coordinately control both photoperiodic flowering and post-flowering plant height. Both genes exhibit intrinsic, rhythmic expression peaking at ZT12, and their encoded proteins physically interact with the florigen proteins FT2a and FT5a. Loss-of-function mutants display delayed flowering under long days (LDs) and increased plant height under both LDs and short days (SDs), whereas overexpression phenocopies the mutant flowering phenotype, indicating revealing a critical dosage requirement for proper function. Mechanistically, GmJMJ19 and GmJMJ20 are recruited by the FT/FD transcriptional complex to directly activate AP1a and AP1c expression through chromatin modulation. Population genomic analyses reveal distinct selection signatures: GmJMJ19 underwent sustained directional selection during cultivation, whereas GmJMJ20 experienced an early domestication sweep with limited subsequent change. Haplotype analysis identifies coordinated latitudinal clines, with the JMJ19H1/JMJ20H1 combination predominating at high latitudes to promote early flowering and limit height, while JMJ19H2/JMJ20H2 and wild JMJ19H3/JMJ20H3 alleles prevail at low latitudes, conferring later flowering and increased height. Collectively, our findings establish GmJMJ19 and GmJMJ20 as central chromatin regulators linking florigen signaling to downstream target expression and provide valuable allelic resources for breeding regionally adapted soybean varieties across a wide range of latitudinal environments.

Histone modulation

Rapid assessment of clinical severity for salmonellosis cases via protein family domain analysis and machine learning.

Salmonella is a common pathogen, infecting more than a million people yearly. Rapid assessment of clinical case severity is essential for improving patient outcomes and optimizing healthcare resources. Advancements in genome sequencing technologies have enabled the analysis of bacterial genomes from many clinical cases, opening up new opportunities for precise and timely diagnosis. This study proposes a genome-based framework for identifying critical Salmonella cases before the onset of critical symptoms and facilitating early medical intervention. By leveraging protein family (Pfam) domains as the representation for genomic data, the complex genetic profiles of Salmonella cases are simplified into interpretable features. The severity levels of cases were investigated through rigorous data analysis, resulting in a set of 70 Pfam domains that could be potentially used as biomarkers. Machine Learning was employed to assess the predictive power of the curated Pfam biomarkers, achieving high accuracy (~93%) in sorting cases into critical, moderate, and mild categories. The results demonstrate the efficacy of the proposed approach. This framework highlights the potential of using bacterial genomic data in clinical decision-making, opening the window for timely personalized interventions for Salmonella infection management.

Domains of unknown function (DUFs)

Creating bottom-up RNA transfer vehicles from synthetic protein assemblies.

Evolution guides biological systems to populate ecological niches, with viruses among the most successful examples of this principle. Viruses evolved over billions of years to efficiently transfer genetic information. Although viruses are highly diverse, most have converged towards remarkable similarity in the size and shape of their capsids1,2. By contrast, generative models for protein design enable the creation of protein architectures that are absent from nature3-5. Here we investigate whether protein assemblies designed by artificial intelligence can be functionalized to construct nucleic acid transport vehicles that are independent of evolutionary trajectories. By combining natural protein domains with synthetic protein assemblies, we create more than 100 bottom-up RNA transfer vehicles with unique sizes and shapes. These vehicles surpass the RNA transfer efficiency of widely used delivery vehicles by several orders of magnitude. In addition, we demonstrate that their tropism can be programmed by incorporation of computationally designed peptide binders and use them to deliver therapeutically relevant cargo RNAs into a wide range of cellular models. We show the in vivo biodistribution of one of these vehicles in a mouse at near-single-cell resolution, confirm its safety, and use it to perform a gene-editing treatment strategy for Duchenne muscular dystrophy in patient-derived cells and a pig. Our work demonstrates how proteins created by generative artificial intelligence can be harnessed for the rational engineering of RNA transport systems with the desired properties by overcoming the limitations of natural protein diversity.

Journal Article

Knockdown Proteomics Reveals USP7 as a Regulator of Cell-Cell Adhesion in Colorectal Cancer via AJUBA.

Ubiquitin-specific protease 7 (USP7) is implicated in many cancers including colorectal cancer in which it regulates cellular pathways such as Wnt signaling and the P53-MDM2 pathway. With the discovery of small-molecule inhibitors, USP7 has also become a promising target for cancer therapy and therefore systematically identifying USP7 deubiquitinase interaction partners and substrates has become an important goal. In this study, we selected a colorectal cancer cell model that is highly dependent on USP7 and in which USP7 knockdown significantly inhibited colorectal cancer cell viability, colony formation, and cell-cell adhesion. We then used inducible knockdown of USP7 followed by LC-MS/MS to quantify USP7-dependent proteins. We identified the Ajuba LIM domain protein as an interacting partner of USP7 through co-IP, its substantially reduced protein levels in response to USP7 knockdown, and its sensitivity to the specific USP7 inhibitor FT671. The Ajuba protein has been shown to have oncogenic functions in colorectal and other tumors, including regulation of cell-cell adhesion. We show that both knockdown of USP7 or Ajuba results in a substantial reduction of cell-cell adhesion, with concomitant effects on other proteins associated with adherens junctions. Our findings underlie the role of USP7 in colorectal cancer through its protein interaction networks and show that the Ajuba protein is a component of USP7 protein networks present in colorectal cancer.

Ubiquitin-Specific Peptidase 7

PDP-Miner: an AI/ML tool to detect prophage tail proteins with depolymerase domains across thousands of bacterial genomes.

MOTIVATION: Antibiotic resistance is predicted to become the leading cause of human mortality by 2050. Despite this, no other major antibiotic class has been approved for medical use since 1987. Nevertheless, phage tail proteins offer a promising alternative, given their depolymerase activity toward outer membrane polysaccharides. Several pathogenic bacteria harbor prophages, thus making these prophages' molecular target already known. RESULTS: We therefore developed a wrapper for an existing machine learning-based phage depolymerase prediction tool (Depolymerase-Predictor), called PDP-Miner, which annotates phage tail proteins ab initio, detects depolymerase activity within this candidate protein subset, and then performs post-hoc validation by annotating protein domains thereby allowing the user to investigate for protein domains indicative of depolymerase activity. This tool allowed identification of 10 high confidence phage depolymerase gene candidates across all 1294 Pseudomonas genomes available on the International Pseudomonas Consortium Database while also accurately reporting depolymerases in known phage genomes, similarly to other software like PhageDPO or DepoScope. AVAILABILITY AND IMPLEMENTATION: Source code, test datasets and documentation are freely available for download at http:///www.github.com/jeffgauthier/pdpminer. This software is free and open source under the GNU General Public License v3.0.

Prophages