PubMed HealthSearch

SEARCH · PubMed Health

Results for “Protein engineering”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Engineering Protein Stability with Small Molecules: A Review of the ecDHFR Destabilizing Domain System.

The E. coli dihydrofolate reductase (ecDHFR) destabilizing domain (DD) is a versatile post-translational tool for the conditional control of protein stability via ligand-induced stabilization. In this system, a DD-tagged protein is rapidly degraded by the proteasome unless stabilized by the antibiotic trimethoprim (TMP), allowing for conditional control of protein abundance. The ecDHFR-DD system has been successfully applied across diverse biological systems, including yeast, invertebrate models such as Drosophila, and mammalian cells, to study a broad spectrum of cellular and developmental processes. Compared with DNA- and RNA-based regulatory approaches, post-translational systems offer faster response times and more precise control, making them valuable for processes that require tight, reversible regulation. In this review, we synthesize current knowledge on the mechanisms, performance, and optimization of the ecDHFR-DD system across organisms and evaluate its advantages and limitations relative to most conditional gene expression systems. We also highlight emerging opportunities for applying the system across diverse areas, ranging from functional genomics and synthetic biology to biomedical research. Additionally, we discuss its potential application in applied biological systems, such as pest and vector management, positioning the ecDHFR-DD system as a broadly applicable platform for the precise and tunable control of protein function across diverse disciplines.

Tetrahydrofolate Dehydrogenase

SARS-CoV-2 Orf3a protein interaction mapping using unnatural amino acid incorporation.

Mapping transient protein-protein interactions remain a major challenge in studying viral host-pathogen interfaces. While some virus-host interactions are stable and readily captured, the majority are highly dynamic, reflecting the need for viral proteins to engage distinct host factors at different stages of the life cycle. Here, we employ a protein engineering strategy based on the site-specific incorporation of the unnatural acid p-azido-L-phenylalanine (AzF) to enable photo-crosslinking proteomic analysis of the SARS-CoV-2 accessory protein Orf3a in live cells. Genetic installation of AzF at residue K198 of Orf3a permitted UV-induced covalent capture of proximal host interacting proteins, overcoming challenges associated with membrane localization and limited protein abundance. A total of 248 high-confidence Orf3a-interacting proteins were reproducibly identified and subjected to gene ontology analysis, revealing enrichment in innate immune signaling, antiviral defense, RNA processing, and viral replication-associated pathways. Orf3a is an accessory protein that functions as a viroporin and traffics across multiple cellular compartments, and was found to interact with host RNA helicases, RNA-binding proteins, immune regulators, and metabolic enzymes implicated in SARS-CoV-2 infection. Together, these results demonstrate that genetically encoded, site-specific photo-crosslinking enables selective capture of transient interactions that are often missed by nonspecific 254 nm UV crosslinking approaches and highlights Orf3a as a multifunctional protein that engages diverse host pathways. More broadly, this study establishes a generalizable framework for leveraging unnatural amino acid-based protein engineering approaches to interrogate dynamic host-pathogen interactions.

Humans

Electrochemical sensor toolkit for simultaneous glutamate detection at edge of cleft and peri-soma.

Simultaneously monitoring glutamate (Glu) dynamic at edge of synaptic cleft and peri-soma is crucial for understanding Glu-related pathology. Here, we created an electrochemical Glu sensors toolkit with spatial resolution of ∼60 nm, combining biologically engineered Glu binding protein for specifically capturing Glu together with chemically designed ferrocene groups for signal labeling. Modulation conjugation approach between GluR and ferrocene significantly improved sensitivity up to 32-folds. More importantly, protein engineering of residue mutation and linker peptides flexibility expanded linear range from 10 μM to 6 mM, accelerated on/off times down to 35/40 ms. This toolkit realized real-time quantifying of Glu both at edge of cleft and peri-soma, we discovered that Glu was almost released through SLC7A11 channels in calyx of held synapse upon oxygen-glucose-deprivation, while Glu was mainly released through hemichannels upon β-amyloid42 stimulation. Our work provided a methodology for investigating Glu release and reuptake and offered insights for Glu related pathology.

Glutamic Acid

Charting the development and engineering of CRISPR base editors: lessons and inspirations.

CRISPR base editors (BEs) have introduced a new chapter in precise genome editing. The brief but fruitful history of BE development documents many case studies that not only lay the foundation of base-editing technology but are also instrumental to future protein engineering efforts. In this review, we summarize the development and engineering of various BEs with a focus on recent progress. These include traditional cytosine and adenine base editors (CBEs and ABEs), novel TadA-derived CBEs, transversion BEs, dual BEs, and CRISPR-free BEs. We discuss each aspect of the workflow and highlight the successes and challenges encountered in the engineering process.

Gene Editing

Resolving cellular signaling in space and time: From organelle proteomics to spatial phosphoproteomics.

Cellular signaling is inherently organized in space and time, requiring coordinated control of protein localization, molecular interactions, and enzymatic activity across subcellular compartments. Recent advances in chemical biology, protein engineering, and quantitative proteomics have made it possible to interrogate these dimensions in an integrated manner. Here, we highlight emerging strategies to resolve signaling organization across three interconnected dimensions: organelle-resolved proteome mapping to define spatial context, proximity labeling to capture local protein interaction networks, and spatially resolved phosphoproteomics to quantify signaling outputs. Developments in proximity labeling, including split, conditionally activated and light-gated enzymes, enable temporally controlled, context-dependent profiling of transient protein assemblies in living cells. Advances in high-throughput and low-input phosphoproteomics, together with improved computational frameworks for kinase activity inference and subcellular enrichment strategies, are enabling spatially resolved measurement of signaling activity. Together, these approaches are shifting the field from static localization maps toward dynamic models of signaling networks.

Proteomics

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans

pH Tunes the DNA Repair Efficiency and Strand Preference of the AlkB Family Enzymes.

AlkB-family Fe(II)/2-oxoglutarate-dependent dioxygenases repair alkylated nucleic acid lesions through oxidative dealkylation and play important roles in genome maintenance. 1-Methyl-2'-deoxyadenosine (1mA) and 3-methyl-2'-deoxycytidine (3mC) are well-established substrates of AlkB, ALKBH2, and ALKBH3. Although these enzymes have been extensively studied, the influence of proton concentration (pH) on their catalytic behavior and strand preference remains poorly defined. Here, we systematically examined how pH modulates the activity of the prototypical bacterial AlkB and the human homologues ALKBH2 and ALKBH3 using defined DNA substrates in both single-stranded (ssDNA) and double-stranded (dsDNA) contexts containing 1mA and 3mC lesions. Across a broad pH range, all three enzymes mainly exhibit bell-shaped activity profiles with distinct optima. The prevailing view in the field is that AlkB preferentially repairs these lesions in ssDNA, ALKBH2 favors dsDNA, and ALKBH3 prefers ssDNA. However, our results demonstrate that pH influences the catalytic efficiency and strand utilization in a substrate- and enzyme-dependent manner. AlkB maintains a consistent ssDNA preference for 3mC but exhibits variable strand preference for 1mA at different pH values. ALKBH2 retains a strong dsDNA preference for 1mA across all conditions but shows a clear pH-dependent strand switch for 3mC, favoring ssDNA under acidic conditions and preferring dsDNA at neutral to alkaline pH conditions. In contrast, ALKBH3 consistently favors ssDNA for 3mC but exhibits pH-dependent strand preference for 1mA. Our results show that the reported strand preferences largely hold at pH 7.0-8.0 but are not complete, as strand utilization and pH optima vary by enzyme and substrate. The observations demonstrate that proton availability strongly influences AlkB-family catalysis and is an important factor in how these enzymes process damaged DNA. These findings may also aid the optimization of AlkB-based protein engineering and sequencing technologies.

Hydrogen-Ion Concentration

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning

Mining metagenomes from extremophiles as a resource for novel glycoside hydrolases for industrial applications.

The exploration of metagenomes from extremophiles has emerged as a promising approach for discovering novel glycoside hydrolases (GHs) with potential industrial applications. Extremophiles, which thrive in harsh conditions such as high salinity, extreme temperatures, and acidic or alkaline environments, produce enzymes naturally adapted to function under these conditions. This unique adaptability makes them highly desirable for industrial processes requiring robust and efficient biocatalysts. These biocatalysts reduce reliance on harsh chemicals and energy-intensive processes, contributing to greener industrial operations. This review underscores the power of metagenomics in bypassing the need to culture large libraries of extremophiles in the lab. High-throughput sequencing and bioinformatics enable the identification of novel GH-encoding genes directly from environmental DNA. While metagenomic mining has yielded promising results, challenges such as the expression of extremophile-derived genes in mesophilic hosts, low activity yields, and scalability remain. Advances in synthetic biology and protein engineering could address these bottlenecks, enabling more efficient utilization of GHs. Additionally, integrating machine learning for predictive functional annotation may accelerate the identification of high-value candidates.

Glycoside Hydrolases

Towards mechanistic models of mutational effects: Deep learning on Alzheimer's Aβ peptide.

Deep Mutational Scanning (DMS) has enabled multiplexed measurement of mutational effects on protein properties, including kinematics and self-organization, with unprecedented resolution. However, potential bottlenecks of DMS characterization include experimental design, data quality, and depth of mutational coverage. Here, we apply deep learning to comprehensively model the mutational effect of the Alzheimer's Disease associated peptide Aβ42 on aggregation-related biochemical traits from DMS measurements. Among tested neural network architectures, Convolutional Neural Networks and Recurrent Neural Networks are found to be the most cost-effective models with high performance even under insufficiently-sampled DMS studies. While sequence features are essential for satisfactory prediction from neural networks, geometric-structural features further enhance the prediction performance. Notably, we demonstrate how mechanistic insights into phenotype may be extracted from the neural networks themselves suitably designed. This methodological benefit is particularly relevant for biochemical systems displaying a strong coupling between structure and phenotype such as the conformation of Aβ42 aggregate and nucleation, as shown here using a Graph Convolutional Neural Network (GCN) developed from the protein atomic structure input. In addition to accurate imputation of missing values (which here ranged up to 55% of all phenotype values at key residues), the mutationally-defined nucleation phenotype generated from a GCN shows improved resolution for identifying known disease-causing mutations relative to the original DMS phenotype. Our study suggests that neural network derived sequence-phenotype mapping can be exploited not only to provide direct support for protein engineering or genome editing but also to facilitate therapeutic design with the gained perspectives from biological modeling.

Alzheimer's disease

In silico genome mining and characterization of putative horse feces-derived bacterial phytases as potential monogastric animal feed additive candidates.

Phytic acid exerts a significant antinutritional effect in poultry, swine, and fish, which can be mitigated by supplementing monogastric feeds with efficient microbial phytases. Accordingly, mining bacterial genomes for novel phytases represents a strategic computational approach to identifying candidates for improving monogastric animal nutrition. In this study, 162 bacterial genomes associated with horse feces were systematically mined using an in silico pipeline to identify and characterize putative phytases.A total of 69 non-redundant sequences were identified and classified as histidine acid phytase (HAPhy) or protein tyrosine phosphatase-like phytase (PTPLPhy). HAPhys were detected in the genomes of Escherichia coli, Klebsiella pneumoniae, Salmonella enterica, Acinetobacter baumannii, and Cutibacterium equinum, whereas PTPLPhys were found in K. pneumoniae, Limosilactobacillus reuteri, Pediococcus acidilactici, Bifidobacterium pseudolongum, and Prescottella equi. Principal component analysis identified glucose-1-phosphatase (CAJ1242485.1) and bifunctional acid phosphatase (NHR17779.1) as the HAPhy candidates exhibiting the most favorable predicted physicochemical properties for potential feed applications. Similarly, among the PTPLPhys, protein tyrosine phosphatase (UNQ40438.1) and a hypothetical protein (CAJ1246072.1) showed the most favorable computational profiles. Biosafety analysis identified potential virulence factors, indicating that sources should be screened prior to feed application. High-quality AlphaFold2 models were obtained for these phytases (90.9-97.2). Molecular docking analysis showed that NHR17779.1 exhibited the strongest binding to phytic acid, whereas CAJ1246072.1 demonstrated the weakest interaction. Overall, this study identifies the horse fecal microbiota as a diverse source of putative phytases that may serve as promising targets for genetic and protein engineering; however, further in vitro and in vivo studies are essential to validate the enzymatic activity and industrial efficacy of these computational candidates.

Bacterial phytase

From feasibility to predictability: prime editing redefines precision breeding in plants.

Originally developed in mammalian systems as a genome editing strategy without double-strand breaks, prime editing (PE) has been adapted for precise genome modifications. However, its deployment revealed key limitations, including reduced efficiency, strong locus dependency, low germline transmission, and somatic chimerism. Consequently, diverse PE variants have emerged, resulting in fragmented landscape of architectures with context-dependent and inconsistent performance. This review consolidates these advances and outlines emerging design principles behind plant PE systems. It evaluates optimization strategies at multiple levels, discusses their applications in monocots and eudicots, and highlights persistent bottlenecks and future directions, including AI-guided protein engineering and improved delivery strategies. These advances position PE as a rapidly evolving platform toward enabling precision breeding in plants.

cis-regulatory engineering

Treatment of a severe vascular disease using a bespoke CRISPR-Cas9 base editor in mice.

Pathogenic missense mutations in the alpha actin isotype 2 (ACTA2) gene cause multisystemic smooth muscle dysfunction syndrome (MSMDS), a genetic vasculopathy that is associated with stroke, aortic dissection and death in childhood. Here we perform mutation-specific protein engineering to develop a bespoke CRISPR-Cas9 enzyme with enhanced on-target activity against the most common MSMDS-causative mutation ACTA2 R179H. To directly correct the R179H mutation, we screened dozens of configurations of base editors to develop a highly precise corrective A-to-G edit with minimal deleterious bystander editing that is otherwise prevalent when using wild-type SpCas9 base editors. We create a murine model of MSMDS that shows phenotypes consistent with human patients, including vasculopathy and premature death, to explore the in vivo therapeutic potential of this strategy. Delivery of the customized base editor via an engineered smooth muscle-tropic adeno-associated virus (AAV-PR) vector substantially prolongs survival and rescues systemic phenotypes across the lifespan of MSMDS mice, including in the vasculature, aorta and brain. Our results highlight how bespoke mutant-specific CRISPR-Cas9 enzymes can improve mutation correction with base editors.

Animals

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

Genomic prospecting and biochemical characterization of a novel thermostable 3-quinuclidinone reductase from hot spring metagenomes for efficient biocatalysis.

This study presents the discovery and characterization of a novel thermophilic 3-quinuclidinone reductase (ScQR) identified through metagenomic mining of hot spring environments. ScQR, a member of the short-chain dehydrogenase/reductase (SDR) superfamily, was heterologously expressed in Escherichia coli, and its catalytic properties were systematically characterized. The enzyme demonstrates exceptional thermal stability, retaining 86% of its activity after 48 hours at 70°C. Furthermore, K+ and Mg²+ ions significantly enhanced ScQR's activity at specific concentrations. Structural analysis revealed that ScQR adopts a typical SDR fold with a conserved catalytic triad (S141-Y155-K159), and it is NAD(H) dependent. Enzyme assays indicated that ScQR is highly stereoselective for (R)-3-quinuclidinol, with no activity against its enantiomer, (S)-3-quinuclidinol. The enzyme exhibits optimal activity at pH 9 and 85°C, making it a promising candidate for industrial applications requiring high thermal stability. Molecular dynamics simulations further revealed that ScQR preserves global structural integrity up to 360 K, whereas higher temperatures induce destabilization, predominantly in the C-terminal region and residues 95-100. In addition, structure-guided computational design enabled by LigandMPNN and UniKP yielded three ScQR variants with improved substrate affinity and catalytic efficiency while maintaining the overall fold and function. This work underscores the power of metagenomics with structure-driven protein design in discovering novel enzymes with unique catalytic properties from extreme environments and establishes ScQR as a promising biocatalyst for biotechnological and pharmaceutical applications.IMPORTANCEThis study reports the discovery of ScQR, a novel thermophilic 3-quinuclidinone reductase identified via metagenomic mining. ScQR represents one of the most heat-resistant members of the SDR superfamily discovered to date, maintaining 86% activity after 48 hours at 70°C. These findings establish ScQR as a robust biocatalyst for high-temperature pharmaceutical applications and demonstrate a scalable workflow for optimizing enzymes from extreme environments, offering significant value to the fields of biocatalysis and protein engineering.

computational design

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms

Engineering an inducible leukemia-associated fusion protein enables large-scale ex vivo production of functional human phagocytes.

Ex vivo expansion of human CD34+ hematopoietic stem and progenitor cells remains a challenge due to rapid differentiation after detachment from the bone marrow niche. In this study, we assessed the capacity of an inducible fusion protein to enable sustained ex vivo proliferation of hematopoietic precursors and their capacity to differentiate into functional phagocytes. We fused the coding sequences of an FK506-Binding Protein 12 (FKBP12)-derived destabilization domain (DD) to the myeloid/lymphoid lineage leukemia/eleven nineteen leukemia (MLL-ENL) fusion gene to generate the fusion protein DD-MLL-ENL and retrovirally expressed the protein switch in human CD34+ progenitors. Using Shield1, a chemical inhibitor of DD fusion protein degradation, we established large-scale and long-term expansion of late monocytic precursors. Upon Shield1 removal, the cells lost self-renewal capacity and spontaneously differentiated, even after 2.5 y of continuous ex vivo expansion. In the absence of Shield1, stimulation with IFN-γ, LPS, and GM-CSF triggered terminal differentiation. Gene expression analysis of the obtained phagocytes revealed marked similarity with naïve monocytes. In functional assays, the novel phagocytes migrated toward CCL2, attached to VCAM-1 under shear stress, produced reactive oxygen species, and engulfed bacterial particles, cellular particles, and apoptotic cells. Finally, we demonstrated Fcγ receptor recognition and phagocytosis of opsonized lymphoma cells in an antibody-dependent manner. Overall, we have established an engineered protein that, as a single factor, is useful for large-scale ex vivo production of human phagocytes. Such adjustable proteins have the potential to be applied as molecular tools to produce functional immune cells for experimental cell-based approaches.

Humans