PubMed HealthSearch

SEARCH · PubMed Health

Results for “computational design”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Using Prime Editing Guide Generator (PEGG) for high-throughput generation of prime editing sensor libraries.

Prime editing enables the generation of nearly any small genetic variant. However, the process of prime editing guide RNA (pegRNA) design is challenging and requires automated computational design tools. We developed Prime Editing Guide Generator (PEGG), a fast, flexible, and user-friendly Python package that enables the rapid generation of pegRNA and pegRNA-sensor libraries. Here, we describe the installation and use of PEGG (https://pegg.readthedocs.io) to rapidly generate custom pegRNA-sensor libraries for use in high-throughput prime editing screens.

Gene Editing

Application of emerging technologies in the antiviral field.

Viral diseases pose a serious threat to global public health, agriculture, and biosecurity. Conventional antiviral strategies are often limited by an incomplete understanding of disease mechanisms, poor targeting precision, and slow response times. Emerging technologies are now reshaping the landscape of antiviral research. This review examines the roles of four key frontiers, including organoid models, gene editing, AI-driven molecular design, and synthetic biology. Organoids provide physiologically relevant platforms that model virus-host interactions and disease progression. Viral infections remain a major challenge to human and animal health, agriculture, and biosecurity. Progress in antiviral research is constrained by the complexity of viral pathogenesis, the diversity and rapid evolution of viruses, and the limited translational relevance of some traditional model systems. Recent advances in organoid technology, gene editing, artificial intelligence, and synthetic biology are expanding the toolkit available for antiviral research and development. In this review, we discuss how these four technological frontiers contribute to disease modeling, target discovery, molecular design, and translational innovation. Organoids, in particular, provide physiologically relevant systems for investigating viral infection, tissue tropism, host responses, and pathogenesis. Gene editing tools, such as CRISPR, enable precise manipulation of host and viral genomes, facilitating the development of resistant organisms and next-generation vaccine platforms. AI technologies, including AlphaFold for structure prediction and platforms for de novo protein design, address long-standing bottlenecks in structural biology and offer powerful means to engineer antiviral proteins, antibodies, and vaccine antigens. Synthetic biology, guided by the Design-Build-Test-Learn cycle, integrates computational design, genetic assembly, and functional validation into a cohesive pipeline. Together, these technologies form a synergistic workflow that spans disease modeling, target discovery, molecular design, construction, testing, and iterative optimization. This integrated approach is shifting antiviral development from traditional empirical methods toward more precise, intelligent strategies. The review also highlights ongoing challenges in integration and scalability, stressing that high-quality biological datasets and stronger interdisciplinary collaboration are essential for realizing translational potential. By presenting a cohesive view of these converging methodologies, this review offers a framework to guide the intelligent evolution of antiviral strategies in both human and animal health.

Antiviral

Genomic prospecting and biochemical characterization of a novel thermostable 3-quinuclidinone reductase from hot spring metagenomes for efficient biocatalysis.

This study presents the discovery and characterization of a novel thermophilic 3-quinuclidinone reductase (ScQR) identified through metagenomic mining of hot spring environments. ScQR, a member of the short-chain dehydrogenase/reductase (SDR) superfamily, was heterologously expressed in Escherichia coli, and its catalytic properties were systematically characterized. The enzyme demonstrates exceptional thermal stability, retaining 86% of its activity after 48 hours at 70°C. Furthermore, K+ and Mg²+ ions significantly enhanced ScQR's activity at specific concentrations. Structural analysis revealed that ScQR adopts a typical SDR fold with a conserved catalytic triad (S141-Y155-K159), and it is NAD(H) dependent. Enzyme assays indicated that ScQR is highly stereoselective for (R)-3-quinuclidinol, with no activity against its enantiomer, (S)-3-quinuclidinol. The enzyme exhibits optimal activity at pH 9 and 85°C, making it a promising candidate for industrial applications requiring high thermal stability. Molecular dynamics simulations further revealed that ScQR preserves global structural integrity up to 360 K, whereas higher temperatures induce destabilization, predominantly in the C-terminal region and residues 95-100. In addition, structure-guided computational design enabled by LigandMPNN and UniKP yielded three ScQR variants with improved substrate affinity and catalytic efficiency while maintaining the overall fold and function. This work underscores the power of metagenomics with structure-driven protein design in discovering novel enzymes with unique catalytic properties from extreme environments and establishes ScQR as a promising biocatalyst for biotechnological and pharmaceutical applications.IMPORTANCEThis study reports the discovery of ScQR, a novel thermophilic 3-quinuclidinone reductase identified via metagenomic mining. ScQR represents one of the most heat-resistant members of the SDR superfamily discovered to date, maintaining 86% activity after 48 hours at 70°C. These findings establish ScQR as a robust biocatalyst for high-temperature pharmaceutical applications and demonstrate a scalable workflow for optimizing enzymes from extreme environments, offering significant value to the fields of biocatalysis and protein engineering.

computational design

Toward life with a 19-amino acid alphabet through generative artificial intelligence design.

Because all known living organisms are made from at least 20 canonical amino acids, the feasibility of life using a more simplified alphabet remains unclear. In this work, we leveraged computational design and synthetic biology to explore building a cell from a 19-amino acid alphabet. Initial analyses suggested that isoleucine (Ile) may be dispensable, which we confirmed by directly replacing Ile residues in essential proteins in Escherichia coli. Critically, protein language models and structure-based models were necessary to redesign functional Ile-less proteins in most cases. We systematically replaced all 382 Ile residues from the ribosome and combined 21 redesigned subunits at a native genomic locus to produce a viable, evolutionarily stable cell. This work provides a roadmap to create the first 19-amino acid organism since early evolution.

Escherichia coli

Generative artificial intelligence for enzyme design and biocatalysis.

Sparked by innovations in generative artificial intelligence (AI), the field of protein design has undergone a paradigm shift with an explosion of new models for optimizing existing enzymes or creating them from scratch. After more than one decade of low success rates for computationally designed enzymes, generative AI models are now frequently used for designing proficient enzymes. Here, we provide a comprehensive overview and classification of generative AI models for enzyme design, highlighting models with experimental validation relevant to real-world settings and outlining their respective limitations. We argue that generative AI models now have the maturity to create and optimize enzymes for industrial applications. Wider adoption of generative AI models with experimental feedback loops can speed up the development of biocatalysts and serve as a community assessment to inform the next generation of models.

Biocatalysis

Creating bottom-up RNA transfer vehicles from synthetic protein assemblies.

Evolution guides biological systems to populate ecological niches, with viruses among the most successful examples of this principle. Viruses evolved over billions of years to efficiently transfer genetic information. Although viruses are highly diverse, most have converged towards remarkable similarity in the size and shape of their capsids1,2. By contrast, generative models for protein design enable the creation of protein architectures that are absent from nature3-5. Here we investigate whether protein assemblies designed by artificial intelligence can be functionalized to construct nucleic acid transport vehicles that are independent of evolutionary trajectories. By combining natural protein domains with synthetic protein assemblies, we create more than 100 bottom-up RNA transfer vehicles with unique sizes and shapes. These vehicles surpass the RNA transfer efficiency of widely used delivery vehicles by several orders of magnitude. In addition, we demonstrate that their tropism can be programmed by incorporation of computationally designed peptide binders and use them to deliver therapeutically relevant cargo RNAs into a wide range of cellular models. We show the in vivo biodistribution of one of these vehicles in a mouse at near-single-cell resolution, confirm its safety, and use it to perform a gene-editing treatment strategy for Duchenne muscular dystrophy in patient-derived cells and a pig. Our work demonstrates how proteins created by generative artificial intelligence can be harnessed for the rational engineering of RNA transport systems with the desired properties by overcoming the limitations of natural protein diversity.

Journal Article

Multitarget interactions of bisphenol A in polycystic ovary syndrome: evidence from integrated network toxicology, mendelian randomization, and molecular docking.

OBJECTIVE: To study the potential pathogenic mechanisms of bisphenol A (BPA) in polycystic ovary syndrome (PCOS) using an integrative computational strategy. DESIGN: Integrative computational study combining network toxicology, Mendelian randomization (MR), and molecular docking. SUBJECTS: For MR analysis, genetic data were sourced from large European-ancestry cohorts, including plasma protein quantitative trait loci data and genome-wide association study summary statistics for PCOS (3,045 cases and 267,780 controls). EXPOSURE: In silico exposure to BPA for target prediction; genetically predicted plasma protein levels for causal inference. MAIN OUTCOME MEASURES: Identification of overlapping targets between BPA and PCOS; functional enrichment pathways; causal effects of prioritized proteins on PCOS risk (odds ratios with 95% confidence intervals); binding affinities between BPA and core targets (kcal/mol). RESULTS: Network toxicology identified 310 overlapping targets between BPA and PCOS. Enrichment analyses revealed significant involvement in endocrine signaling, inflammatory pathways (eg, IL-17), and cellular processes. MR demonstrated that genetically elevated levels of RET, CXCL8, HTR6, MMP1, MMP9, NTRK1, and TNNI2 were significantly associated with increased PCOS risk, whereas higher PSAP and SHBG levels were protective. Molecular docking confirmed stable binding between BPA and all nine key targets, with strongest affinity for SHBG (-8.4 kcal/mol), followed by NTRK1, TNNI2, and RET. CONCLUSION: This integrative investigation suggests that BPA may contribute to PCOS pathogenesis through multitarget interactions involving inflammatory mediators, endocrine regulators, and tissue remodeling proteins. The findings provide prioritized targets and mechanistic insights for future experimental validation and environmental risk assessment.

Female

Application of Three-Dimensionally Printed Surgical Guides in Precise Sacral Tumor Excision and Defect Reconstruction.

OBJECTIVE: Precise resection of sacral tumors remains technically demanding due to their deep anatomical location and close proximity to critical neurovascular structures. Conventional freehand techniques often result in suboptimal resection margins, excessive blood loss, and compromised lumbopelvic stability. This study evaluated whether patient-specific three-dimensional (3D)-printed guiding templates improve surgical accuracy and perioperative outcomes in sacral tumor resection and reconstruction. METHODS: Nineteen patients undergoing en bloc sacral tumor resection (S1-S3 involvement) with spinopelvic reconstruction (2006-2020) were retrospectively analyzed. Patients were divided into a 3D-printing group (n&#x2009;=&#x2009;10) and a conventional freehand group (n&#x2009;=&#x2009;9). In the 3D-printing group, computer-aided design and 3D-printed templates were used for osteotomy, screw placement, and defect reconstruction. Perioperative metrics, surgical accuracy, and complications were compared between groups using Welch's t-test and the Hodges-Lehmann method; oncologic events during follow-up were recorded descriptively. RESULTS: The 3D-printing group demonstrated significantly shorter operative time (456.5&#x2009;&#xb1;&#x2009;62.36 vs. 574.44&#x2009;&#xb1;&#x2009;114.58&#x2009;min, p&#x2009;=&#x2009;0.012), reduced blood loss (4081.40&#x2009;&#xb1;&#x2009;838.99 vs. 5090.0&#x2009;&#xb1;&#x2009;1059.67&#x2009;mL, p&#x2009;=&#x2009;0.034), and fewer fluoroscopic exposures (4.2&#x2009;&#xb1;&#x2009;0.79 vs. 10.0&#x2009;&#xb1;&#x2009;1.58, p&#x2009;<&#x2009;0.001) compared with the conventional group. Osteotomy accuracy was also superior in the 3D-printing group, with significantly lower angular deviation (3.33&#xb0;&#x2009;&#xb1;&#x2009;0.45&#xb0; vs. 6.79&#xb0;&#x2009;&#xb1;&#x2009;2.16&#xb0;, p&#x2009;=&#x2009;0.0012). Postoperative complication rates were comparable (30% vs. 44.4%, p&#x2009;=&#x2009;0.649), but hospital stay was significantly shorter in the 3D-printing group (10.7&#x2009;&#xb1;&#x2009;2.71 vs. 18.11&#x2009;&#xb1;&#x2009;4.01&#x2009;days, p&#x2009;<&#x2009;0.001). CONCLUSION: Patient-specific 3D-printed guiding templates enhance precision in sacral tumor excision and reconstruction, improving surgical efficiency and perioperative safety. This computer-assisted, template-guided approach represents a valuable advancement for complex sacral oncologic surgery.

Humans

Shielding performance and clinical applicability of lead-free materials in computed tomography.

Owing to the high radiation exposure associated with computed tomography (CT) examinations and the image quality degradation caused by conventional radiation shielding materials, this study evaluated the dose reduction performance and image quality maintenance potential of a newly developed lead-free composite shielding material. This material was composed of bismuth, tungsten, tungsten carbide, aluminium, and polyurethane. Phantom-based dose measurements demonstrated that the shielding material achieved dose reduction rates ranging from 17.6% to 37.6%, depending on tube voltage. Signal-to-noise ratio (SNR), contrast-to-noise ratio (CNR), and changes in tube current-time product (mAs) under a scout-based automatic exposure control (AEC) protocol were analysed according to the presence or absence of the shielding material across regions. For the clinical evaluation, CT scans were performed on four patients. Furthermore, the images were reviewed to evaluate whether this material affected image quality. The shielding material exhibited radiation reduction levels comparable to those reported in previous studies. SNR and CNR analyses showed minor statistical variations in certain regions; however, most differences were not statistically significant, and even significant differences remained within a range that did not compromise diagnostic image quality. Under the scout-based AEC protocol, the use of the shielding material resulted in less than 1% variation in mAs values. No visually perceptible artefacts or clinically significant image quality degradation were observed. The proposed composite shielding material demonstrated the potential to mitigate some limitations of conventional shielding materials and showed preliminary clinical feasibility as an adjunctive strategy for radiation dose reduction in CT examinations.

Radiation Protection

Predicting and comparing transcription start sites in single cell populations.

The advent of 5' single-cell RNA sequencing (scRNA-seq) technologies offers unique opportunities to identify and analyze transcription start sites (TSSs) at a single-cell resolution. These technologies have the potential to uncover the complexities of transcription initiation and alternative TSS usage across different cell types and conditions. Despite the emergence of computational methods designed to analyze 5' RNA sequencing data, current methods often lack comparative evaluations in single-cell contexts and are predominantly tailored for paired-end data, neglecting the potential of single-end data. This study introduces scTSS, a computational pipeline developed to bridge this gap by accommodating both paired-end and single-end 5' scRNA-seq data. scTSS enables joint analysis of multiple single-cell samples, starting with TSS cluster prediction and quantification, followed by differential TSS usage analysis. It employs a Binomial generalized linear mixed model to accurately and efficiently detect differential TSS usage. We demonstrate the utility of scTSS through its application in analyzing transcriptional initiation from single-cell data of two distinct diseases. The results illustrate scTSS's ability to discern alternative TSS usage between different cell types or biological conditions and to identify cell subpopulations characterized by unique TSS-level expression profiles.

Transcription Initiation Site

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, &#x3008;MAE&#x3009; = 0.11 eV, and &#x3008;RMSE&#x3009; = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by &#x223c;10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature

FIERCE: reconstructing dynamic trajectories from the differentiation potency of single cells.

MOTIVATION: Since the introduction of single-cell RNA sequencing (scRNA-seq), numerous computational approaches have been developed to reconstruct dynamic cellular processes from static transcriptional profiles. These methods order cells along continuous trajectories by assessing their similarity in the gene-expression space. However, they rely on several assumptions, such as prior knowledge of the structure and directionality of the expected genealogy. These assumptions can limit their application to complex cellular systems with poorly understood developmental paths. RESULTS: To address this challenge, we introduce FIERCE (Framework for InfERence of the veloCity of Entropy), a novel computational pipeline designed to predict the changes in the differentiation potency of single cells during dynamic processes. Through a fully unsupervised approach, FIERCE enables the inference of cell lineages directly on the differentiation landscape of the biological system, thus eliminating the need for prior specification of developmental parameters. We demonstrate the efficacy of FIERCE by reconstructing three well-known mouse differentiation systems and by quantifying its accuracy on simulated data. AVAILABILITY AND IMPLEMENTATION: The FIERCE R package is available on GitHub at https://github.com/bicciatolab/FIERCE.

Cell Differentiation

REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics Research.

SUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR.

Regulon

ImmunoTar-integrative prioritization of cell surface targets for cancer immunotherapy.

MOTIVATION: Cancer remains a leading cause of mortality globally. Recent improvements in survival have been facilitated by the development of targeted and less toxic immunotherapies, such as chimeric antigen receptor (CAR)-T cells and antibody-drug conjugates (ADCs). These therapies, effective in treating both pediatric and adult patients with solid and hematological malignancies, rely on the identification of cancer-specific surface protein targets. While technologies like RNA sequencing and proteomics exist to survey these targets, identifying optimal targets for immunotherapies remains a challenge in the field. RESULTS: To address this challenge, we developed ImmunoTar, a novel computational tool designed to systematically prioritize candidate immunotherapeutic targets. ImmunoTar integrates user-provided RNA-sequencing or proteomics data with quantitative features from multiple public databases, selected based on predefined criteria, to generate a score representing the gene's suitability as an immunotherapeutic target. We validated ImmunoTar using three distinct cancer datasets, demonstrating its effectiveness in identifying both known and novel targets across various cancer phenotypes. By compiling diverse data into a unified platform, ImmunoTar enables comprehensive evaluation of surface proteins, streamlining target identification and empowering researchers to efficiently allocate resources, thereby accelerating the development of effective cancer immunotherapies. AVAILABILITY AND IMPLEMENTATION: Code and data to run and test ImmunoTar are available at https://github.com/sacanlab/immunotar.

Humans

ProgModule: A novel computational framework to identify mutation driver modules for predicting cancer prognosis and immunotherapy response.

BACKGROUND: Cancer originates from dysregulated cell proliferation driven by driver gene mutations. Despite numerous algorithms developed to identify genomic mutational signatures, they often suffer from high computational complexity and limited clinical applicability. METHODS: Here, we presented ProgModule, an advanced computational framework designed to identify mutation driver modules for cancer prognosis and immunotherapy response prediction. In ProgModule, we introduced the Prognosis-Related Mutually Exclusive Mutation (PRMEM) score, which optimizes the balance between exclusive mutation coverage and the incorporation of mutation combination mechanisms critical for cancer prognosis. RESULTS: Applying to BLCA and HNSC cohorts, ProgModule successfully identified driver modules that stratify patients into distinct prognostic subgroups, and the combination of these modules could serve as an effective prognostic biomarker. Extending our method to diverse cancers, ProgModule presented robust prognostic performance and stability across model parameters, including stopping criteria and network topology. Moreover, our analysis suggested that driver modules can predict immunotherapeutic benefit more effectively than existing signatures. Further analyses based on published CRISPR data indicated that genes within these modules may serve as potential therapeutic targets. CONCLUSIONS: Altogether, ProgModule emerges as a powerful tool for identifying mutation driver modules as prognostic and immunotherapy response biomarkers, and genes within these modules may be used as potential therapeutic targets for cancer, offering new insights into precision oncology.

Humans

Scorpion venom peptides: Novel therapeutic approaches for inflammatory and hepatic disorders.

Chronic hepatic disorders, such as metabolic dysfunction associated steatohepatitis (MASH), alcohol associated liver disease (ALD), and viral hepatitis (Hepatitis B virus [HBV]/Hepatitis C virus [HCV]), are primarily driven by persistent immune-mediated inflammation and hepatic stellate cell activation leading to fibrosis, yet conventional therapies lack tissue and molecular specificity. Scorpion venom peptides, refined through evolutionary selection, provide highly potent, target specific scaffolds capable of modulating intrahepatic inflammatory networks. Recent in vivo preclinical studies indicate that voltage gated potassium (Kv1.3) channel blocking peptides, such as BmKK2, significantly reduce macrophage activation and inhibit downstream cytokine production, effectively ameliorating diet-induced steatohepatitis and tissue scarring in murine models. Engineered hepatotropic candidates, such as Smp76 and Mucroporin-M1, demonstrate dual therapeutic functions: they neutralize extracellular Hepatitis C particles and suppress key host transcription factors necessary for Hepatitis B replication. This review systematically examines scorpion venom peptides organized by disease category, covering their historical development, structural classification into disulfide-bridged and non-disulfide-bridged families, ion channel specificity, hepatic anti-inflammatory and antiviral mechanisms, and translational challenges including nano-formulation delivery strategies and computational drug design. These target-specific peptides are ultimately positioned as promising molecular leads that may bridge targeted immunomodulation with the resolution of chronic, progressive liver injury.

Anti-inflammatory effects

Simulation of CRISPR/Cas9-mediated gene editing for the&#xa0;Vitellogenin gene in Apis mellifera.

CRISPR/Cas9 genome editing provides a powerful framework for interrogating gene function in Apis mellifera. Yet, empirical application remains challenging due to biological constraints, including haplodiploid genetics, narrow embryonic injection window, and the social rearing requirements that complicate functional validation. These constraints necessitate in silico pre-screening to maximize editing success before resource-intensive wet-lab implementation. Within the omnigenic framework, which distinguishes core regulatory genes from peripheral loci buffered by network effects, vitellogenin (Vg) represents an optimal target which is ancestrally dedicated to yolk provisioning; it has been co-opted to orchestrate diverse non-reproductive functions including longevity, stress resistance, immunity, and social behavior. We developed a computational pipeline to design a list of 57 and 56 candidate guide RNAs&#xa0;(gRNA) for targeted Vg knockout, evaluating candidate sites in both functional exons 2 and 3 based on structural accessibility and frameshift efficiency. Comparative analysis revealed complementary strengths in two top-best candidates from initial target pool of predicted gRNAs. The gRNA targeting exon 2 exhibits weaker secondary structure (&#x394;G&#x2009;=&#x2009;-0.25&#x202f;kcal/mol versus -2.10&#x202f;kcal/mol for exon 3), aligning with empirical evidence that sites with &#x394;G&#x2009;>&#x2009;-1.0&#x202f;kcal/mol achieve 2-5&#x2009;&#xd7;&#x2009;higher Cas9 binding efficiency. This site yielded moderate frameshift frequency (77.8%; 61.9 percentile). Conversely, the predicted editing outcome for the gRNA targeting exon 3, despite stronger structural constraints, demonstrated superior functional disruption metrics demonstrating very high frameshift frequency (88.3%; 95.2 percentile), high in silico editing precision, minimal microhomology-mediated repair bias, and reproducible outcomes wherein nearly all predicted indels disrupt the coding sequence. Protein structure and domain analyses further predict that frameshift edits will generate a truncated protein missing all downstream functional domains. We recommend parallel empirical validation of both exon 2 and exon 3 targets to resolve the trade-off between structural accessibility (favoring higher editing rates) and frameshift efficacy (favoring complete loss-of-function). This dual-target strategy accommodates uncertainty in in vivo performance while maximizing the probability of generating informative phenotypes. Our in silico framework enables rational CRISPR design in non-model organisms by computationally balancing biophysical accessibility with functional impact, accelerating functional genomics in species where empirical optimization faces substantial biological constraints.

Animals

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning