PubMed HealthSearch

PubMed · 40600900

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

Abstract

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yalin Hou, Yiming Li, Ruiqing Zheng, Fuhao Zhang, Fei Guo, Min Li, Min Zeng. 2025-07-01. Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.. https://doi.org/10.1093/bioinformatics%2Fbtaf385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A noncontiguous code for RNA-guided DNA recognition at the origin of CRISPR-Cas.

CRISPR-Cas provides RNA-mediated adaptive immunity, but how its first RNA-guided effector arose is unclear. In this study, we report the discovery of Viral Interference Programmable Repeat (VIPR) systems consisting of a Vipr protein ancestral to the earliest CRISPR-Cas effectors and VIPR RNAs (vrRNAs) comprising alternating GGY/NN motifs. Unlike canonical guide RNAs that pair with target nucleic acids through contiguous complementarity, vrRNAs recognize double-stranded DNA through a noncontiguous code in which the variable NN dinucleotides collectively specify a gapped target sequence. Natural vrRNA targets suggest that VIPR systems act against competing phages, and we demonstrate programmable phage defense by redirecting the complex for transcriptional repression. These results suggest that adaptive immunity originated from ancient warfare between viruses, revealing a previously unidentified logic for encoding information in sequence.

CRISPR-Cas Systems

An inducer-independent, single-plasmid CRISPR-Cas9 system for genome editing in Bacillus species.

Advances in molecular biology tools are essential for streamlining and accelerating genetic engineering of cells across industrial and academic applications. While CRISPR-Cas improves genome editing efficiency, current systems have limitations and are often host specific, which restricts their versatility. This study describes a versatile CRISPR-Cas9 system for genome editing in industrially relevant Bacillus species. By adapting the well-established pJOE8999 vector-based CRISPR-Cas9 genome editing system, we constructed an inducer-independent, broad-host-range genome editing system. It maintains the benefits of low toxicity to the target cell and the cloning host as well as the ease to use of a single-plasmid CRISPR-Cas9 system. We utilized the constitutive Sigma70-type promoter from the conserved veg gene of Bacillus, to develop and test the suitability of promoter variants of different strengths for Cas9 expression. Successful gene deletions in three different Bacillus species demonstrated the versatility of the modified system for this industrially important genus. This was further confirmed by the integration of a reporter gene fusion and the introduction of a single point mutation in the genome of Bacillus licheniformis. This one-step CRISPR-based transformation protocol developed in this study enables fast genome editing workflows with minimal hands-on time. KEY POINTS: • Editing and screening of promoter variants for balanced Cas9 expression in Bacillus. • Development of a versatile inducer-independent, single-plasmid CRISPR-Cas-based system. • Verification of the modified CRISPR-based system for genome editing in different Bacilli.

CRISPR-Cas Systems

SpacerScope: binary-vectorized, genome-wide off-target profiling for RNA-guided nucleases without prior candidate-site bias.

The precision of CRISPR/Cas systems is fundamental to their application in plant and animal biotechnology. However, comprehensive sequence-based off-target candidate discovery remains a computational bottleneck, particularly in large and complex genomes. Here we developed SpacerScope, an off-target candidate discovery framework that enables unbiased, genome-wide discovery by leveraging binary vectorization, bitwise filtering, and right-end-anchored alignment. Benchmarking against human CIRCLE-seq data demonstrated that SpacerScope recovered 100% of validated off-target sites (6142/6142), matching the sensitivity of exhaustive algorithms. Crucially, SpacerScope achieved this maximum candidate recovery while substantially reducing computational overhead. In large-genome evaluations, SpacerScope maintained low peak memory usage of 2.20 GiB and achieved substantial runtime improvements over indel-aware comparator tools, including more than 50-fold speedup relative to Cas-OFFinder 3 (544 s versus 29 185 s). Furthermore, comparative analyses in polyploid species, such as the octoploid strawberry, revealed that SpacerScope identified larger sequence-compatible candidate burdens than standard web-based design platforms. Our results establish SpacerScope as a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes. The source code and program was publicly available at https://github.com/charlesqu666/SpacerScope. Short Abstract CRISPR/Cas sequence-based off-target candidate discovery remains computationally challenging in large, repetitive, and polyploid genomes. Existing tools either miss indel-containing candidate sites or incur prohibitive runtime and memory costs. We developed SpacerScope, a binary-vectorized framework that enables unbiased, genome-wide off-target candidate discovery without pre-selected candidate sites. By integrating bitwise filtering with right-end-anchored alignment, SpacerScope recovered 100% of validated off-target sites in human CIRCLE-seq data while using only 2.20 GiB of memory and achieving more than 10-fold speedup over indel-aware alternatives. Evaluation in plant genomes, including rice and octoploid strawberry, further demonstrated SpacerScope's capacity to identify larger sequence-compatible candidate burdens overlooked by standard tools. SpacerScope thus provides a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes, supporting downstream prioritization.

CRISPR-Cas Systems