SynTnpBs: Redesigned CRISPR-Like Nucleases Using AI-Based Protein Design with Evolutionary Constraints
A paper from Jennifer Doudna's lab (Doudna won the 2020 Nobel Prize in Chemistry, alongside Emmanuelle Charpentier, for co-developing CRISPR-Cas genome editing) introduces SynTnpBs: computationally designed, non-natural variants of TnpB, a compact CRISPR-Cas12-like RNA-guided nuclease.
Redesigning an RNA-guided nuclease is a lot harder than redesigning a typical enzyme. TnpB has to coordinate guide-RNA binding, target-DNA recognition, recognition of a short activation motif (the TAM), RNA-DNA heteroduplex formation, conformational transitions, and catalytic DNA cleavage, all in one compact protein. That's a long list of constraints to preserve while still moving the sequence somewhere new.
Why structure alone wasn't enough
The team first tried unconstrained sequence design with ESM-IF1, an inverse-folding model conditioned on the TnpB structure. It reproduced the overall fold, kept the catalytic DED residues intact, and generated sequences only 50–60% identical to wild type. But without being shown the actual RNA/DNA complex, it also altered several residues involved in TAM recognition, guide-RNA binding, and RNA-DNA interface formation. You can't just hand a model the backbone and expect it to know which residues are secretly load-bearing for a completely different molecule, the RNA/DNA guide.
To fix this, the authors built a masking strategy that folds evolutionary information into inverse folding. They derived two signals from natural TnpB homologs: positional conservation (how strongly a residue is conserved across the family) and protein-nucleic acid coupling strength (which residues coevolve with specific RNA/DNA positions). Positions above a conservation or coupling threshold were fixed to wild type; everything else was redesigned by ESM-IF1. The overall workflow: TnpB structure and evolutionary sequence data go in, conserved and coevolving residues get identified and masked, ESM-IF1 generates sequences, consensus sequences get calculated, and the results get screened experimentally for activity, first in bacteria, then in human and plant cells.
A modular design lesson
The protein was also split into two lobes: REC (DNA recognition and heteroduplex positioning) and NUC (RNA binding and the catalytic core), designed somewhat independently. REC turned out to be far more fragile: only 1 of 44 REC-lobe variants stayed functional when paired with wild-type NUC, versus 13 of 16 NUC-lobe variants staying active when paired with wild-type REC. Across all 1,980 REC-NUC combinations tested, about 24% showed detectable activity, and roughly 8% of those exceeded wild-type activity in the bacterial screen. That's a useful principle for multidomain proteins generally: different regions may need very different amounts of constraint.
How well did it work?
Nine diverse, active variants (SynTnpB-v1 through v9) went on to human and plant cells. In a BFP knockout assay in HEK293T cells, wild-type TnpB hit about 28% editing; v1 reached 46% and v5 reached 50%. At endogenous loci (RUNX1, NIBAN1, EMX1, AGBL1), several variants also beat wild type, most notably v1 and v5 at EMX1 (3.8-fold and 3.1-fold over wild type, respectively). The nine variants also worked in Arabidopsis, with v1 outperforming wild-type TnpB at nearly every tested plant target.
A highly divergent variant, v7 (about 77% identity to wild type overall, 85 redesigned positions), was solved by cryo-EM in two conformational states, including a TAM-bound intermediate that hadn't been captured before for TnpB. The structures showed new electrostatic and hydrogen-bonding contacts at the RNA-DNA interface, and, more importantly, the redesigned protein still made the conformational moves needed for cleavage. The model wasn't just fitting one static shape; it was producing sequences compatible with the whole conformational cycle.
Worth knowing before you get too excited
Higher activity didn't come free. v5 and v7 showed more detectable off-target sites than wild type, while v1's specificity stayed broadly comparable; all variants kept the canonical TTGAT TAM preference. Only about 24% of the 1,980 combinatorial designs showed detectable bacterial activity, so this is still a design-build-test-learn pipeline leaning heavily on screening, not a one-shot generative solution.
The strategy has so far been demonstrated on one TnpB scaffold (ISDra2); whether it generalizes to other TnpBs, Cas proteins, or unrelated multistate enzymes is still open. The framing in some coverage, "AI-designed nuclease outperforms nature," is also worth reading carefully: no single SynTnpB was universally better across every target and property, and the designs still depend heavily on the natural fold and many fixed, evolutionarily conserved residues. This builds on evolution rather than replacing it.
For similar tools, see the BioMoDes Biomolecular Design page.
References
- Skopintsev, P., Esain-Garcia, I., DeTurk, E.C., et al. Structure and evolution-guided design of minimal RNA-guided nucleases. Science (2026).
- SynTnpBs code (GitHub).
- Preprint.
- C&EN. AI nuclease design for CRISPR.
- Fierce Biotech. Nobel laureate Jennifer Doudna enters AI protein design arena.
- 36Kr Global. Coverage of SynTnpBs.
- The Scientist. A synthetic nuclease boosts CRISPR editing efficiency.
This post is an AI-reworded, expanded version, in my own voice, of a summary I originally posted on LinkedIn.
I try my best to make the information on this website as accurate as possible.
If you find any errors in the contents of this page or any other page on this website,
I would greatly appreciate that you kindly get in touch with me at
contact[at]abeebyekeen[dot]com.
If you are interested in joining my free weekly “BioModes and Top Reads” newsletter, please subscribe below.
