SEED | 给基因编辑器留住骨架,改写序列 SEED | Keep the fold, rewrite the genome editor AI-assisted · reviewed
加州大学伯克利分校 Innovative Genomics Institute 的 Petr Skopintsev、Isabel Esain-Garcia、Evan C. DeTurk 与 Jennifer A. Doudna 团队近期在 Science 报道,将 ESM-IF1 结构逆折叠与序列保守性、蛋白-核酸共进化约束结合,设计出一组非天然小型 RNA 引导核酸酶 SynTnpB;这些变体不仅与天然序列显著分化,还在细菌、拟南芥原生质体和人细胞中保留或超过野生型编辑活性,并用 cryo-EM 解释了部分人工残基如何稳定动态 RNA-DNA 界面。

为什么复杂核酸酶很难用生成模型重写
RNA 引导核酸酶不是一个只要折叠正确就能工作的静态蛋白。它必须识别 guide RNA、搜索特定 DNA motif、形成 RNA-DNA heteroduplex,再通过多个构象状态完成切割。普通序列语言模型可以生成活性蛋白,但往往仍与天然序列高度相似;纯结构逆折叠能扩大序列空间,却可能改坏催化位点或核酸接触残基。
作者选择 TnpB 作为测试平台。TnpB 是 CRISPR-Cas12 的小型转座子祖先,蛋白约 46 kDa,天然具备 RNA 引导 DNA 切割能力,也因体积小而适合有限载荷的递送系统。真正的问题不是能否给它换几个氨基酸,而是能否大幅改写序列,同时保留多结构域协作与构象运动。
真正的新意:让结构模型服从进化留下的功能边界
ESM-IF1 单独使用时,能够生成预测折叠接近 ISDra2 TnpB、且仅有约 50% 至 60% 序列同一性的蛋白,也能保留 RuvC 的 DED 催化三联体;但它会误改 TAM 识别和 RNA/DNA 接触位点。作者因此不让模型完全自由生成,而是用两个维度构建 residue mask:一是天然 TnpB 多序列比对得到的位置保守性 Ci,二是 Potts/GREMLIN 模型从 TnpB-RNA 或 TnpB-DNA 配对序列中估计的共进化耦合强度 σi。
超过阈值的关键残基保留为野生型,其余位置交给 ESM-IF1 生成。研究又把 TnpB 拆成 DNA 识别的 REC lobe 与 RNA 结合/催化的 NUC lobe,分别设计后再组合筛选。这不是“AI 从零发明一个核酸酶”,而是用结构扩大搜索空间,再用进化信息划出不可轻易跨越的功能边界。
数据强在从 1,980 个组合走到跨系统与结构验证
作者先从每个条件下 10,000 条 ESM-IF1 生成序列中计算 consensus,再构建 44 个 REC lobe 与 45 个 NUC lobe 的全组合,共筛选 1,980 个融合体。细菌 ccdB survival assay 中,466 个组合出现可检测 enrichment,占 24%;按论文表述,约 8% 的设计活性超过野生型。结果也揭示不对称容忍度:单参数条件下,16 个 NUC 变体中有 13 个保持活性,而 REC lobe 对序列改写明显更敏感。
9 个兼顾多样性与活性的 SynTnpB 进入细胞验证。在 HEK293T 的 BFP knockout 模型中,野生型平均编辑率为 28%,v1 与 v5 分别达到 46% 和 50%。在 RUNX1、NIBAN1、EMX1、AGBL1 四个内源位点上,多数变体接近野生型;v1 与 v5 在 EMX1 分别达到野生型的 3.8 倍和 3.1 倍。9 个变体还在拟南芥原生质体的 4 个 AtPDS3 靶点测试,v1 在几乎所有靶点上优于野生型。
结构层面的闭环来自 v7。按论文排除固定的无序 C 端计算,它与野生型只有 77% 序列同一性,85 个位置由模型生成,却在 RUNX1 位点达到 44% 编辑。作者以 2.8 Å cryo-EM 解析 TAM-bound 与 R-loop-formed 两种状态,并用 reversal mutagenesis 支持多个人工残基的功能贡献。
最重要的一点:可以设计的不是一张结构,而是一组运动状态
最有价值的结论不是“AI 找到一个更强编辑器”,而是复杂蛋白的构象运动也可能在大幅改写序列后被保留下来。v7 中的人工正电残基增强了 RNA-蛋白界面的电势,并在 seed pairing、RNA pseudoknot、完整 heteroduplex 与 lid/bridge helix 区域形成新的接触网络。
尤其重要的是,cryo-EM 捕获到此前未报道的 TAM-bound 中间态:DNA 已经结合,但 guide RNA 与 target DNA 只形成一个 seed base pair。随后进入完整 R-loop 状态时,模型生成的残基仍能配合 lid domain 和 bridge helix 的运动。换句话说,设计目标不应只是“最后折成什么样”,还应包括蛋白在工作过程中能否沿着正确的构象路径移动。
这也解释了为什么进化约束有价值。静态结构未必看见短暂接触,而天然序列中的保守性与共进化信号,可能记录了这些瞬时状态对功能的长期选择压力。
批判性地读:活性提升与可用编辑器之间还有三道关
第一,筛选的底层 scaffold 仍是 ISDra2 TnpB,关键残基由野生型 mask 保留,C 端也固定不变。这是一种强有力的半理性生成策略,但距离完全 de novo 的 RNA 引导核酸酶仍有明显距离。论文证明的是“在远离天然序列的同时保住功能”,不是已经自由设计出任意新机制。
第二,活性与特异性没有同步改善。v1 的 genome-wide specificity 与野生型相近,但 v5 和 v7 出现更多可检测脱靶位点。v7 的体外切割动力学还略低于野生型。更高 on-target activity 若伴随更宽松的识别,不能简单等价为更好的编辑器。
第三,验证仍限于细菌筛选、HEK293T 细胞和拟南芥原生质体。研究没有展示哺乳动物体内递送、组织特异编辑、免疫原性、长期安全性或治疗性终点。TnpB 的小体积有递送吸引力,但 SynTnpB 的序列新颖性也可能改变表达、稳定性与免疫风险。
论文还披露多位作者已就相关工作申请专利,Doudna、Jacobsen 与 Banfield 等作者与多家基因编辑、农业生物技术或投资机构存在公司及顾问关系。这些关系已公开,不削弱实验本身,但未来平台比较仍需要独立复现与统一 benchmark。
下一步要优化的不是单一分数,而是一组互相牵制的指标
接下来应把活性、特异性、TAM 范围、热稳定性、guide 兼容性、递送表达和免疫原性纳入同一多目标设计框架。对 REC 与 NUC lobe 分开设定约束是一个可推广原则:REC 负责精细的 DNA 搜索和识别,似乎更怕改;NUC lobe 容忍更大的序列变化。未来可利用这种不对称性,先在高容忍模块扩大探索,再对识别模块做更谨慎的局部优化。
更关键的验证是把 SynTnpB 带入哺乳动物体内,在多组织、多位点和长期随访中比较 on-target、off-target、递送效率与免疫反应。若这些指标能够同时过关,这项工作才会从“复杂酶可以被重新设计”的方法学证明,走向真正可用的小型编辑器平台。
Yang 的信号评级:High
轴一,信号强度:High。 论文把生成模型、进化约束、1,980 组合筛选、细菌/人/植物验证、全基因组脱靶分析、2.8 Å cryo-EM 与反向突变串成了完整证据链。它不仅给出活性变体,也解释了人工残基如何支持多构象核酸结合。
轴二,技术成熟度:Medium。 平台已越过纯计算与单一筛选,证明非天然 SynTnpB 可在多种细胞系统工作;但还缺哺乳动物体内递送、长期安全性和系统性特异性优化,最强变体也并非在所有指标上优于野生型。
一句话总结:真正被扩大的不只是 TnpB 的序列空间,而是我们设计会运动、会识别、会切割的复杂蛋白的能力。
Petr Skopintsev, Isabel Esain-Garcia, Evan C. DeTurk, Jennifer A. Doudna and colleagues at UC Berkeley’s Innovative Genomics Institute report in Science a design strategy that combines ESM-IF1 structure-conditioned inverse folding with sequence conservation and protein-nucleic-acid coevolution. The resulting non-natural minimal RNA-guided nucleases, termed SynTnpBs, diverge substantially from natural sequences yet retain or exceed wild-type activity in bacteria, Arabidopsis protoplasts and human cells; cryo-EM further shows how designed residues stabilize a dynamic RNA-DNA interface.

Why complex nucleases are difficult to rewrite
An RNA-guided nuclease is not a static protein that works once its final fold is correct. It must recognize a guide RNA, search for a DNA motif, form an RNA-DNA heteroduplex and move through several conformational states before cleavage. Sequence language models can generate active proteins, but their outputs often remain close to natural sequences. Pure structure-conditioned inverse folding explores farther, but can corrupt catalytic residues or nucleic-acid contacts.
The authors chose TnpB as the test platform. TnpBs are small transposon-encoded ancestors of CRISPR-Cas12, with a protein mass of roughly 46 kDa and natural RNA-guided DNA-cleavage activity. Their compact size is attractive for delivery systems with limited cargo. The real challenge is therefore not to change a few amino acids, but to rewrite a large fraction of the sequence without breaking multidomain coordination and conformational motion.
What is truly new: letting evolution define the functional boundary
On its own, ESM-IF1 generated proteins predicted to retain the ISDra2 TnpB fold at only about 50 to 60% sequence identity and preserved the RuvC DED catalytic triad. It nevertheless altered residues involved in TAM recognition and RNA/DNA contacts. The authors therefore constrained generation with two residue-level signals: positional conservation, Ci, from natural TnpB multiple-sequence alignments; and coevolutionary coupling strength, σi, estimated with Potts/GREMLIN models from paired TnpB-RNA or TnpB-DNA sequences.
Residues above chosen thresholds remained fixed at their wild-type identities, while ESM-IF1 generated the remaining positions. The team also split TnpB into a DNA-recognition REC lobe and an RNA-binding/catalytic NUC lobe, designed them separately, and recombined them experimentally. This is not an AI system inventing a nuclease from nothing. Structure expands the search space, while evolution marks the functional boundaries that should not be crossed casually.
Where the evidence is strongest: 1,980 designs, three cell systems and structure
The authors computed consensus sequences from 10,000 ESM-IF1 generations per condition, then combined 44 REC lobes with 45 NUC lobes to screen 1,980 fusions. In a bacterial ccdB survival assay, 466 combinations had detectable enrichment, or 24% of the library; in the paper’s wording, about 8% of designs exceeded wild-type activity. The screen also revealed asymmetric tolerance. Under single-parameter conditioning, 13 of 16 NUC variants remained active, whereas the REC lobe was much more sensitive to sequence divergence.
Nine SynTnpBs that balanced activity and diversity advanced to cellular testing. In a HEK293T BFP knockout assay, wild-type editing averaged 28%, while v1 and v5 reached 46% and 50%. Across four endogenous loci - RUNX1, NIBAN1, EMX1 and AGBL1 - most variants were comparable to wild type; at EMX1, v1 and v5 reached 3.8-fold and 3.1-fold the wild-type level. The nine variants were also tested at four AtPDS3 sites in Arabidopsis protoplasts, where v1 outperformed wild type at nearly every target.
Variant v7 closed the structural loop. Using the paper’s calculation that excludes the fixed unstructured C terminus, it had only 77% sequence identity to wild type and 85 model-generated positions, yet reached 44% editing at RUNX1. The authors determined TAM-bound and R-loop-formed structures at 2.8 Å and used reversal mutagenesis to support functional contributions from multiple designed residues.
The most important point: the target is an ensemble of moving states
The paper’s most important result is not simply that AI found a stronger editor. It shows that conformational motion in a complex enzyme can survive substantial sequence rewriting. In v7, designed basic residues increased positive electrostatic potential at the RNA-protein interface and created new contacts around seed pairing, the RNA pseudoknot, the full heteroduplex, and the lid and bridge-helix regions.
Cryo-EM captured a previously unreported TAM-bound intermediate in which DNA is bound but the guide RNA and target DNA have formed only one seed base pair. As the complex proceeds to the full R-loop state, model-generated residues remain compatible with lid-domain and bridge-helix motion. Protein design should therefore optimize not only the final structure, but whether the protein can travel through the correct functional trajectory.
This is also why evolutionary constraints matter. Static structures may miss transient contacts, whereas conservation and coevolution can preserve a long-term record of selection on those short-lived functional states.
How to read it critically: three gaps remain before a usable editor
First, the design still begins from the ISDra2 TnpB scaffold. Critical residues were preserved with wild-type masks, and the unstructured C terminus was kept fixed. This is a powerful semirational generative strategy, but it is not a fully de novo RNA-guided nuclease. The study shows that function can survive movement far from natural sequences, not that arbitrary new mechanisms can already be designed.
Second, activity and specificity did not improve together. Variant v1 had genome-wide specificity comparable to wild type, whereas v5 and v7 had more detectable off-target sites. V7 also showed moderately slower in vitro cleavage kinetics than wild type. Higher on-target activity is not automatically a better editor if target recognition becomes more permissive.
Third, validation remains limited to bacterial selection, HEK293T cells and Arabidopsis protoplasts. The study does not test mammalian in vivo delivery, tissue-specific editing, immunogenicity, long-term safety or therapeutic endpoints. Compact size is attractive for delivery, but sequence novelty may also alter expression, stability and immune recognition.
The paper discloses patent filings covering aspects of the work, and several authors have company, advisory or investment relationships in genome editing, agricultural biotechnology and related fields. These interests are transparent, but future platform comparisons will still need independent replication and standardized benchmarking.
The next step is multi-objective design, not one higher score
Future work should optimize activity, specificity, TAM range, thermal stability, guide compatibility, delivery expression and immunogenicity together. Treating the REC and NUC lobes differently is a potentially general principle: the REC lobe performs finely tuned DNA search and recognition and appears less tolerant of change, whereas the NUC lobe accepts greater sequence divergence. Exploration can be broader in tolerant modules and more conservative in recognition modules.
The decisive validation will be mammalian in vivo testing across tissues and targets, with longitudinal measurements of on-target editing, off-target activity, delivery and immune responses. If these constraints can be satisfied simultaneously, this work may move from a proof that complex enzymes can be redesigned to a practical platform for compact genome editors.
Yang’s signal rating: High
Axis 1, signal strength: High. The paper connects generative modeling, evolutionary constraints, a 1,980-combination screen, bacterial/human/plant validation, genome-wide off-target profiling, 2.8-Å cryo-EM and reversal mutagenesis. It provides active proteins and a mechanism for how designed residues support multistate nucleic-acid binding.
Axis 2, technical maturity: Medium. The platform has moved beyond computation and a single screening assay, and non-natural SynTnpBs work across several cell systems. It still lacks mammalian in vivo delivery, long-term safety assessment and systematic specificity optimization, and the strongest variants are not superior to wild type on every axis.
One-sentence summary: The work expands not only TnpB sequence space, but our ability to design complex proteins that must move, recognize and catalyze.