Yang Liu

← Thinking

AI 从头设计的 CRISPR 剪刀:赢的是酶工程那一段最便宜的活 AI-Designed CRISPR Scissors: The Win Is in the Cheap-to-Validate Enzyme Step

Jennifer Doudna 的实验室在 Science 上发表了一件在这个领域里迟早会发生、但真发生时仍值得停下来看的事:他们用 AI 设计出自然界里不存在的 RNA 引导核酸酶,其中最好的变体在人细胞里的编辑效率反超天然酶、脱靶还更低。基因编辑和蛋白从头设计这两条线,第一次在一把“分子剪刀”上正式合流。

判断可以先放在最前面:这项工作真正的信号,不在“AI 造出了新酶”,而在论文标题自己承认的那个词——structure and evolution-guided design。AI 赢的是酶工程这一段最便宜、最能被客观验证的活;而从人细胞里的漂亮数字,到人体内的安全有效,中间那道最贵的门,它一步都还没碰。

把基因编辑接到蛋白从头设计这条线上

过去二十年,扩充 CRISPR 工具箱主要靠“去自然界里找”:从细菌、古菌、噬菌体的基因组里挖新的 Cas 家族,再改造。蛋白从头设计是另一条线——先由 David Baker 这类实验室证明可以用模型凭结构生成自然界没有的蛋白,2024 年 Profluent 又把它推到基因编辑门口,用蛋白语言模型生成了被媒体称作“CRISPR 版 ChatGPT”的 OpenCRISPR-1。

Doudna 这篇把两条线正式接上了。他们没有去大海捞针,而是拿一个已经研究得很透的骨架——TnpB,一种极其紧凑、被认为是 Cas12 进化祖先的最小 RNA 引导核酸酶——用一个 inverse folding(逆折叠)模型 ESM-IF1 反向推断:给定想要的三维结构,序列应该长什么样。关键的一步是,他们没让模型自由发挥,而是加上了来自进化信息的约束,避免在那些对功能至关重要的结合区域引入破坏性突变。生成的变体被命名为 SynTnpB。

这就是“合流”的具体含义:不再把编辑酶当成只能从自然界继承的现成零件,而是当成可以在计算里被重新设计、再拿回湿实验里筛选的对象。

被 AI 改写的是一个已知骨架,不是凭空造剪刀

值得把成果说具体,因为魔鬼和诚实都在数字里。

团队一共设计了约 1,980 个候选,在细菌里筛出 466 个有可检测活性的变体,再挑其中编辑功能最强的几个拿到植物和人细胞里深测。在人细胞里,两个变体分别达到 46% 和 50% 的编辑效率,而天然 TnpB 只有 28%;在某些人类 DNA 靶点上,编辑效率最高约为天然酶的四倍,脱靶反而更低。这些“胜出”的变体并不是对天然序列的小修小补——它与 DNA、RNA 相互作用的结构域,跟天然对应物的序列一致性只有 83% 和 72%,已经相当发散。冷冻电镜进一步给出了机制解释:最发散的那个活性变体,在引导 RNA–DNA 界面上形成了新的静电和氢键相互作用。

把这些放在一起,是一个干净的进展:AI 能在一个理解得很好的骨架上,生成序列上明显偏离自然、功能上却更强的编辑酶,而且改进能被结构解释。这不是宣传,是真战果。

这类信号不止一例。几乎同时,另一项工作(enFanzor)用不依赖实验训练数据的深度学习框架改造紧凑的 Fanzor 核酸酶,把编辑效率提高约十倍、并把配套的 ωRNA 大幅缩短,在人造血干祖细胞里也有活性。两件事指向同一个结论:在“造出一个有活性的小型核酸酶”这一步上,AI 已经能突破进化给定的性能天花板。

但要看清 AI 到底动了哪一段。它动的是“设计—构建—测试”这个闭环里的设计那一步,把原本要靠经验和运气去猜的序列空间,压缩成一次有方向的生成;真正判定成败的,仍然是后面那一长串湿实验——细菌筛活性、人细胞测效率、冷冻电镜看机制。AI 收窄了搜索,没有取代验证。

标题里的 guided,才是这篇论文最该被读到的地方

看 AI×生物的进展,有一把尺子很好用:它快在能被便宜、客观、可重复验证的地方,慢在验证昂贵又主观的地方。酶的活性恰好是前者——把一个变体放进细菌或人细胞,测它切没切、切得准不准,便宜、客观、可自动化。所以 AI 在酶工程上落地最快,一点都不意外,这和抗体设计、虚拟筛选是同一类甜区。

也正因为如此,这篇论文最该被读到的,是它没有走“纯黑箱从头生成”那条更激进的路,而是老老实实叫 guided:由结构引导、由进化约束。换句话说,让模型真正跑得动、跑得准的,恰恰是人类事先塞进去的领域知识——TnpB 的结构、哪些残基不能碰、进化允许的变异边界。Doudna 传递出的判断也是这个意思(据相关报道转述,非逐字原话):AI 工具让设计快了很多,但要把这些模型用到极致,仍然需要真正懂分子机制的人。这不是谦辞,是这篇论文的方法本身在替她说话——把领域知识抽掉,模型大概率生成一堆不折叠、不结合、不切割的死序列。

对“AI 会不会取代科学家”这个问法,这篇是一个很好的反例:它不是替代了懂机制的人,而是放大了懂机制的人。

从人细胞里的高效率,到人体内的安全,隔着最贵的一段

接下来是限制,也是读这类新闻最需要保持清醒的地方。

第一,“在人细胞里高效、低脱靶”和“在人体内安全有效”是两件事,中间隔着基因治疗最难、最贵的一整段:怎么递送到目标细胞、在全基因组尺度上脱靶到底有多低、一个自然界不存在的蛋白会不会引发免疫反应、长期基因组稳定性如何。这些恰恰落在尺子昂贵的那一端,这篇论文一个都没有回答,也没打算回答。

第二,编辑效率的纪录是在便于测量的报告系统和培养细胞里刷出来的,它是“这把剪刀锋不锋利”的证据,不是“这把剪刀能不能治病”的证据。TnpB 极其紧凑本身是个真实优势——越小越好递送,尤其对 AAV 这类载量吃紧的载体——但优势要兑现,还得过前面那一整排贵门。

第三,方法的普适性还有待观察。它在 TnpB 这个被理解得很透、结构清楚的骨架上成立;换到机制更不清楚、约束更难写的蛋白上,“structure and evolution-guided”里那两个引导还剩多少,是个真问题。

判断

这是一件真进展,正好落在基因编辑与蛋白设计的交叉口上。但它的价值要放对位置:AI 压缩的是酶设计这段可廉价验证的前端,让“造一把更好的分子剪刀”从大海捞针变成有方向的生成;它没有、也没声称压缩掉后面递送、体内特异性、免疫原性、长期安全那一整段昂贵验证

所以值得这样跟踪它:该盯的不是又一个“AI 设计的编辑酶在培养皿里刷新了效率纪录”,那一端会越来越拥挤、也越来越便宜;真正的分水岭,是 AI 设计出来的编辑酶第一次干净地跨过那些贵门——在体内被验证既高效又安全。同样值得记住的是论文标题里那个 guided:它提醒的是,模型放大的是懂分子机制的人,而不是取代他们。在一个“会用 AI 做科研”正在变成基础能力的时代,这句话对每一个还在一线的人,都是好消息。

主要信源

Jennifer Doudna’s lab published something in Science that was always going to happen in this field, yet still worth stopping for when it did: they used AI to design RNA-guided nucleases that do not exist in nature, and the best of them out-edit the natural enzyme in human cells with lower off-target activity. Gene editing and de novo protein design — two separate tracks — met for the first time on a single pair of “molecular scissors.”

Put the judgment first. The real signal here is not “AI made a new enzyme.” It is the word the paper’s own title concedes: structure and evolution-guided design. AI won the cheapest, most objectively verifiable step of the whole pipeline — enzyme engineering. The expensive gate between a nice number in human cells and being safe and effective in a human body, it has not touched at all.

Wiring gene editing into the de novo protein design track

For two decades, expanding the CRISPR toolbox mostly meant prospecting nature: mining new Cas families from bacterial, archaeal, and phage genomes, then engineering them. De novo protein design was a separate track — labs like David Baker’s first showed you can generate proteins nature never made from structure alone; in 2024 Profluent pushed it to the doorstep of gene editing with OpenCRISPR-1, a protein-language-model-generated editor the press called “ChatGPT for CRISPR.”

Doudna’s paper formally joins the two tracks. They did not go fishing in the ocean. They took a well-understood scaffold — TnpB, an extremely compact minimal RNA-guided nuclease believed to be the evolutionary ancestor of Cas12 — and used an inverse-folding model, ESM-IF1, to reason backward: given the target 3D structure, what sequence should produce it? The crucial move: they did not let the model run free. They added constraints from evolutionary information to avoid destructive mutations in regions critical to function. The resulting variants are called SynTnpBs.

That is what “convergence” concretely means here: the editing enzyme is no longer treated as a finished part you can only inherit from nature, but as an object you can redesign computationally and then screen back in the wet lab.

What AI rewrote was a known scaffold, not scissors from thin air

The result is worth stating concretely, because both the devil and the honesty are in the numbers.

The team designed roughly 1,980 candidates, screened out 466 with detectable activity in bacteria, then took the few with the strongest editing into plant and human cells for deeper testing. In human cells, two variants reached 46% and 50% editing efficiency versus 28% for natural TnpB; at some human DNA targets, editing was up to ~4-fold higher than the natural enzyme, with lower off-target activity. These winners were not minor tweaks of the natural sequence — their DNA- and RNA-interacting domains share only 83% and 72% sequence identity with their natural counterparts, i.e. quite divergent. Cryo-EM supplied the mechanism: the most divergent active variant formed new electrostatic and hydrogen-bond interactions at the guide RNA–DNA interface.

The signal is not unique. Almost concurrently, a separate effort (enFanzor) used a deep-learning framework that needs no experimental training data to reengineer the compact Fanzor nuclease — roughly a tenfold gain in editing efficiency, a much-shortened ωRNA, and activity in human hematopoietic stem and progenitor cells. Both point to the same conclusion: on the single step of “making an active, compact nuclease,” AI can now break past the ceiling evolution handed us.

Put together, this is a clean advance: on a well-understood scaffold, AI can generate editors that are clearly non-natural in sequence yet stronger in function, and the improvement is structurally explainable. That is not hype; it is a real result.

But be precise about which step AI touched. It touched the design step of the design-build-test loop, compressing a sequence space you used to grope through by intuition and luck into one directed generation. What actually decided success was still the long chain of wet-lab work that followed — activity screens in bacteria, efficiency in human cells, mechanism by cryo-EM. AI narrowed the search; it did not replace validation.

The word “guided” in the title is the most honest thing in this paper

AI-for-biology progress reads cleanly through one ruler: it is fast where results can be verified cheaply, objectively, and reproducibly, and slow where verification is expensive and subjective. Enzyme activity is squarely the former — drop a variant into bacteria or human cells and measure whether it cuts, and how precisely: cheap, objective, automatable. So AI landing fastest in enzyme engineering is no surprise; it is the same sweet spot as antibody design and virtual screening.

For exactly that reason, the thing to read in this paper is that it did not take the more radical “pure black-box de novo generation” route. It is honestly called guided: structure-guided, evolution-constrained. In other words, what actually made the model work was the domain knowledge humans put in beforehand — the TnpB structure, which residues you cannot touch, the mutational boundaries evolution permits. This is also the judgment attributed to Doudna in coverage of the work (paraphrased, not verbatim): AI tools make design much faster, but getting the most out of these models still requires people who truly understand molecular mechanism. That is not modesty; the paper’s own method says it for her — strip the domain knowledge out and the model will mostly generate dead sequences that do not fold, bind, or cut.

To the question “will AI replace scientists,” this paper is a good counterexample: it did not replace the person who understands mechanism — it amplified them.

From high efficiency in human cells to safety in a human body, the most expensive stretch remains

Now the limits — the place where reading news like this most requires a clear head.

First, “efficient and low-off-target in human cells” and “safe and effective in a human body” are two different things, separated by the hardest, most expensive stretch of gene therapy: how you deliver it to the target cells, how low the off-target rate really is at genome scale, whether a protein that does not exist in nature triggers an immune response, and how stable the genome stays long-term. These sit at the expensive end of the ruler, and this paper answers none of them — nor did it set out to.

Second, the efficiency records were set in easy-to-measure reporter systems and cultured cells. They are evidence that the scissors are sharp, not evidence that the scissors can cure. TnpB being extremely compact is a genuine advantage — smaller is easier to deliver, especially for payload-limited vectors like AAV — but cashing that advantage in still means clearing the whole row of expensive gates ahead.

Third, generality is unproven. It works on TnpB, a scaffold that is well understood and structurally clear. On proteins whose mechanism is murkier and whose constraints are harder to write down, how much of the “structure and evolution-guided” leverage survives is an open question.

The judgment

This is a real advance, and it lands exactly at the gene-editing / protein-design intersection that matters here. But place its value correctly: AI compressed the cheap-to-validate front end of enzyme design, turning “build a better pair of molecular scissors” from a needle-in-a-haystack search into a directed generation. It did not — and did not claim to — compress the whole expensive stretch of delivery, in-vivo specificity, immunogenicity, and long-term safety that follows.

So here is how to track it: the thing to watch is not another “AI-designed editor breaks an efficiency record in a dish” — that end will only get more crowded and cheaper. The real watershed is the first time an AI-designed editor cleanly clears the expensive gates and is validated as both efficient and safe in vivo. And remember the “guided” in the title: it is a reminder that the model amplifies the people who understand molecular mechanism, rather than replacing them. In an era where “knowing how to do science with AI” is becoming basic literacy, that sentence is good news for everyone still working at the bench.

Sources