Yang Liu

← Thinking

虚拟细胞:数据、算力、社区都到位了,能验证的结果还没到 Virtual Cells: The Data, Compute, and Community Are Here — the Validated Results Aren't Yet

单细胞与基础模型领域的领军者之一 Fabian Theis 在 X 上分享了 Helmholtz 慕尼黑本周举办的虚拟细胞工作坊——主题是 AI 和基础模型如何重塑健康研究。这是一个方向性信号:它延续了近月 DeepMind、Arc、Chan Zuckerberg Initiative 都在押的“虚拟细胞”范式,是 AI-for-bio 的主线之一。

判断可以先放在最前面:虚拟细胞是一个值得持续跟踪的真北极星,但今天它更多是“北极星加基础设施”。把尺子对准它最核心的能力——预测一个细胞对扰动会作何反应——会看到一幅不太舒服的图景:数据、算力、社区都已经到位,能被验证的产出还没跟上。

理念:一个能在硅里做实验的细胞模型

“AI 虚拟细胞”议程被引用最多的奠基文献,是 2024 年发在 Cell 上的一篇展望——“How to build the virtual cell with artificial intelligence”,作者包括 Bunne、Quake、Regev、Theis 等人。它的核心主张是:用基础模型学出一个跨分子、细胞、多细胞尺度的通用“细胞状态表征”,再配上一组能操纵或解码这个表征的神经网络接口——这样研究者就能在计算机里“做实验”,比如预测一个细胞在基因敲除或加药之后会怎么变,而不必先做那个湿实验。

如果它成立,改变会很大:湿实验从“每一步都做”,变成“先在硅里筛,只把最值得的候选拿到台面上做”。这正是它能同时吸引顶尖实验室和最大金主的原因。

成果:谁在下注,建了什么

这条线不缺投入,而且投入是实打实的。

Arc 研究所发布了它的第一个虚拟细胞模型 STATE:用超过 1.67 亿个观测细胞加逾 1 亿个扰动细胞、覆盖 70 种人类细胞背景训练,用来预测转录组对基因、化学、细胞因子扰动的反应。它还办了 2025 年的虚拟细胞挑战赛,明确把它框成“虚拟细胞的图灵测试”:预测一个留出细胞类型(H1 人胚胎干细胞)里的单基因扰动反应。响应远超预期——来自 114 个国家的五千多人报名、一千两百多支队伍提交、三百多支完成最终提交。

CZI 也在铺基础设施:TranscriptFormer,用 1.12 亿个细胞、跨越 15 亿年演化的 12 个物种训练;rBio,一个“在虚拟细胞上做推理”的模型,用“软验证”——从模拟预测里学习,而不是每一步都跑真实验——并声称能省掉一部分湿实验;再加上庞大的 CELLxGENE 单细胞数据库,以及 2025 年 10 月一个由英伟达加速的虚拟细胞平台。再叠加像 Helmholtz 这样专门办工作坊的机构,可以看到数据、算力、人才、社区正在快速汇聚。

这些都是真进展,尤其在数据和社区动员上。但要说清楚它们证明了什么、又没证明什么。

评测之争:结果到底有多硬

最值得读的,恰恰是这个领域对自己做的诚实评测。

在自己那场挑战赛的总结里,Arc 直说:扰动预测模型“还没能在所有指标上稳定超过朴素基线”——它们在“区分不同扰动”“识别差异表达基因”这些子能力上明显变好,但核心问题(准确预测一个未见扰动会把细胞变成什么样)并没有干净地赢过简单基线。这是主办方自己的结论,不是外部唱衰者的话。

更早、更狠的一击,是 Kedzierska 等人发在 Genome Biology 上的评测:在零样本设定下测两个被引用最多的单细胞基础模型 Geneformer 和 scGPT,发现在“细胞类型聚类”这种基础任务上,它们竟然打不过“只挑高变基因(HVG)、再用 Harmony、scVI 这些老方法”——HVG 在所有指标上都赢了两个大模型。换句话说,在某些核心任务上,重度训练的基础模型还没能稳定超过一个几乎不用训练的基线。

把两件事放在一起,图景就清楚了:这个领域有世界级的数据、算力、社区和热度,但在它的旗舰任务——扰动预测——上,还没跨过“能不能赢一个笨基线”这道线。

判断

虚拟细胞是“输入远远跑在验证过的输出前面”的一个最干净的例子:数据、算力、顶尖实验室、工作坊、几千人的挑战赛,全都到位;而那件核心的事(可靠地预测一个细胞在未见扰动下会怎么变)还没被证明。

所以值得盯的信号,不是又一个基础模型发布,也不是又一场工作坊——那一端只会越来越响。真正的分水岭有两半:扰动预测能不能在分布之外(新细胞类型、新扰动)稳定且明确地赢过朴素基线;以及有没有哪一次硅里的预测,真的改变了一个真实实验的设计。在那之前,每一个“虚拟细胞”的说法都值得用同一把尺子去量:表征学习(便宜、可基准化)有真进展;预测因果(昂贵,也是真正在卖的东西)还没兑现。Theis 办一场工作坊,是这个领域往哪投的方向性信号——不是细胞已经可被模拟的证据。

主要信源

Fabian Theis, one of the leaders in single-cell and foundation-model work, shared on X the Virtual Cell Workshop held this week at Helmholtz Munich — on how AI and foundation models are reshaping health research. It is a directional signal: it continues the “virtual cell” paradigm that DeepMind, Arc, and the Chan Zuckerberg Initiative have all been betting on in recent months, one of the main threads of AI-for-bio.

Put the judgment first. The virtual cell is a real north star worth tracking, but today it is mostly north star plus infrastructure. Hold the ruler to its most central capability — predicting how a cell responds to a perturbation — and you see an uncomfortable picture: the data, compute, and community are all in place, and the validated output has not yet caught up.

The idea: a cell model you can run experiments on in silico

The most-cited founding document of the “AI virtual cell” agenda is a 2024 perspective in Cell — “How to build the virtual cell with artificial intelligence,” by Bunne, Quake, Regev, Theis, and colleagues. Its core proposal: use foundation models to learn a universal “cell-state representation” spanning molecular, cellular, and multicellular scales, paired with a set of neural-network interfaces that manipulate or decode that representation — so a researcher can “run experiments” on a computer, e.g. predict what a cell will do after a gene knockout or a drug, without first doing that wet experiment.

If it works, the change is large: wet-lab work goes from “do every step” to “screen in silico first, and only take the most worthwhile candidates to the bench.” That is why it draws both elite labs and the biggest funders at once.

The results: who is betting, and what they have built

This thread does not lack investment, and the investment is real.

Arc Institute released its first virtual cell model, STATE, trained on more than 167 million observational cells and over 100 million perturbational cells across 70 human cell contexts, built to predict transcriptomic responses to genetic, chemical, and cytokine perturbations. It also ran the 2025 Virtual Cell Challenge, framed as “a Turing test for the virtual cell”: predict single-gene perturbation responses in a held-out cell type (H1 hESC). The response far exceeded expectations — over 5,000 registrants across 114 countries, more than 1,200 teams submitting, and over 300 completing final submissions.

CZI is laying down infrastructure too: TranscriptFormer, trained on 112 million cells across 12 species spanning 1.5 billion years of evolution; rBio, a model that “reasons over virtual cells” using “soft verification” — learning from simulated predictions instead of running a real experiment at every step — and claims to bypass some wet-lab work; plus the vast CELLxGENE single-cell database, and an October 2025 virtual-cell platform accelerated with NVIDIA. Add institutions like Helmholtz convening dedicated workshops, and you see data, compute, talent, and community converging fast.

These are real advances, especially the data and community mobilization. But be clear about what they prove and what they do not.

The evaluation debate: how hard are the results, really?

The thing most worth reading is precisely the honest evaluation the field gives of itself.

In the wrap-up of its own challenge, Arc says outright that perturbation-prediction models “do not yet consistently outperform naive baselines across all metrics” — they improved markedly on sub-capabilities like discriminating between perturbations and identifying differentially expressed genes, but the central question (accurately predicting what an unseen perturbation turns a cell into) has not been cleanly won against simple baselines. That is the organizer’s own conclusion, not an outside detractor’s.

An earlier, harder blow is the evaluation by Kedzierska et al. in Genome Biology: testing the two most-cited single-cell foundation models, Geneformer and scGPT, in a zero-shot setting, they found that on a basic task like cell-type clustering the models could not beat “just pick highly variable genes (HVG) and use old methods like Harmony and scVI” — HVG outperformed both large models across all metrics. In other words, on some core tasks, heavily trained foundation models do not yet reliably beat an almost-training-free baseline.

Put the two together and the picture is clear: the field has world-class data, compute, community, and buzz, but on its flagship task — perturbation prediction — it has not yet cleared the “can it beat a dumb baseline” bar.

The judgment

The virtual cell is one of the cleanest examples of a field whose inputs can run far ahead of its validated outputs: data, compute, elite labs, workshops, a several-thousand-person challenge — all in place; while the central thing (reliably predicting how a cell changes under an unseen perturbation) has not been demonstrated.

So the signal to watch is not another foundation-model release, nor another workshop — that end will only get louder. The real watershed has two parts: whether perturbation prediction can stably and clearly beat naive baselines out of distribution (new cell types, new perturbations); and whether any in-silico prediction ever actually changes how a real experiment gets designed. Until then, every “virtual cell” claim deserves the same ruler: representation learning (cheap, benchmarkable) shows genuine progress; predicting causality (expensive, and the thing actually being sold) has not been cashed in. Theis convening a workshop is a directional signal about where the field is investing — not evidence the cell is yet simulable.

Sources