Yang Liu

← Journal

SEED | UCE:把跨物种细胞放进同一张地图 SEED | UCE maps cells across datasets and species AI-assisted · reviewed

Paper
Yanay Rosen, Yusuf Roohani, Ayush Agrawal, Leon Samotorčan, ..., Stephen R. Quake & Jure Leskovec · Nature, 2026

Stanford University 的 Yanay Rosen、Yusuf Roohani 与 Stephen R. Quake、Jure Leskovec 团队联合 Tabula Sapiens Consortium 近期在 Nature 发表 Article,提出 Universal Cell Embedding (UCE):用蛋白语言模型把不同物种的 protein-coding genes 编码到共同词典,再以超过 3,600 万个细胞自监督训练一个 6.5 亿参数 transformer,使新单细胞数据无需标注、模型重训或微调即可进入同一个 1,280 维细胞空间,为跨数据集、组织和物种的细胞图谱整合提供了新的基础模型框架。

Content infographic

单细胞图谱越来越大,却仍缺一套稳定坐标系

单细胞 RNA 测序已经产生跨组织、时间、疾病和物种的海量数据,但每个数据集往往仍是一座独立小岛。基因集合不同、测序技术不同、批次效应强,同一种细胞在不同实验中可能比同批次的其他细胞看起来更远。scVI、scArches 等方法可以整合数据,却通常需要为新数据重新训练或微调;Geneformer、scGPT 等 foundation models 尝试学习通用表示,但跨新数据集和新物种的稳定性仍有限。

真正需要的不是每次重新画一张 UMAP,而是一个固定坐标系:今天训练的参考图谱、明天新测的患者样本、甚至训练时从未见过的物种,都能直接落进同一个表示空间;在这个空间中,batch 被压低,cell type、cell state、谱系和功能关系仍被保留。只有坐标系保持不变,基于它训练的细胞分类器、检索工具和疾病比较才可能跨实验复用。

UCE 因而问了一个比“能否提高某个分类准确率”更基础的问题:能否从没有 cell-type labels 的表达数据中,学习一种对数据集和物种都相对稳定、又能承载生物学关系的 universal cell representation?

真正的新意:用蛋白序列把不同物种的基因翻译成共同语言

UCE 把一个细胞视作 “bag of RNA”。对每个细胞,它按照表达量加权,从非零表达的 protein-coding genes 中有放回地抽样 1,024 个 gene tokens。每个基因不是用固定 gene ID 表示,而是先把其蛋白产物的氨基酸序列送入 150 亿参数 ESM2 protein language model,得到 5,120 维 protein embedding;多个蛋白产物时取平均,再压缩为 transformer 输入。基因按染色体和 genomic location 排列,并加入 chromosome start/end 与 CLS token。

主体是 33 层、超过 6.5 亿参数的 transformer。训练时遮蔽一部分已表达基因,让模型根据 cell embedding 与 gene embedding 判断某基因在该细胞中是否表达。整个训练不使用 cell-type 或 dataset labels,最终 CLS 输出成为 1,280 维 cell embedding。模型在 24 张 A100 80 GB GPU 上训练 40 天,数据来自 300 多个数据集和超过 3,600 万个细胞。

这个设计最关键的不是 transformer,而是 protein sequence 充当跨物种基因词典。传统跨物种整合常先寻找 homologous genes;UCE 只需要新物种的 protein-coding sequence,就能把基因投进 ESM2 的通用蛋白空间。即使某个物种从未进入训练集,它的细胞仍可被编码。这样一来,模型不必为每个基因组建立新的 one-to-one vocabulary,也不必为每个新实验改变坐标系。

作者据此建立 Integrated Mega-scale Atlas (IMA):约 3,600 万个细胞、超过 1,000 个命名 cell types、约 300 个数据集、50 种组织和 8 个物种。UCE 的核心平台承诺是:IMA 不是一次性的汇总图,而是一个可让未来数据直接接入的参照空间。

数据强在真正留出集、未见物种和多尺度生物结构

第一层证据来自训练时尚未公开的 Tabula Sapiens v2。该数据包含 581,430 个细胞、27 种组织、167 个 batches 和 162 个 cell types,且同时包含 droplet 与 plate-based 测序。UCE 在不微调的情况下,相比次优的 Geneformer,overall score 提高 13.9%,biological conservation 提高 16.2%,batch correction 提高 10.1%;表现也略优于需要 cell-type labels 和 dataset-specific training 的 scVI 与 scArches。卵巢数据中,45,757 个 10x 3′ v3.1 细胞与 3,610 个 Smart-seq2.5 细胞被混合到同一生物结构中,同时保留更清晰的 cell-type separation。

第二层是严格的跨物种零样本测试。训练数据覆盖人、小鼠、mouse lemur、斑马鱼、猪、rhesus macaque、crab-eating macaque 和 western clawed frog。对训练时未见的 green monkey,17 个 cell-type centroids 中有 13 个最近的跨物种邻居是同类细胞,放宽到前三近邻时达到 17/17;naked mole rat 为 17/24;chicken heart 的 15 类细胞中,12 类可在前两个跨物种近邻中匹配,而训练集中没有任何鸟类。与需要新数据标签并重新训练的 SATURN、SAMap 相比,UCE 在四个跨物种数据集中的三个表现更好。

第三层是表示空间自身形成了多尺度生物结构。Tabula Sapiens v2 中,90/97 个 mesoderm-derived cell-type centroids 的最近邻仍来自中胚层,endoderm 为 46/56,ectoderm 为 22/30;用 held-out cell types 预测 germ-layer origin 的 classifier 准确率超过 80%。细胞间 embedding distance 与 Cell Ontology tree distance 在前五级关系内显著相关,说明空间不只分开细胞类型,也编码了较粗到较细的层级与发育关系。

第四层是跨图谱检索的生物学案例。作者用 UCE 找到肾脏 erythropoietin-producing Norn cells,并让一个简单 logistic classifier 在原始数据中以 98.3% 复现其标签;随后把 classifier 扫过 IMA,在 13 个肾脏、6 个肺和 6 个心脏数据集中发现表达 Dcn、Lpar1、Col1a1、Cxcl12、Cfh 等 markers 的 Norn-like cells。肺部疾病数据中,这类细胞在 COPD、IPF 和对照组均存在,并呈现疾病相关 collagen 与 oxygen-sensing gene 差异。这个例子展示了固定坐标系如何把“某数据集中的新细胞”变成跨组织检索问题。

最重要的一点:模型的产品不是分类器,而是可共享的细胞坐标

UCE 最重要的贡献,不是某个 benchmark 增加了十几个百分点,而是它尝试改变单细胞分析的工作单位。传统流程是“拿到新数据—重新选基因—重新整合—重新训练—得到一个只对本项目有效的空间”;UCE 希望变成“拿到新数据—直接嵌入—与所有已有细胞比较”。如果成功,embedding 本身可以成为研究基础设施,而不是项目内部的中间文件。

这对可复用下游模型尤其关键。因为 UCE 明确不做 fine-tuning,空间坐标不会随着每次新数据而漂移;在一个参考数据集上训练的 logistic classifier,可以直接用于 Tabula Sapiens v2,甚至 green monkey。分类器、nearest-neighbour search、异常细胞检测和疾病状态检索因此可能像基因组坐标一样跨研究共享。

蛋白语言模型的引入还提供了一种有意思的多尺度先验:细胞表示建立在基因表达之上,而基因 token 又带有蛋白序列演化信息。跨物种泛化不完全依赖 RNA 表达矩阵的同源基因交集,而是借助 protein space 把进化关系带入细胞空间。这可能是 UCE 相比纯 gene-rank 或 binned-expression 模型更容易处理未见物种的根本原因。

但这仍是 universal embedding,不是完整 virtual cell。它能组织、检索和转移标签,却尚不能可靠预测基因扰动、药物响应、时间演化或跨模态状态,更没有生成一个此前不存在且可实验验证的细胞状态。把坐标系做好是重要底座,但坐标系本身不等于因果模型。

怎样批判性地读:通用性最容易被标签基准高估

第一,主要 benchmarks 仍围绕 expert-annotated cell types、batch correction 和 ontology proximity。Cell labels 的粒度和边界本来就主观;粗粒度免疫细胞或跨物种同源细胞匹配得好,不代表模型能区分连续状态、短暂过渡、微弱疾病反应或药物扰动。论文也明确承认,perturbation response 与 cross-modality integration 尚未被充分评价。

第二,IMA 的宏大规模不能全部视作独立验证。IMA 大量由训练语料构成,labels 虽未用于自监督训练,但 atlas 内的组织结构属于训练内表征。Tabula Sapiens v2 和未见物种实验提供了更重要的 out-of-sample 证据;即使如此,UCE 训练中包含 Tabula Sapiens v1 和大量相近的人类、小鼠组织,因此“新数据”不等于“新生物学分布”。

第三,UCE 仍是黑箱。模型能把细胞放到合理位置,却很难回答哪个基因、蛋白先验或 attention pattern 驱动了距离。若研究者从 embedding proximity 推断功能或机制,缺乏解释层会让 batch、共享 stress programme 或测序深度被误读成生物学同源性。作者也强调 UCE 不替代传统 batch correction;已知技术混杂仍需显式处理。

第四,输入表示有信息损失。每个细胞只抽样 1,024 个非唯一 gene tokens,表达量被转换为抽样概率,训练目标又是 binary expressed/not-expressed。这样便于跨平台稳健,却会丢失部分定量表达差异,可能压平稀有 transcript、剂量效应和细微 cell states。随机抽样还会使同一细胞在不同 seed 下得到略有差异的 embedding,尽管固定 seed 可复现。

第五,跨物种结果主要证明 coarse alignment,而不是进化机制闭环。训练数据明显偏向哺乳动物,尤其人、小鼠和 brain;与高度发散物种进行细粒度 cell-type transfer 仍困难。protein embeddings 带来强 prior,但 ESM2 本身也来自不均衡的蛋白序列语料,不能自动消除数据偏差。

最后,Norn-like cells 是有价值的 hypothesis-generation 展示,却没有功能性闭环。marker enrichment 和跨数据集复现支持“相似转录状态存在”,但 Epo transcript 本身低且常缺失,论文没有用 EPO protein、lineage tracing、hypoxia perturbation 或组织定位证明这些肺和心脏细胞就是功能性 Norn equivalents。疾病差异应被视为待验证假设,而不是新机制定论。

下一步:从稳定坐标走向可检验的细胞预测模型

最优先的验证是把 benchmark 从“恢复标签”推进到“预测未见干预”。可以在固定 UCE 空间中测试 CRISPR perturbation、药物处理、感染、分化时间序列和疾病进展,要求模型在完全留出的 donor、platform、species 和 tissue 上预测方向与幅度。对连续状态,应使用 trajectory、response vector 和 calibrated uncertainty,而不是只看 cluster purity。

第二,应把 transcriptome 与 spatial、chromatin accessibility、surface protein、morphology 和 lineage information 对齐,同时保持坐标稳定。如果每增加一种模态就必须重新训练并移动全部旧坐标,UCE 的“通用参照系”优势会被削弱;需要明确 versioning、backward compatibility 和 reference-anchor 策略。

第三,需要发展 mechanistic interpretability:追踪特定 protein embeddings、gene sampling 和 attention heads 如何影响 cell distance,用 perturbation data 检验重要特征是否真有因果意义。对于 Norn-like cells,下一步应在肺与心脏中做原位定位、低氧反应、EPO protein 检测和功能干预,验证 embedding 相似性是否对应共同生理功能。

第四,工程上需要降低推理和训练成本。原模型训练消耗 24 张 A100、40 天,1,024-token transformer 仍有二次复杂度。更高效的 state-space 或 sparse architectures、预计算 protein tokens、分层检索与小模型蒸馏,决定 UCE 能否从少数大型中心的基础设施变成普通实验室日常使用的工具。

论文已开放 UCE 代码、模型权重与 处理后数据,为独立实验室检验跨平台稳定性、数据泄漏、成本和真实生物学发现率提供了条件。下一阶段最有说服力的结果,不会是更大的 UMAP,而是 UCE 在未知扰动与实验验证中持续给出正确、可解释的预测。

Yang 的信号评级:High

轴一,信号强度:High。 UCE 提出清晰的平台级架构:以 protein language model 构建跨物种 gene vocabulary,在 3,600 万细胞上学习固定、无需微调的 cell embedding。证据覆盖严格留出的 581,430-cell atlas、多技术 batch、未见物种、发育谱系、ontology alignment 和跨组织细胞检索,并开放代码、权重与数据。

轴二,技术成熟度:Medium。 UCE 已经能作为单细胞整合、标签转移和检索工具使用,但训练与推理成本高,数据偏向人和小鼠,解释性有限,细微状态和定量表达可能丢失;对 perturbation、multimodal prediction 与新机制的能力仍缺少前瞻性功能验证。

一句话总结:UCE 把单细胞 foundation model 最重要的产品从“一个更好的分类器”变成“一套所有新细胞都能直接进入的共享坐标系”,但从地图走到可预测、可干预的 virtual cell 仍有一段关键距离。

Yanay Rosen, Yusuf Roohani, Stephen R. Quake, Jure Leskovec and colleagues at Stanford University, with the Tabula Sapiens Consortium, recently published an Article in Nature introducing Universal Cell Embedding (UCE). The system uses a protein language model to encode protein-coding genes from different species into a shared vocabulary, then self-supervises a 650-million-parameter transformer on more than 36 million cells. New single-cell data can enter the same 1,280-dimensional cell space without labels, model retraining or fine-tuning, providing a new foundation-model framework for integrating cell atlases across datasets, tissues and species.

Content infographic

Cell atlases keep growing but still lack a stable coordinate system

Single-cell RNA sequencing has produced enormous datasets across tissues, developmental time, disease and species, but each dataset often remains an isolated island. Gene sets differ, sequencing technologies differ and batch effects can be strong enough that the same cell type across experiments looks farther apart than unrelated cells within one batch. Methods such as scVI and scArches can integrate new data, but usually require retraining or fine-tuning. Foundation models such as Geneformer and scGPT seek general representations, yet stable transfer to new datasets and species remains limited.

The unmet need is not another UMAP generated anew for every project. It is a fixed coordinate system in which today’s reference atlas, tomorrow’s patient sample and even a species absent from training can be placed directly. In that space, batch should be suppressed while cell type, state, lineage and functional relationships remain. Only if the coordinates stay fixed can classifiers, retrieval tools and disease comparisons trained on them be reused across experiments.

UCE therefore asks a more foundational question than whether a model can improve one classification score: can unlabelled expression data support a universal cell representation that remains relatively stable across datasets and species while retaining biological relationships?

The key idea is translating genes across species through protein sequence

UCE treats a cell as a “bag of RNA”. For each cell, it samples with replacement 1,024 protein-coding gene tokens from non-zero expressed genes, weighted by expression. A gene is not represented by a fixed ID. Instead, the amino-acid sequence of its protein product is passed through the 15-billion-parameter ESM2 protein language model to produce a 5,120-dimensional protein embedding. Embeddings are averaged when a gene encodes multiple proteins and compressed for transformer input. Genes are ordered by chromosome and genomic position, with chromosome start/end tokens and a CLS token added.

The core model is a 33-layer transformer with more than 650 million parameters. During training, a subset of expressed genes is masked and the model uses the cell and gene embeddings to predict whether a gene was expressed in that cell. No cell-type or dataset labels are used. The final CLS output becomes a 1,280-dimensional cell embedding. Training across more than 300 datasets and 36 million cells took 40 days on 24 A100 80 GB GPUs.

The most consequential design choice is not the transformer but the use of protein sequence as a cross-species gene dictionary. Conventional cross-species integration often begins by finding homologous genes. UCE only needs the protein-coding sequences from a new species and can place its genes in the ESM2 protein space. Cells from a species absent from training can therefore be encoded without constructing a new one-to-one vocabulary or moving the cell coordinate system for every experiment.

The authors use this architecture to build the Integrated Mega-scale Atlas (IMA): roughly 36 million cells, more than 1,000 named cell types, around 300 datasets, 50 tissues and eight species. The platform claim is that IMA is not a one-off aggregate plot but a reference space into which future data can be inserted directly.

The strongest evidence combines a true holdout, unseen species and biological hierarchy

The first layer of evidence is Tabula Sapiens v2, which was unpublished when the model was trained. It contains 581,430 cells, 27 tissues, 167 batches and 162 cell types, including both droplet- and plate-based sequencing. Without fine-tuning, UCE improves over the next-best Geneformer by 13.9% on the overall score, 16.2% on biological conservation and 10.1% on batch correction. It also performs slightly better than scVI and scArches, which use cell-type labels and dataset-specific training. In the ovary data, 45,757 10x 3′ v3.1 cells and 3,610 Smart-seq2.5 cells mix into the same biological structure while retaining clearer cell-type separation.

The second layer is strict zero-shot transfer to species absent from training. Training covers human, mouse, mouse lemur, zebrafish, pig, rhesus macaque, crab-eating macaque and western clawed frog. For green monkey, the nearest cross-species neighbour matches the same cell type for 13 of 17 cell-type centroids, rising to 17 of 17 within the three nearest neighbours. Naked mole rat reaches 17 of 24. In chicken heart, 12 of 15 cell types match within the two nearest cross-species neighbours despite no bird being present in training. UCE outperforms SATURN and SAMap, which retrain with new data and labels, on three of four cross-species datasets.

The third layer is multi-scale biological organization within the space. In Tabula Sapiens v2, 90 of 97 mesoderm-derived cell-type centroids have another mesoderm-derived type as their nearest neighbour; the corresponding counts are 46 of 56 for endoderm and 22 of 30 for ectoderm. A classifier trained to predict germ-layer origin for held-out cell types exceeds 80% accuracy. Embedding distance also tracks Cell Ontology tree distance across the first five levels, suggesting that the space encodes hierarchy and developmental relationships rather than merely separating labels.

The fourth layer is a cross-atlas biological retrieval case. UCE identifies kidney erythropoietin-producing Norn cells, and a simple logistic classifier recaptures their original labels 98.3% of the time. The classifier is then applied across IMA and finds Norn-like cells expressing Dcn, Lpar1, Col1a1, Cxcl12 and Cfh in 13 kidney, six lung and six heart datasets. In lung disease datasets, these cells occur in COPD, IPF and control samples and show disease-associated differences in collagen and oxygen-sensing genes. The case demonstrates how a novel cell in one dataset can become a search query across tissues.

The product is not a classifier but a shareable coordinate system

UCE’s most important contribution is not a double-digit benchmark gain. It changes the intended unit of single-cell analysis. The conventional workflow is to obtain new data, select genes, reintegrate, retrain and produce a project-specific space. UCE aims to reduce that to obtaining new data, embedding it directly and comparing it with all prior cells. The embedding can become infrastructure rather than an internal intermediate file.

This matters for reusable downstream models. Because UCE is deliberately used without fine-tuning, its coordinates do not drift with every new dataset. A logistic classifier trained in one reference dataset can be applied directly to Tabula Sapiens v2 and even green monkey. Classifiers, nearest-neighbour search, anomalous-cell detection and disease-state retrieval can therefore become portable in a way that resembles genomic coordinates.

The protein language model also introduces a multi-scale prior: a cell representation is built from gene expression, while each gene token carries evolutionary information from protein sequence. Cross-species transfer is not restricted to the intersection of homologous genes in two RNA matrices. Instead, protein space contributes evolutionary structure to cell space. This is likely a central reason UCE handles unseen species better than pure gene-rank or binned-expression models.

UCE is nevertheless a universal embedding, not a complete virtual cell. It organizes, retrieves and transfers labels, but does not yet reliably predict genetic perturbation, drug response, temporal evolution or multimodal state, nor does it generate a previously unseen cell state that is then experimentally validated. A stable map is an important substrate, but a map is not itself a causal model.

Universality can be overestimated by label-recovery benchmarks

First, the main benchmarks still emphasize expert-annotated cell types, batch correction and ontology proximity. Cell-label granularity and boundaries are subjective. Good performance on coarse immune classes or cross-species homologous types does not establish sensitivity to continuous states, transient transitions, subtle disease responses or drug perturbations. The paper explicitly notes that perturbation response and cross-modality integration are not yet adequately evaluated.

Second, the scale of IMA should not all be treated as independent validation. Much of IMA is also training material. Its labels are held out from self-supervision, but organization within the atlas remains an in-training representation. Tabula Sapiens v2 and the unseen-species experiments provide the more important out-of-sample evidence; even there, training includes Tabula Sapiens v1 and extensive related human and mouse tissues, so a new dataset is not necessarily a new biological distribution.

Third, UCE remains a black box. It can place cells in plausible locations without explaining which gene, protein prior or attention pattern drives the distance. If researchers infer function or mechanism from embedding proximity, batch, shared stress programmes or sequencing depth may be mistaken for biological homology. The authors emphasize that UCE does not replace conventional batch correction; known technical confounders still require explicit handling.

Fourth, the input representation loses information. Each cell is represented by only 1,024 non-unique gene tokens, expression becomes a sampling probability and the training objective is binary expressed versus not expressed. This supports robustness across technologies but discards some quantitative variation, potentially smoothing rare transcripts, dose effects and fine cell states. Random sampling also gives a cell slightly different embeddings under different seeds, although a fixed seed is reproducible.

Fifth, the cross-species results mostly establish coarse alignment, not evolutionary mechanism. Training is heavily biased toward mammals, especially human, mouse and brain tissues, and fine-grained transfer across distant species remains difficult. Protein embeddings provide a powerful prior, but ESM2 itself is trained on an uneven protein-sequence corpus and cannot remove data bias automatically.

Finally, the Norn-like-cell case is a useful demonstration of hypothesis generation rather than functional closure. Marker enrichment and replication across datasets support the existence of a related transcriptional state, but Epo transcripts are sparse and often absent. The study does not use EPO protein, lineage tracing, hypoxia perturbation or spatial functional assays to show that lung and heart cells are functional Norn equivalents. Disease differences should remain hypotheses, not mechanistic conclusions.

The next step is moving from stable coordinates to testable prediction

The highest-priority benchmark is unseen intervention rather than label recovery. UCE should be tested on CRISPR perturbations, drugs, infection, differentiation time courses and disease progression, with donor, platform, species and tissue held out completely. Continuous states need trajectory, response-vector and calibrated-uncertainty metrics, not cluster purity alone.

Second, transcriptomes need to align with spatial context, chromatin accessibility, surface proteins, morphology and lineage information while preserving coordinate stability. If every new modality requires retraining and moves all historical coordinates, the universal-reference advantage weakens. Versioning, backward compatibility and stable reference anchors should be explicit parts of the platform.

Third, UCE needs mechanistic interpretability. Researchers should be able to trace how protein embeddings, gene sampling and attention influence cell distance, then test those features with perturbation data. For Norn-like cells, in situ localization, hypoxia response, EPO protein measurement and functional intervention in lung and heart are the natural next tests.

Fourth, engineering costs must fall. The original training run used 24 A100 GPUs for 40 days, while a 1,024-token transformer retains quadratic complexity. Efficient state-space or sparse architectures, precomputed protein tokens, hierarchical retrieval and distilled models will determine whether UCE becomes a routine laboratory tool rather than infrastructure available only to large centres.

The authors have released the UCE code, model weights and processed data, enabling independent laboratories to test cross-platform stability, data leakage, cost and true biological discovery rate. The next persuasive result will not be a larger UMAP; it will be correct and interpretable UCE predictions under previously unseen interventions, followed by experimental validation.

Yang’s signal rating: High

Axis 1, signal strength: High. UCE presents a clear platform architecture: a protein-language-model gene vocabulary supporting a fixed, no-fine-tuning cell embedding learned from 36 million cells. Evidence spans a true 581,430-cell holdout atlas, multi-technology batches, unseen species, developmental lineages, ontology alignment and cross-tissue cell retrieval, with code, weights and data released.

Axis 2, technical maturity: Medium. UCE can already support single-cell integration, label transfer and retrieval, but training and inference are expensive, data are biased toward human and mouse, interpretation is limited and fine cell states or quantitative expression may be smoothed. Perturbation, multimodal prediction and new-mechanism claims still lack prospective functional validation.

One-sentence summary: UCE shifts the most important product of a single-cell foundation model from a better classifier to a shared coordinate system that any new cell can enter directly, while the journey from mapping to a predictive and actionable virtual cell remains unfinished.