Yang Liu

← Journal

SEED | 抗体发现的 ML-ready 短码 SEED | An ML-ready short code for antibody discovery AI-assisted · reviewed

Paper
Deepash Kothiwal, Aaron W. Kollasch, Murali Anuganti, Nicholas Hollmer, Anita Ghosh, Roushu Zhang, Haiying Li, Steffanie B. Paul, Ruitong Li, Yvrick Zagar, Mina Abdollahi, Zachary Anderson, Filmawit Belay, Matt Salotto, Sophia Ulmer, Youssef Atef Abdelalim, Aditi Kachare, Satyendra Kumar, Mahesh Vangala, Chang Yang, Alain Chedotal, Joseph G. Jardine, Andre A.R. Teixeira, Deborah J. Moshinsky, Haisun Zhu, Shaotong Zhu, Timothy A. Springer, Debora S. Marks & Rob Meijers · Cell Systems, 2026

Institute for Protein Innovation 等机构的 Deepash Kothiwal 与通讯作者 Rob Meijers 团队近期在 Cell Systems 报道,构建了以 VH1-69 scaffold 和 CDRH3-based antigen recognition module (ARM) 为核心的 synthetic Fab yeast display library,在 10 个 cell surface antigens 上并行筛选,并用 logistic regression 从早期 sorting NGS 数据中找回被实验选择漏掉的 ROBO2N 和 PD-L2 binders,为 hybrid in silico/experimental antibody discovery 提供了一个 ML-ready 数据平台。

Content infographic

问题不是再做一个抗体库,而是让筛选数据可被学习

抗体发现已经有 phage display、yeast display、immunization、single B-cell cloning 等成熟路径。真正的瓶颈不只是能不能找到 binder,而是能不能把发现过程变成一个可规模化、可复用、可被机器学习读取的数据系统。

传统 display library 往往同时改变多个 CDR 区域。这样的多样性对发现抗体有利,但对机器学习不友好:序列空间太大,heavy/light-chain pairing 信息复杂,selection 过程又会受到表达量、增殖速度、display efficiency 和 sorting gate 的影响。一个抗体在最终 FACS round 中频率高,不一定说明它最强;一个在早期出现、后来被淘汰的 clone,也不一定没有价值。

这篇文章问的就是这个问题:能不能设计一个足够简单、足够大、又足够可测序的抗体库,让实验筛选和机器学习真正接上?

作者选择把抗原识别重点压缩到 CDRH3。CDRH3 是抗体识别中最重要、也最可变的区域之一。团队设计了一个短于 100 nt 的 ARM,覆盖 heavy-chain CDRH3 和邻近 light-chain barcode,使每一个 Fab 的关键识别模块可以通过 NGS 直接追踪。这样,筛选不只是产生候选抗体,也同时产生可供 ML 学习的 antibody-antigen interaction dataset。

新意在于 ARM:把抗原识别压成一段短码

这项研究的技术新意有三层。

第一层是库设计。作者使用 VH1-69 heavy-chain scaffold,并配对 4 条 light chains,其中 VK1-39、VK3-15、VK3-20 提供较平坦的 paratope,VK4-1 提供更 concave 的 paratope。CDRH3 长度主要覆盖 11-17 aa,位置特异性氨基酸频率来自 Observed Antibody Space 中 9.5 million unique naive B-cell CDRH3 sequences。设计时还剔除了 cysteine、methionine 和可能导致 aggregation、polyreactivity、chemical liability 的 motifs。

第二层是 ARM 表达。CDRH3 后面接 light-chain barcode,形成一个短核苷酸模块。这个模块既代表主要 paratope,也记录 heavy/light-chain pairing。相比需要完整读出六个 CDR 的复杂 library,这种 ARM 结构更适合 deep sequencing,也更适合后续用 k-mer、distance 或更复杂模型学习 selection pattern。

第三层是研究设计。作者没有只在一个靶点上做 proof-of-concept,而是在 10 个 cell surface glycoprotein antigens 上并行筛选,包括 PD-L1、PD-L2、TIGIT、LOX1、DKK1、IL-23R、DCC、ROBO1、ROBO2 和 syncytin-2。筛选流程包括 1 轮 MACS、3 轮 FACS,以及每一轮后的 ARM NGS tracking。

因此,这篇文章的新意不是“用了机器学习”本身,而是把 antibody discovery 的实验设计提前改造成 ML-compatible:抗体库、barcode、NGS、screening 和模型输入从一开始就是一套管线。

强数据来自 10 个靶点、424 个抗体和 ML 找回漏网克隆

第一组强数据来自 library scale 和 sequence quality。两个 VH1-69 sub-libraries 的 deep sequencing 读出 390 million unique CDRH3 sequences,每个 sub-library 的 molecular diversity 估计约为 2.5 x 10^8 unique clones。四个 sub-libraries 合并后,作者估计 VH1-69 library 的真实 molecular diversity 接近 1 billion unique Fabs。ARM library 的 2-mer 和 3-mer usage 与 naive B-cell CDRH3 repertoire 的 Spearman correlation 分别为 0.78 和 0.76,同时显著减少了不利 motifs。

第二组强数据来自 10 个靶点的并行筛选。作者根据最终 FACS round 中 enriched CDRH3 clusters 选择 429 条 heavy chains,最终表达出 424 个有足够产量的 IgG1 antibodies。其中 301 个通过 SEC aggregation/degradation 检查,354 个不显示对 avidin、DNA、insulin 或 Sf9 membrane preparation 的 polyreactivity,285 个同时通过 SEC 和 polyreactivity 两类检查。

第三组强数据来自功能验证。作者对 335 个 experimentally verified antibodies 做 SPR,对 317 个抗体做 cell-display titration。结果中有 103 个抗体达到 K_D < 10 nM,118 个抗体达到 EC50 < 25 nM。部分 TIGIT 和 LOX1 抗体在 flow cytometry 应用中与商业抗体相当;ROBO1、ROBO2 和 syncytin-2 代表性抗体的 melting temperature 超过 70 C;ROBO1/ROBO2 rabbit chimera antibodies 还能在小鼠胚胎神经组织 IHC 中给出清晰染色。

第四组强数据是 ML rescue。ROBO2N 的最终 FACS 被单一 clone 主导,早期 FACS1 中仍保留大量多样性。作者用 FACS1 over MACS 的 k-mer enrichment 训练 regularized logistic regression model,在 1,909 个 FACS1 ROBO2N ARMs 中选择 29 个高分、但在后续 sorting 中流失的候选。结果,top 10 ML-derived antibodies 中有 9 个显示强 binding kinetics;cell display 证实 11 个 ROBO2N binders;这 11 个也都结合 ROBO1,且未观察到对 PD-L1、PD-L2、DCC、LOX1 的 off-target cross-reactivity。

第五组强数据是 PD-L2 的复现。PD-L2 experimental campaign 原本只得到 2 个好 binder,false positive rate 为 86%。同样的 LR strategy 从早期 FACS 数据中选出 33 个 ARM motifs,其中 17 个抗体在 cell display 中显示 comparable potency,把 false positive rate 降到 48%,虽然 SPR properties 相对弱一些。这说明 ML rescue 不是只在 ROBO2N 这个案例里有信号。

最重要的信息是:好抗体可能藏在早期筛选数据里

这篇文章最值得记住的一点,不是 logistic regression 多先进。相反,作者用的是相对简单、可解释的模型。真正重要的是它揭示了 display selection 的一个结构性问题:最终 enrichment 不是抗体质量的完整代理。

在 yeast display 中,clone 的最终频率会受到很多非目标因素影响,包括 yeast display efficiency、heavy/light-chain pairing、growth during amplification、sorting stringency 和 antigen format。一个低频 clone 可能因为表达、扩增或 gate 选择而消失,但它本身仍然可以是好 binder。ROBO2N 和 PD-L2 的结果说明,早期 sorting round 的 NGS 数据中保留了这类被实验流程漏掉的候选。

ARM 的价值正是在这里。它让每个候选的关键 paratope 可以被短码追踪,让 selection trajectory 变成模型可以学习的序列-富集关系。机器学习在这里不是替代实验,而是从实验没有充分利用的数据中重新排序候选。

这也改变了抗体发现平台的评价方式。一个平台不只应该看最终挑出的 top clone 有多好,还应该看它是否留下足够丰富、足够标准化、足够可复用的数据,让后续模型能够不断提高发现效率。

不能把它读成通用 AI 抗体设计已经解决

这项研究很强,但必须避免过度解读。

第一,它不是 zero-shot antibody design。模型没有从任意抗原序列直接设计抗体,而是在真实 yeast display selection 和 NGS 数据基础上做 candidate rescue。这里的 ML 更接近 selection-data mining,而不是完全体外的 de novo antibody design。

第二,库的多样性是有意约束的。VH1-69 scaffold、固定 CDR1/CDR2、有限 light-chain set 和 CDRH3-centered design 让数据更干净、更适合机器学习,但也意味着某些需要其他 CDR 区域参与的 paratopes 无法被访问。作者在 Discussion 中也承认,ROBO epitope hotspots 可能既反映 antigen feature,也反映 library diversity 的限制。

第三,抗原形式仍然是工程化 ectodomain 或 purified construct。cell-surface antigen 的天然膜环境、glycosylation、构象状态、cell-type context 和 tissue accessibility 都可能影响真实抗体表现。文章有 cell display、flow cytometry 和 IHC 验证,但离 therapeutic-grade validation 仍有距离。

第四,ML 部分还很早期。ROBO2N 和 PD-L2 两个案例证明了方向,但模型主要是 logistic regression 和 k-mer enrichment,尚未展示跨靶点泛化、跨 scaffold 泛化、结构感知设计或真正的 prospective large-scale validation。PD-L2 的 ML rescue 也降低了 false positive rate,但 SPR properties 仍较弱。

第五,研究虽声明 authors declare no competing interests,并公开了 Zenodo NGS 数据和代码,但要把这个平台变成社区级资源,还需要更清楚的 reagent access、更多靶点家族、更多失败案例和不同实验室的独立复现。

下一步是从可学习数据集走向可泛化模型

这项研究对抗体发现领域最大的启发,是应该把实验管线设计成数据管线。未来抗体库不只是为了筛出 binder,也应该为了产生高质量的正例、负例、低频候选、enrichment trajectory、biophysical characterization 和 downstream assay 数据。

第一步是扩大靶点范围。作者已经证明 10 个 cell surface antigens 可行,但真正训练可泛化模型需要更多 receptor families、orthologs、paralogs、不同 structural classes 和更困难的 membrane proteins。特别是 ROBO1/ROBO2 这种高同源 paralog 体系,很适合学习 cross-reactivity、species specificity 和 epitope hotspot。

第二步是把模型从 rescue 走向 prediction。现在 LR model 的价值是从早期 selection 中找回漏网 clone。下一阶段可以测试更强模型是否能预测 affinity、off-rate、polyreactivity、developability、epitope bin 和 cross-reactivity,而不仅仅是预测 enrichment。

第三步是引入结构和功能终点。ARM short code 很适合 NGS 和 k-mer 模型,但抗体识别最终仍是三维 interaction。把 ARM trajectory 与 AlphaFold-style structural modeling、epitope binning、SPR kinetics、cell binding、IHC/flow application 和 developability 数据连接起来,才可能让模型真正学到抗体功能。

第四步是社区化。作者公开了超过 68,000 个与靶点相关、至少有 2 个 ARM counts 的 unique ARM sequences,以及 486 个抗体的 characterization。这类数据如果持续扩展,可能比单个 antibody discovery campaign 更有长期价值。

Yang 的信号评级:High

轴一,信号强度:High。 理由是这篇文章把 antibody discovery 的实验平台和机器学习输入结构真正连了起来:ARM short code、high-throughput yeast display、NGS tracking、10 个 cell-surface antigens、424 个重组抗体,以及 ROBO2N/PD-L2 的 ML rescue,共同支持“早期筛选数据可以找回被实验选择漏掉的功能性抗体”这个核心结论。

轴二,成熟度:Medium。 理由是这已经是可运行的平台和公开数据集,不只是概念图;但它仍依赖一个受限 scaffold 和 CDRH3-centered library,ML 主要在两个靶点上验证,治疗转化还需要更广泛的靶点、独立实验室复现、结构/功能泛化和更严格的 developability/safety 评估。

一句话总结:这篇文章的核心价值,是把抗体发现从“筛到什么算什么”推进到“每一轮筛选都生成可被模型重新挖掘的数据”。

Deepash Kothiwal, corresponding author Rob Meijers and colleagues from the Institute for Protein Innovation and collaborating institutions recently reported in Cell Systems a synthetic Fab yeast-display library built around a VH1-69 scaffold and a compact CDRH3-based antigen recognition module (ARM). They screened the library in parallel against 10 cell-surface antigens and used logistic regression on early sorting NGS data to recover ROBO2N and PD-L2 binders missed by experimental selection, creating an ML-ready platform for hybrid in silico and experimental antibody discovery.

Content infographic

The real problem is making antibody screening learnable

Antibody discovery already has mature routes, including phage display, yeast display, immunization and single B-cell cloning. The bottleneck is not only whether a binder can be found. It is whether the discovery process can become a scalable, reusable data system that machine learning can actually read.

Traditional display libraries often vary multiple CDR regions at once. That diversity helps discovery, but it is hard for machine learning: the sequence space is large, heavy/light-chain pairing is complex, and selection is affected by expression, proliferation, display efficiency and sorting gates. A clone that is frequent in the final FACS round is not necessarily the best antibody, and a clone that appears early but disappears later is not necessarily useless.

This paper asks a specific question: can an antibody library be designed to be simple enough, large enough and sequenceable enough to connect experimental screening with machine learning?

The authors focus antigen recognition on CDRH3, one of the most important and variable antibody recognition regions. They design an ARM shorter than 100 nucleotides that covers heavy-chain CDRH3 plus an adjacent light-chain barcode, allowing the key recognition module of each Fab to be tracked directly by NGS. Screening therefore produces not only candidate antibodies, but an antibody-antigen interaction dataset that can be used for ML.

The novelty is ARM, a compact antigen-recognition code

The technical novelty has three layers.

The first is library design. The authors use a VH1-69 heavy-chain scaffold and pair it with four light chains. VK1-39, VK3-15 and VK3-20 provide relatively flat paratopes, while VK4-1 provides a more concave paratope. CDRH3 lengths mainly cover 11-17 amino acids, and position-specific amino acid frequencies are derived from 9.5 million unique naive B-cell CDRH3 sequences in the Observed Antibody Space. Cysteine, methionine and motifs associated with aggregation, polyreactivity or chemical liabilities are removed.

The second is the ARM representation. CDRH3 is followed by a light-chain barcode, creating a short nucleotide module. This module represents the main paratope and records heavy/light-chain pairing. Compared with libraries that require reading all six CDRs, this ARM structure is easier to deep sequence and easier to model with k-mers, distances or more complex algorithms.

The third is study design. The authors do not stop at a single proof-of-concept target. They screen the library in parallel against 10 cell-surface glycoprotein antigens: PD-L1, PD-L2, TIGIT, LOX1, DKK1, IL-23R, DCC, ROBO1, ROBO2 and syncytin-2. The workflow uses one round of MACS, three rounds of FACS and ARM NGS tracking after each round.

The novelty is therefore not simply that the paper uses machine learning. It redesigns antibody discovery itself to be ML-compatible from the beginning, aligning the antibody library, barcode, NGS readout, screening workflow and model input.

The strongest data are 10 targets, 424 antibodies and ML-rescued clones

The first strong dataset is library scale and sequence quality. Deep sequencing of two VH1-69 sub-libraries detected 390 million unique CDRH3 sequences, and each sub-library had an estimated molecular diversity of about 2.5 x 10^8 unique clones. With four sub-libraries combined, the authors estimate that the VH1-69 library approaches 1 billion unique Fabs. The ARM library’s 2-mer and 3-mer usage correlates with the naive B-cell CDRH3 repertoire, with Spearman correlations of 0.78 and 0.76, while reducing unfavorable motifs.

The second strong dataset is the parallel screen across 10 targets. Based on enriched CDRH3 clusters in the final FACS round, the authors ordered 429 heavy chains and expressed 424 IgG1 antibodies with sufficient yield. Of these, 301 passed SEC checks for aggregation or degradation, 354 showed no reactivity to avidin, DNA, insulin or Sf9 membrane preparation, and 285 passed both SEC and polyreactivity filters.

The third strong dataset is functional validation. The authors performed SPR on 335 experimentally verified antibodies and cell-display titration on 317 antibodies. They found 103 antibodies with K_D < 10 nM and 118 antibodies with EC50 < 25 nM. Selected TIGIT and LOX1 antibodies performed comparably to commercial antibodies in flow cytometry. Representative ROBO1, ROBO2 and syncytin-2 antibodies had melting temperatures above 70 C. ROBO1 and ROBO2 rabbit-chimera antibodies also produced clear IHC staining in mouse embryonic neural tissue.

The fourth strong dataset is ML rescue. The final ROBO2N FACS round was dominated by a single clone, while early FACS1 still contained substantial diversity. The authors trained a regularized logistic regression model using k-mer enrichment from FACS1 over MACS and scored 1,909 ROBO2N ARMs with FACS1 counts above 5. They selected 29 high-scoring ARMs that were depleted in later sorting rounds. Nine of the top 10 ML-derived antibodies showed strong binding kinetics; cell display confirmed 11 ROBO2N binders; all 11 also bound ROBO1, and no off-target cross-reactivity was observed against PD-L1, PD-L2, DCC or LOX1.

The fifth strong dataset is PD-L2 replication. The experimental PD-L2 campaign originally produced only two good binders and had an 86% false-positive rate. Applying the same LR strategy to early FACS data selected 33 ARM motifs, of which 17 antibodies showed comparable potency in cell display, lowering the false-positive rate to 48%, although their SPR properties were weaker. This suggests that ML rescue is not limited to the ROBO2N case.

Good antibodies can be hidden in early sorting data

The most important point is not that logistic regression is a sophisticated model. It is deliberately simple and interpretable. The important message is that display selection has a structural limitation: final enrichment is not a complete proxy for antibody quality.

In yeast display, final clone frequency is shaped by many non-target factors, including display efficiency, heavy/light-chain pairing, growth during amplification, sorting stringency and antigen format. A low-frequency clone may disappear because of expression, amplification or gating, while still being a good binder. The ROBO2N and PD-L2 results show that early sorting NGS data can preserve candidates that the experimental workflow later loses.

That is where ARM matters. It lets each candidate’s key paratope be tracked as a short code and turns selection trajectories into sequence-enrichment relationships that models can learn. Machine learning here does not replace experiments. It reorders candidates from experimental data that would otherwise remain underused.

This also changes how antibody discovery platforms should be judged. A platform should not only be evaluated by the top clones it returns at the end. It should also be evaluated by whether it leaves behind rich, standardized and reusable data that models can mine later.

This is not yet general AI antibody design

The study is strong, but it should not be overread.

First, this is not zero-shot antibody design. The model does not design antibodies directly from arbitrary antigen sequences. It rescues candidates from real yeast-display selection and NGS data. The ML component is closer to selection-data mining than fully in silico de novo antibody design.

Second, the library diversity is deliberately constrained. The VH1-69 scaffold, fixed CDR1/CDR2 regions, limited light-chain set and CDRH3-centered design make the dataset cleaner and more ML-friendly, but they also exclude paratopes that depend on other CDR regions. The authors acknowledge that the ROBO epitope hotspots may reflect both antigen features and library-diversity limits.

Third, the antigens are engineered ectodomains or purified constructs. Native membrane context, glycosylation, conformational state, cell-type context and tissue accessibility can all affect antibody performance. The paper includes cell display, flow cytometry and IHC validation, but it is still far from therapeutic-grade validation.

Fourth, the ML evidence is early. ROBO2N and PD-L2 demonstrate the direction, but the model is mainly logistic regression on k-mer enrichment. The paper does not yet show cross-target generalization, cross-scaffold generalization, structure-aware design or large-scale prospective validation. PD-L2 rescue lowered the false-positive rate, but SPR properties remained weaker.

Fifth, although the authors declare no competing interests and release the Zenodo NGS data and code, turning this into a community resource will require clearer reagent access, more target families, more negative cases and independent replication across laboratories.

The next step is from learnable datasets to generalizable models

The biggest lesson for antibody discovery is that experimental pipelines should be designed as data pipelines. Future libraries should not only find binders. They should generate high-quality positives, negatives, low-frequency candidates, enrichment trajectories, biophysical characterization and downstream assay data.

The first next step is target expansion. The authors show feasibility across 10 cell-surface antigens, but generalizable models will need more receptor families, orthologs, paralogs, structural classes and difficult membrane proteins. High-homology paralog systems such as ROBO1/ROBO2 are especially useful for learning cross-reactivity, species specificity and epitope hotspots.

The second step is moving from rescue to prediction. Here, the LR model rescues overlooked clones from early selection. The next question is whether stronger models can predict affinity, off-rate, polyreactivity, developability, epitope binning and cross-reactivity, not merely enrichment.

The third step is adding structural and functional endpoints. The ARM short code is well suited for NGS and k-mer models, but antibody recognition is ultimately a three-dimensional interaction. Linking ARM trajectories to structural modeling, epitope binning, SPR kinetics, cell binding, IHC or flow cytometry performance and developability data is what could let models learn antibody function.

The fourth step is community scale. The authors release more than 68,000 target-associated unique ARM sequences with at least two ARM counts, as well as characterization for 486 antibodies. If expanded over time, this dataset may be more valuable than any single antibody discovery campaign.

Yang’s signal rating: High

Axis 1, signal strength: High. The paper meaningfully connects an antibody discovery platform to a machine-learning input structure. The ARM short code, high-throughput yeast display, NGS tracking, 10 cell-surface antigens, 424 recombinant antibodies and ROBO2N/PD-L2 ML rescue all support the central claim that early selection data can recover functional antibodies missed by experimental selection.

Axis 2, maturity: Medium. This is already a working platform and public dataset, not just a concept. But it still depends on a constrained scaffold and CDRH3-centered library, the ML evidence is mainly validated on two targets, and therapeutic translation will require broader targets, independent replication, structural and functional generalization, and stronger developability and safety assessment.

One-sentence summary: The core value of this paper is that it moves antibody discovery from “take whatever survives screening” toward “make every selection round produce data that models can mine again.”