Yang Liu

← Thinking

当 AI 科研工作台变成默认配置 When AI Science Workbenches Become the Default

Anthropic 在 2026 年 6 月 30 日发布 Claude Science,把数据库、分析工具、可视化、计算资源和多 agent 工作流打包成一个面向科学家的 AI workbench。这个事件真正重要的地方,不是又多了一个科研 AI 产品,而是它开始把“会调数据库、会跑分析、会做图、会把工具串起来”这件事平台化。

我的判断是:“会用 AI 做科研”会迅速从差异化能力变成基础能力。真正的护城河会转移到三个地方:提出值得验证的问题、判断证据是否足够、把结果接到真实科研或转化决策里。

工具能力开始被打包成默认科研工作台

Anthropic 官方对 Claude Science 的定位很清楚:它不是一个新模型,而是一个科学工作台。它把科学家平时分散使用的数据库、代码环境、图表生成、计算资源、可视化和审稿式检查放进一个统一环境里。

官方信息里有几个关键点:Claude Science 可在本地 macOS、Linux 或远程机器 / HPC login node 上运行;它预置超过 60 个科学 skills 和 connectors,覆盖 genomics、single-cell、proteomics、structural biology、cheminformatics 等方向;它能原生渲染 3D protein structures、genome browser tracks、chemical structures;它还能生成带代码、环境和 message history 的可审计 artifact。

这意味着过去很多“技术门槛”正在被产品化。以前一个研究者要完成一个相对完整的 AI-assisted workflow,至少要懂数据库、懂脚本、懂可视化、懂模型接口、懂文件格式、懂计算环境。现在这些东西开始变成工作台的默认组件。

对一线研究者来说,这不是小改进。这是在降低科研中间层的摩擦。

三巨头争夺的不是一次查询,而是科研入口

TechFundingNews 把 Anthropic、Google、OpenAI 放在同一条竞争线上看,认为“通用科研底座”正在变成三巨头的新战场。这个判断有一定道理,但需要拆开看:TechFundingNews 是产业报道,不是三家公司共同确认的战略声明;真正可依赖的事实仍然要回到各家公司自己的发布。

Google 在 I/O 2026 推出 Gemini for Science,官方说法是用 Hypothesis Generation、Computational Discovery、Science Skills 等工具加速科学方法本身,其中 Science Skills 集成了 30+ 生命科学数据库和工具,包括 UniProt、AlphaFold Database、AlphaGenome API、InterPro 等。

OpenAI 在 2026 年 4 月发布 GPT-Rosalind,把它定位为面向 biology、drug discovery、translational medicine 的 frontier reasoning model。后续更新又强调 LifeSciBench、medicinal chemistry、genomics、wet lab troubleshooting,以及 Codex 里的 Life Sciences Research / NGS Analysis plugins。

所以这不是单点功能竞争,而是入口竞争。谁能成为科学家每天打开的工作台,谁就有机会占据三个高价值位置:数据入口、工作流入口、判断辅助入口。

但这也说明一个反向结论:当平台厂商都在做“默认科研工作台”,单纯会把 AI 接到 PubMed、PDB、UniProt、scRNA-seq pipeline、CRISPR screen design 上,就很难长期构成个人护城河。

护城河从“会不会用工具”转到“能不能验真”

如果数据库检索、初步分析、可视化和报告生成都被平台打包,研究者之间的差异不会消失,而是换位置。

第一层差异,是能不能提出好问题。AI workbench 可以更快回答问题,但它不能自动保证问题值得回答。在科研里,真正稀缺的不是“我能不能查到资料”,而是“这个问题是否切在机制、模型、实验条件和转化瓶颈的交叉点上”。

第二层差异,是能不能判断证据强度。Claude Science、Gemini for Science、GPT-Rosalind 都在强调 citations、artifact provenance、reviewer agents、auditable workflow。这些设计很重要,但它们解决的是可追踪性,不自动等于科学真实性。一个结果有引用、有代码、有图,并不代表实验设计合理、对照充分、统计解释稳健、临床转化路径成立。

第三层差异,是能不能把结果接到真实决策。AI 可以帮你生成一组候选靶点、候选机制、候选实验设计,但下一步选哪个、为什么选、先用什么系统验证、失败后如何解释,仍然依赖领域判断。

这就是护城河转移的方向:从操作熟练度,转向问题定义、证据分层和验证路径设计。

对一线研究者:威胁不是被替代,而是被平均化

很多研究者会把这类平台理解成“AI 会不会替代科研人员”。我觉得这个问法不够准确。

更现实的威胁是:AI 会把一批原本看起来很专业的中间能力平均化。

会写脚本做常规分析,门槛会下降。会查多个数据库,门槛会下降。会把结果做成图,门槛会下降。会快速写出综述草稿,门槛也会下降。

这些能力仍然重要,但不再足以证明一个研究者有独特判断。未来更能区分人的,是看见 AI output 后能不能问:

这个结论依赖了哪个数据库偏差?

这个分析有没有把 batch effect 当成生物学差异?

这个 candidate target 的表达证据和功能证据是否在同一个细胞类型里闭合?

这个 CRISPR screen hit 是否只是 viability artifact?

这个蛋白结构预测对真实构象、复合物状态、细胞内环境的假设是否过强?

这个 wet-lab validation 应该先做最便宜的 falsification,还是直接上昂贵系统?

这类问题不会因为工具更强而消失。相反,工具越强,错误越容易被包装得像正确结果。

对内容和咨询:价值不在工具介绍,而在判断框架

如果“会用 AI 做科研”本身被平台化,那么围绕 AI for Science 的内容和咨询也要重新定位。内容端的价值不能停留在“介绍某个工具怎么用”。这类内容会越来越快过期,也会越来越容易被平台文档、教程和别人的演示替代。

更有价值的内容方向,是持续回答:

这类平台把科研链条里的哪一段商品化了?

它真正降低的是搜索成本、分析成本、可视化成本,还是验证成本?

哪些能力会被平台吃掉,哪些能力反而更稀缺?

一线研究者应该如何重新定义自己的专业性?

AI output 到 wet-lab truth 中间还缺哪些判断?

咨询端也是一样。未来真正有价值的不是“帮团队接一个 AI 工具”,而是帮研究团队重构科研工作流里的判断层:什么任务应该自动化,什么任务必须人审,什么结果可以进入内部决策,什么结果只能当 hypothesis,什么地方必须设验证闸门。

也就是说,面向 AI for Science 的内容和咨询,核心不应是“我比别人更会用 AI”。更稳的定位是:帮助研究者、团队和机构判断 AI 在科研链条里到底改了哪一段,以及下一步该怎么验真。

这类平台仍然没有越过科学的最后一公里

Claude Science 的官方案例里,有 single-cell RNA-seq analysis、CRISPR screen design、protein structure prediction、cheminformatics,也有 Manifold Bio、Allen Institute、UCSF Brain Tumor Center 等早期使用案例。Google 和 OpenAI 的官方材料也都强调 hypothesis generation、evidence synthesis、NGS analysis、medicinal chemistry、wet lab troubleshooting。

这些都是真实进展。但不能误读成“AI 已经能独立完成科学发现”。

它们更像把科学工作中的一大段中间层压缩了:文献、数据库、代码、图表、计算、初步推理、报告。最后是否成立,还要回到实验系统、真实数据、可重复结果、机制解释和临床 / 转化路径。

所以,通用 agentic 科研平台的落地,不会让领域专家变得不重要。它会让低质量的“专家感”更容易被拆穿,也会让真正有判断力的人放大得更快。

我的判断

“会用 AI 做科研”在 2024-2025 年还可能是一种明显优势;到 Claude Science、Gemini for Science、GPT-Rosalind 这类平台陆续出现后,它会越来越像科研工作者的默认 literacy。

护城河不会消失,但会迁移。

旧护城河是:我会用工具。

新护城河是:我知道该问什么,知道什么证据不能信,知道下一步怎么验证。

对一线研究者,这是威胁,因为很多操作层能力会被平台平均化。对内容创作者、科研顾问和团队服务者,这是机会,因为市场会更需要一种“AI for Science 判断层”:不是报道工具,而是解释工具改变了科研链条里的哪一段;不是教人点按钮,而是帮助人把 AI output 接到可信证据和真实决策里。

主要信源

On June 30, 2026, Anthropic launched Claude Science, an AI workbench for scientists that packages databases, analysis tools, visualization, compute access, and multi-agent workflows into a single research environment. The important point is not that another AI product for science has appeared. It is that the work of querying databases, running analyses, making figures, and chaining tools together is starting to become platform infrastructure.

My view is that “knowing how to do research with AI” will quickly move from a differentiating capability to baseline literacy. The real moat shifts elsewhere: asking questions worth validating, judging whether evidence is strong enough, and connecting outputs to real research or translational decisions.

Tool fluency is being packaged into the default workbench

Anthropic’s own positioning is clear: Claude Science is not a new model, but a scientific workbench. It brings the databases, coding environments, figure generation, compute resources, visualization, and reviewer-like checks that researchers normally use across separate tools into one environment.

Several details matter. Claude Science can run locally on macOS and Linux, or on a remote machine or HPC login node. It comes with more than 60 scientific skills and connectors for genomics, single-cell analysis, proteomics, structural biology, cheminformatics, and other domains. It can natively render 3D protein structures, genome browser tracks, and chemical structures. It also generates auditable artifacts that include code, environment details, and message history.

This means many previous technical barriers are being productized. Before this kind of workbench, a researcher who wanted to run a serious AI-assisted workflow needed to understand databases, scripts, visualization, model interfaces, file formats, and compute environments. Those capabilities are now becoming default components of the workspace.

For frontline researchers, this is not a minor feature improvement. It lowers friction across the middle layer of scientific work.

The big platforms are competing for the research entry point

TechFundingNews frames Anthropic, Google, and OpenAI as entering the same race, with general scientific infrastructure becoming a new battleground. That framing is directionally useful, but it should be read carefully: it is an industry interpretation, not a joint strategic statement from the companies. The stronger facts come from each company’s own releases.

At Google I/O 2026, Google introduced Gemini for Science, including Hypothesis Generation, Computational Discovery, and Science Skills. Google says Science Skills integrates insights from more than 30 major life-science databases and tools, including UniProt, AlphaFold Database, AlphaGenome API, and InterPro.

OpenAI introduced GPT-Rosalind in April 2026 as a frontier reasoning model for biology, drug discovery, and translational medicine. Its later update emphasized LifeSciBench, medicinal chemistry, genomics, wet lab troubleshooting, and Codex-based Life Sciences Research and NGS Analysis plugins.

This is not a competition over isolated features. It is a competition over the research entry point. The company that becomes the workspace scientists open every day gains leverage over data access, workflow execution, and judgment support.

But this also implies the opposite conclusion: if every major platform is building a default science workbench, simply connecting AI to PubMed, PDB, UniProt, scRNA-seq pipelines, or CRISPR screen design will not remain a durable personal moat.

The moat moves from tool use to truth testing

If database search, first-pass analysis, visualization, and report generation are packaged by platforms, the difference between researchers does not disappear. It moves.

The first layer of difference is whether one can ask good questions. An AI workbench can answer questions faster, but it cannot guarantee that the question is worth answering. In science, the scarce ability is not merely finding information. It is identifying whether a question sits at the right intersection of mechanism, model system, experimental condition, and translational bottleneck.

The second layer is evidence judgment. Claude Science, Gemini for Science, and GPT-Rosalind all emphasize citations, artifact provenance, reviewer agents, and auditable workflows. These are important, but they solve traceability, not truth by themselves. A result can have citations, code, and figures and still rest on weak experimental design, inadequate controls, unstable statistics, or an unrealistic clinical path.

The third layer is decision connection. AI can produce candidate targets, mechanisms, or experiment designs. But choosing what to do next, why to choose it, which system to validate in first, and how to interpret failure still depends on domain judgment.

This is the direction of the moat: away from operational fluency and toward question definition, evidence stratification, and validation-path design.

For frontline researchers, the threat is not replacement but averaging

Many researchers will interpret these platforms through the question, “Will AI replace scientists?” I think that is the wrong starting point.

The more realistic threat is that AI will average out a class of middle-layer skills that used to look highly professional.

Writing scripts for standard analyses will become easier. Searching multiple databases will become easier. Turning results into figures will become easier. Producing a first review draft will become easier.

These skills will still matter, but they will no longer be enough to prove distinctive judgment. What will matter more is whether a researcher can look at AI output and ask:

Which database bias does this conclusion depend on?

Did this analysis mistake a batch effect for a biological difference?

Are the expression evidence and functional evidence for this candidate target closed in the same cell type?

Is this CRISPR screen hit a viability artifact?

Does this protein structure prediction over-assume the relevant conformation, complex state, or intracellular environment?

Should the next wet-lab validation be the cheapest falsification test, or is it worth moving directly into a more expensive system?

These questions will not disappear as tools improve. If anything, stronger tools make wrong conclusions easier to package as polished outputs.

For content and consulting, the value is judgment, not tool coverage

If “knowing how to do research with AI” becomes platformized, content and consulting around AI for Science also need to be repositioned. The content layer should not stop at explaining how a tool works. That kind of content will expire quickly and will be easy for product docs, tutorials, and demos to replace.

The more durable content layer asks different questions:

Which part of the research chain has this platform commoditized?

Is it lowering search cost, analysis cost, visualization cost, or validation cost?

Which capabilities will the platform absorb, and which capabilities become more scarce?

How should frontline researchers redefine their own expertise?

What judgment still sits between AI output and wet-lab truth?

The consulting layer follows the same logic. The future value is not helping a team plug in an AI tool. It is helping research teams redesign the judgment layer of the research workflow: which tasks should be automated, which tasks require human review, which outputs can inform internal decisions, which outputs should remain hypotheses, and where validation gates must be placed.

In other words, AI for Science content and consulting should not be built around “I use AI better than others.” A more durable position is helping researchers, teams, and institutions understand which part of the scientific workflow AI has changed, and how to test whether the output is true enough to act on.

These platforms still do not cross science’s last mile

Anthropic’s official examples include single-cell RNA-seq analysis, CRISPR screen design, protein structure prediction, cheminformatics, and early users such as Manifold Bio, the Allen Institute, and the UCSF Brain Tumor Center. Google’s and OpenAI’s materials similarly emphasize hypothesis generation, evidence synthesis, NGS analysis, medicinal chemistry, and wet lab troubleshooting.

These are real advances. But they should not be misread as evidence that AI can independently complete scientific discovery.

What these tools compress is a large middle layer of science: literature, databases, code, figures, compute, first-pass reasoning, and reports. Whether the result is true still depends on experimental systems, real data, reproducible findings, mechanism, and translational path.

So the arrival of general agentic science platforms will not make domain experts irrelevant. It will make low-quality expert-like performance easier to expose, and it will amplify people with real judgment.

My read

“Knowing how to do research with AI” was a visible advantage in 2024 and 2025. After Claude Science, Gemini for Science, and GPT-Rosalind, it will increasingly become baseline literacy for researchers.

The moat does not disappear. It migrates.

The old moat was: I know how to use the tools.

The new moat is: I know what to ask, what evidence not to trust, and how to validate the next step.

For frontline researchers, this is a threat because many operational skills will be averaged out by platforms. For content creators, scientific consultants, and service providers around research teams, it is an opportunity because the market will need an AI for Science judgment layer: not tool coverage, but interpretation of which part of the scientific workflow the tool changes; not button training, but helping people connect AI output to credible evidence and real decisions.

Sources