5 min read

Anthropic and NVIDIA target the lab bench

Anthropic and NVIDIA target the lab bench
Nº 01 · The Lede X Field report

Claude Science lands on the bench

Claude Science lands on the bench
Fig. IX · Filed 01 Jul 2026.

Anthropic launched Claude Science, a research workbench that wraps Claude in the tools scientists actually use — bioinformatics packages, notebooks, and flexible compute — and produces auditable artifacts for every step from literature review through data analysis. The workspace is customizable per lab, meaning teams can pin their own environments rather than negotiate with a generic coding agent. It's Anthropic's clearest bid yet to sit on the researcher's desktop, not just in the API.

Read the source

NVIDIA draws the AI-scientist line
Fig. IIX · Filed 01 Jul 2026.
Nº 02 X Field report

NVIDIA draws the AI-scientist line

NVIDIA released BioNeMo with a pointed argument: most "AI for science" is a general coding agent aimed at biology and told to find a drug, which is not the same as an AI scientist. The framing stakes out domain-specialized agents as the category to beat, and pressures every generalist workbench — Claude Science included — to defend why a general model plus tools is enough.

Read more
Self-evolving agents run wet-lab protocols
Fig. IIIarXiv · Filed 01 Jul 2026.
Nº 03 arXiv Agents · Infrastructure

Self-evolving agents run wet-lab protocols

A self-evolving agentic system generates and executes biological protocols end-to-end, iterating on its own procedures based on execution feedback rather than shipping a fixed pipeline. Moves protocol-writing agents from static templates toward closed-loop experimentation, where the agent that plans the experiment is the same one that revises it after the run.

Read more
Also Filed · Three Briefs from the queue
Nº 04 arXiv Agents · Infrastructure

HealthAgentBench pressures frontier agents

HealthAgentBench arrived as a unified suite of realistic agentic healthcare environments designed to break frontier models. Establishes a shared reference for what "clinically deployable" means at the agent level — the same problem LifeSciBench targeted for life science research tasks — where until now every vendor graded their own homework.

Read
Nº 05 bioRxiv Field report

OpenAI ships GeneBench-Pro

OpenAI introduced GeneBench-Pro, a benchmark for multistage statistical reasoning across genomics, quantitative biology, and translational biomedicine, built on complex real-world datasets. Anchors a new reference score for genomics AI — and lands the same week as HealthAgentBench, marking benchmarks as the fastest-growing category in bio-AI.

Read
Nº 06 bioRxiv Field report

Retrieval plus curation scales AOP mapping

Semantic retrieval, LLM refinement, and expert curation combine in a new pipeline for adverse outcome pathway gene mapping. Shows the hybrid pattern — retrieval augmented generation (letting the model pull from a curated knowledge base) paired with human review — winning where pure LLM extraction still misses on toxicology-grade precision.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 47  ·  01 Jul 2026

Editor's Note

Two vendor plays for the scientist's desktop, two new benchmarks to grade them by — the agentic-biology stack is picking sides.

 

Nº 01 · The Lede  —  X  —  Field report

Claude Science lands on the bench

Claude Science lands on the bench

Fig. I  X · Filed 01 Jul 2026.

Anthropic launched Claude Science, a research workbench that wraps Claude in the tools scientists actually use — bioinformatics packages, notebooks, and flexible compute — and produces auditable artifacts for every step from literature review through data analysis. The workspace is customizable per lab, meaning teams can pin their own environments rather than negotiate with a generic coding agent. It's Anthropic's clearest bid yet to sit on the researcher's desktop, not just in the API.

Read the source →

Why it matters

Frontier vendors are done treating science as a downstream use case — the bench itself is now a product surface, and auditability is shipping as a default feature rather than an afterthought bolted on for regulators.

 

Nº 02  —  X  —  Field report

NVIDIA draws the AI-scientist line

Fig. II  X · Filed 01 Jul 2026.

NVIDIA draws the AI-scientist line

NVIDIA released BioNeMo with a pointed argument: most "AI for science" is a general coding agent aimed at biology and told to find a drug, which is not the same as an AI scientist. The framing stakes out domain-specialized agents as the category to beat, and pressures every generalist workbench — Claude Science included — to defend why a general model plus tools is enough.

Read more →

 

Nº 03  —  arXiv  —  Agents · Infrastructure

Self-evolving agents run wet-lab protocols

Fig. III  arXiv · Filed 01 Jul 2026.

Self-evolving agents run wet-lab protocols

A self-evolving agentic system generates and executes biological protocols end-to-end, iterating on its own procedures based on execution feedback rather than shipping a fixed pipeline. Moves protocol-writing agents from static templates toward closed-loop experimentation, where the agent that plans the experiment is the same one that revises it after the run.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  arXiv  —  Agents · Infrastructure

HealthAgentBench pressures frontier agents

HealthAgentBench arrived as a unified suite of realistic agentic healthcare environments designed to break frontier models. Establishes a shared reference for what "clinically deployable" means at the agent level — the same problem LifeSciBench targeted for life science research tasks — where until now every vendor graded their own homework.

Read →

Nº 05  —  bioRxiv  —  Field report

OpenAI ships GeneBench-Pro

OpenAI introduced GeneBench-Pro, a benchmark for multistage statistical reasoning across genomics, quantitative biology, and translational biomedicine, built on complex real-world datasets. Anchors a new reference score for genomics AI — and lands the same week as HealthAgentBench, marking benchmarks as the fastest-growing category in bio-AI.

Read →

Nº 06  —  bioRxiv  —  Field report

Retrieval plus curation scales AOP mapping

Semantic retrieval, LLM refinement, and expert curation combine in a new pipeline for adverse outcome pathway gene mapping. Shows the hybrid pattern — retrieval augmented generation (letting the model pull from a curated knowledge base) paired with human review — winning where pure LLM extraction still misses on toxicology-grade precision.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.