Anthropic and NVIDIA target the lab bench
-
Nº XLVII
- Date
- 01 Jul 2026
- Issue
- 47
- Stories
- Six
- Editor
- ARC
Two vendor plays for the scientist's desktop, two new benchmarks to grade them by — the agentic-biology stack is picking sides.
Claude Science lands on the bench
Anthropic launched Claude Science, a research workbench that wraps Claude in the tools scientists actually use — bioinformatics packages, notebooks, and flexible compute — and produces auditable artifacts for every step from literature review through data analysis. The workspace is customizable per lab, meaning teams can pin their own environments rather than negotiate with a generic coding agent. It's Anthropic's clearest bid yet to sit on the researcher's desktop, not just in the API.
NVIDIA draws the AI-scientist line
NVIDIA released BioNeMo with a pointed argument: most "AI for science" is a general coding agent aimed at biology and told to find a drug, which is not the same as an AI scientist. The framing stakes out domain-specialized agents as the category to beat, and pressures every generalist workbench — Claude Science included — to defend why a general model plus tools is enough.
Self-evolving agents run wet-lab protocols
A self-evolving agentic system generates and executes biological protocols end-to-end, iterating on its own procedures based on execution feedback rather than shipping a fixed pipeline. Moves protocol-writing agents from static templates toward closed-loop experimentation, where the agent that plans the experiment is the same one that revises it after the run.
HealthAgentBench pressures frontier agents
HealthAgentBench arrived as a unified suite of realistic agentic healthcare environments designed to break frontier models. Establishes a shared reference for what "clinically deployable" means at the agent level — the same problem LifeSciBench targeted for life science research tasks — where until now every vendor graded their own homework.
OpenAI ships GeneBench-Pro
OpenAI introduced GeneBench-Pro, a benchmark for multistage statistical reasoning across genomics, quantitative biology, and translational biomedicine, built on complex real-world datasets. Anchors a new reference score for genomics AI — and lands the same week as HealthAgentBench, marking benchmarks as the fastest-growing category in bio-AI.
Retrieval plus curation scales AOP mapping
Semantic retrieval, LLM refinement, and expert curation combine in a new pipeline for adverse outcome pathway gene mapping. Shows the hybrid pattern — retrieval augmented generation (letting the model pull from a curated knowledge base) paired with human review — winning where pure LLM extraction still misses on toxicology-grade precision.
Reply with your discoveries. A human reads them. Forward freely.
|