6 min read

Benchmarks come for PhD-level biology

Benchmarks come for PhD-level biology
Nº 01 · The Lede bioRxiv Field report

A harder exam for models

A harder exam for models
Fig. IbioRxiv · Filed 20 Aug 2026.

LifeSciBench scores language models on realistic, expert-level life-science tasks, released as a bioRxiv preprint. The design brief is to test multi-step work a practicing scientist actually does rather than exam-style recall. Realistic-task suites are harder to assemble and harder to saturate, and that difficulty is the feature: it keeps a benchmark discriminating long after multiple-choice sets go flat. The practical effect is that "this model handles biology" acquires a number someone else can reproduce and contest. Suites like this become the reference floor for accuracy claims across the field, which is how a marketing adjective turns into a falsifiable one.

Read the source

OpenAI holds the retention line
Fig. IIOpenAI · Filed 20 Aug 2026.
Nº 02 OpenAI Field report

OpenAI holds the retention line

OpenAI reaffirmed Zero Data Retention for eligible API customers on frontier models, and previewed Private Safety Processing, its approach to running safety checks on longer autonomous work without exposing the underlying content. Zero Data Retention means API inputs and outputs aren't kept on OpenAI's servers after a call completes; it's the contractual hook most regulated deployments hang on. As agents take on longer jobs with wider tool access, safety monitoring and data confidentiality pull against each other. Holding both is what keeps patient-adjacent and IP-sensitive biology workloads viable on frontier models at all.

Read more

Also discussed on X.

Routing beats one big agent
Fig. IIIarXiv · Filed 20 Aug 2026.
Nº 03 arXiv Agents · Infrastructure

Routing beats one big agent

A meta-agent picks the agent in Eureka, an arXiv proposal for task-conditioned orchestration of scientific discovery. Rather than one generalist attempting everything, a controller reads the task and routes it to specialists and tools assembled for that job. Orchestration is where most multi-agent science systems break, because the routing decision sets the ceiling, not the individual agents. If task-conditioned routing holds up, the design question for discovery platforms moves from "which model" to "which composition" — and purpose-built bio agents become worth building instead of folding into a generalist.

Read more
Also Filed · Two Briefs from the queue
Nº 04 bioRxiv Field report

Proteomics beats parameters

Adding proteomics beats scaling parameters for single-cell foundation models, according to a new bioRxiv preprint. Cross-modal training outperforming raw model size resets the cost calculus for virtual-cell work, where the default path to better performance has been buying bigger.

Read
Nº 05 Axios Field report

FDA gets a nominee

Trump will nominate Heidi Overton to lead the FDA, a senior administration official told Axios, filling a seat vacant since Marty Makary's exit unsettled the agency. Leadership stability at the top of the review pipeline sets the pace for every novel-modality and AI-enabled submission queued behind it.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 81  ·  20 Aug 2026

Editor's Note

A benchmark that bites, a data-retention promise, and a router that picks your agent for you.

 

Nº 01 · The Lede  —  bioRxiv  —  Field report

A harder exam for models

A harder exam for models

Fig. I  bioRxiv · Filed 20 Aug 2026.

LifeSciBench scores language models on realistic, expert-level life-science tasks, released as a bioRxiv preprint. The design brief is to test multi-step work a practicing scientist actually does rather than exam-style recall. Realistic-task suites are harder to assemble and harder to saturate, and that difficulty is the feature: it keeps a benchmark discriminating long after multiple-choice sets go flat. The practical effect is that "this model handles biology" acquires a number someone else can reproduce and contest. Suites like this become the reference floor for accuracy claims across the field, which is how a marketing adjective turns into a falsifiable one.

Read the source →

Why it matters

Expert-task evaluation is where claims about AI in biology get settled; "trust us, it's PhD-level" stops working the moment a competitor publishes a LifeSciBench number.

 

Nº 02  —  OpenAI  —  Field report

OpenAI holds the retention line

Fig. II  OpenAI · Filed 20 Aug 2026.

OpenAI holds the retention line

OpenAI reaffirmed Zero Data Retention for eligible API customers on frontier models, and previewed Private Safety Processing, its approach to running safety checks on longer autonomous work without exposing the underlying content. Zero Data Retention means API inputs and outputs aren't kept on OpenAI's servers after a call completes; it's the contractual hook most regulated deployments hang on. As agents take on longer jobs with wider tool access, safety monitoring and data confidentiality pull against each other. Holding both is what keeps patient-adjacent and IP-sensitive biology workloads viable on frontier models at all.

Read more →

The Bench NoteFrom Heureka Labs

The terms under which sensitive biology runs on these systems are becoming written commitments rather than assumptions.

Two places to check. A provider's terms govern a call once it leaves; Bench adds a control on your side, where Privacy Mode blocks every off-machine tool at the tool level: web, literature and reference lookups, compound resolution, cloud jobs and sync.
The receipt. While it is on, the chat shows each blocked attempt as it happens, and a plain summary lands in a log inside the project you can open.
What it buys. A guarantee about where files go, not a discount — the agent still thinks, so a locked project draws credits at the usual rate.

What we’re watching: whether researchers start expecting a per-run record of what left the machine from every tool in their stack

 

Nº 03  —  arXiv  —  Agents · Infrastructure

Routing beats one big agent

Fig. III  arXiv · Filed 20 Aug 2026.

Routing beats one big agent

A meta-agent picks the agent in Eureka, an arXiv proposal for task-conditioned orchestration of scientific discovery. Rather than one generalist attempting everything, a controller reads the task and routes it to specialists and tools assembled for that job. Orchestration is where most multi-agent science systems break, because the routing decision sets the ceiling, not the individual agents. If task-conditioned routing holds up, the design question for discovery platforms moves from "which model" to "which composition" — and purpose-built bio agents become worth building instead of folding into a generalist.

Read more →

 

Also Filed  ·  Two Briefs from the queue

Nº 04  —  bioRxiv  —  Field report

Proteomics beats parameters

Adding proteomics beats scaling parameters for single-cell foundation models, according to a new bioRxiv preprint. Cross-modal training outperforming raw model size resets the cost calculus for virtual-cell work, where the default path to better performance has been buying bigger.

Read →

Nº 05  —  Axios  —  Field report

FDA gets a nominee

Trump will nominate Heidi Overton to lead the FDA, a senior administration official told Axios, filling a seat vacant since Marty Makary's exit unsettled the agency. Leadership stability at the top of the review pipeline sets the pace for every novel-modality and AI-enabled submission queued behind it.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.