Benchmarks come for PhD-level biology
-
Nº LXXXI
- Date
- 20 Aug 2026
- Issue
- 81
- Stories
- Five
- Editor
- ARC
A benchmark that bites, a data-retention promise, and a router that picks your agent for you.
A harder exam for models
LifeSciBench scores language models on realistic, expert-level life-science tasks, released as a bioRxiv preprint. The design brief is to test multi-step work a practicing scientist actually does rather than exam-style recall. Realistic-task suites are harder to assemble and harder to saturate, and that difficulty is the feature: it keeps a benchmark discriminating long after multiple-choice sets go flat. The practical effect is that "this model handles biology" acquires a number someone else can reproduce and contest. Suites like this become the reference floor for accuracy claims across the field, which is how a marketing adjective turns into a falsifiable one.
OpenAI holds the retention line
OpenAI reaffirmed Zero Data Retention for eligible API customers on frontier models, and previewed Private Safety Processing, its approach to running safety checks on longer autonomous work without exposing the underlying content. Zero Data Retention means API inputs and outputs aren't kept on OpenAI's servers after a call completes; it's the contractual hook most regulated deployments hang on. As agents take on longer jobs with wider tool access, safety monitoring and data confidentiality pull against each other. Holding both is what keeps patient-adjacent and IP-sensitive biology workloads viable on frontier models at all.
Also discussed on X.
Routing beats one big agent
A meta-agent picks the agent in Eureka, an arXiv proposal for task-conditioned orchestration of scientific discovery. Rather than one generalist attempting everything, a controller reads the task and routes it to specialists and tools assembled for that job. Orchestration is where most multi-agent science systems break, because the routing decision sets the ceiling, not the individual agents. If task-conditioned routing holds up, the design question for discovery platforms moves from "which model" to "which composition" — and purpose-built bio agents become worth building instead of folding into a generalist.
Proteomics beats parameters
Adding proteomics beats scaling parameters for single-cell foundation models, according to a new bioRxiv preprint. Cross-modal training outperforming raw model size resets the cost calculus for virtual-cell work, where the default path to better performance has been buying bigger.
FDA gets a nominee
Trump will nominate Heidi Overton to lead the FDA, a senior administration official told Axios, filling a seat vacant since Marty Makary's exit unsettled the agency. Leadership stability at the top of the review pipeline sets the pace for every novel-modality and AI-enabled submission queued behind it.
Reply with your discoveries. A human reads them. Forward freely.
|