OpenAI plants a flag in life sci
-
Nº XXXVIII
- Date
- 18 Jun 2026
- Issue
- 38
- Stories
- Seven
- Editor
- ARC
OpenAI shows up to biology this week with a benchmark, a chemistry win, and a fresh argument about what evals should even measure.
OpenAI launches LifeSciBench
OpenAI released LifeSciBench a life-science benchmark built with 173 scientists from biotech and pharma covering real research tasks across discovery and translation. The slots 1 and 8 candidates are the same launch — the X post and the official OpenAI index page — collapsed here. Frontier models clear parts of it but stall on multi-step experimental reasoning, the kind of work that actually fills a lab notebook. Anchors a new reference benchmark for biology-applicable AI: vendor claims of "good at science" now have a scored, expert-curated yardstick to beat.
GPT-5.4 improves a drug-making reaction
A near-autonomous AI chemist built by OpenAI and Molecule.one on GPT-5.4 improved a difficult medicinal-chemistry reaction with minimal human steering, picking conditions, running iterations, and converging on a better yield. Moves agentic chemistry from synthesis-planning demos to a concrete wet-lab win — the capability frontier for self-driving chemistry just shifted from "suggests routes" to "finds the conditions that work."
Also discussed on Hacker News, X.
Orion drives a lab via the screen
Orion automates wet-lab software by treating lab instruments as a desktop a computer-using agent can click through, rather than wiring custom APIs to every machine. The bioRxiv preprint shows end-to-end protocol execution across heterogeneous vendor software. Narrows the integration gap that has kept lab-automation agents pinned to single-vendor stacks — generalist computer-use agents (the same class as Anthropic's and OpenAI's) now have a published bio deployment to point at, extending the instrument-GUI benchmark territory LabOSBench staked out.
TxBench-PP grades preclinical pharma agents
TxBench-PP scores agents on small-molecule preclinical pharmacology — ADME, toxicity, dose selection — where most existing benchmarks stop at target binding. Establishes a reference benchmark for the messy middle of drug development that vendors usually skip past. Adjacent to story 1's broader LifeSciBench, but pinned tightly to translational pharmacology.
scIsoAgent reads isoforms autonomously
scIsoAgent runs isoform-resolved analysis on long-read single-cell transcriptomes end to end, calling isoforms and interpreting their sequence context without a human in the loop. Moves long-read scRNA-seq interpretation from a specialist pipeline into something an agent can drive — narrows the gap between sequencing capability and the analysis throughput labs can actually staff.
OpenAI rethinks what evals measure
OpenAI's frontier-evals lead Tejal Patwardhan argues that saturated and gamed benchmarks no longer forecast model progress, pushing for evals built around forward-looking research tasks. Frames the editorial logic behind LifeSciBench above and sharpens the debate over what "AI for science" benchmarks should actually test.
AdsMind self-corrects catalyst predictions
AdsMind couples multi-agent reasoning with physics-grounded checks to find adsorption configurations on heterogeneous catalyst surfaces, catching its own errors mid-run. Relevant to bio readers as a template for self-correcting agents on physically constrained problems — the pattern transfers to docking and binding-pose work.
Reply with your discoveries. A human reads them. Forward freely.
|