6 min read

OpenAI plants a flag in life sci

OpenAI plants a flag in life sci
Nº 01 · The Lede X Field report

OpenAI launches LifeSciBench

OpenAI launches LifeSciBench
Fig. IX · Filed 18 Jun 2026.

OpenAI released LifeSciBench a life-science benchmark built with 173 scientists from biotech and pharma covering real research tasks across discovery and translation. The slots 1 and 8 candidates are the same launch — the X post and the official OpenAI index page — collapsed here. Frontier models clear parts of it but stall on multi-step experimental reasoning, the kind of work that actually fills a lab notebook. Anchors a new reference benchmark for biology-applicable AI: vendor claims of "good at science" now have a scored, expert-curated yardstick to beat.

Read the source

GPT-5.4 improves a drug-making reaction
Fig. IIOpenAI · Filed 18 Jun 2026.
Nº 02 OpenAI Drug discovery · Computational

GPT-5.4 improves a drug-making reaction

A near-autonomous AI chemist built by OpenAI and Molecule.one on GPT-5.4 improved a difficult medicinal-chemistry reaction with minimal human steering, picking conditions, running iterations, and converging on a better yield. Moves agentic chemistry from synthesis-planning demos to a concrete wet-lab win — the capability frontier for self-driving chemistry just shifted from "suggests routes" to "finds the conditions that work."

Read more

Also discussed on Hacker News, X.

Orion drives a lab via the screen
Fig. IIIbioRxiv · Filed 18 Jun 2026.
Nº 03 bioRxiv Field report

Orion drives a lab via the screen

Orion automates wet-lab software by treating lab instruments as a desktop a computer-using agent can click through, rather than wiring custom APIs to every machine. The bioRxiv preprint shows end-to-end protocol execution across heterogeneous vendor software. Narrows the integration gap that has kept lab-automation agents pinned to single-vendor stacks — generalist computer-use agents (the same class as Anthropic's and OpenAI's) now have a published bio deployment to point at, extending the instrument-GUI benchmark territory LabOSBench staked out.

Read more
Also Filed · Four Briefs from the queue
Nº 04 arXiv Clinical AI · Evaluation

TxBench-PP grades preclinical pharma agents

TxBench-PP scores agents on small-molecule preclinical pharmacology — ADME, toxicity, dose selection — where most existing benchmarks stop at target binding. Establishes a reference benchmark for the messy middle of drug development that vendors usually skip past. Adjacent to story 1's broader LifeSciBench, but pinned tightly to translational pharmacology.

Read
Nº 05 bioRxiv Agents · Infrastructure

scIsoAgent reads isoforms autonomously

scIsoAgent runs isoform-resolved analysis on long-read single-cell transcriptomes end to end, calling isoforms and interpreting their sequence context without a human in the loop. Moves long-read scRNA-seq interpretation from a specialist pipeline into something an agent can drive — narrows the gap between sequencing capability and the analysis throughput labs can actually staff.

Read
Nº 06 X Field report

OpenAI rethinks what evals measure

OpenAI's frontier-evals lead Tejal Patwardhan argues that saturated and gamed benchmarks no longer forecast model progress, pushing for evals built around forward-looking research tasks. Frames the editorial logic behind LifeSciBench above and sharpens the debate over what "AI for science" benchmarks should actually test.

Read
Nº 07 arXiv Field report

AdsMind self-corrects catalyst predictions

AdsMind couples multi-agent reasoning with physics-grounded checks to find adsorption configurations on heterogeneous catalyst surfaces, catching its own errors mid-run. Relevant to bio readers as a template for self-correcting agents on physically constrained problems — the pattern transfers to docking and binding-pose work.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 38  ·  18 Jun 2026

Editor's Note

OpenAI shows up to biology this week with a benchmark, a chemistry win, and a fresh argument about what evals should even measure.

 

Nº 01 · The Lede  —  X  —  Field report

OpenAI launches LifeSciBench

OpenAI launches LifeSciBench

Fig. I  X · Filed 18 Jun 2026.

OpenAI released LifeSciBench a life-science benchmark built with 173 scientists from biotech and pharma covering real research tasks across discovery and translation. The slots 1 and 8 candidates are the same launch — the X post and the official OpenAI index page — collapsed here. Frontier models clear parts of it but stall on multi-step experimental reasoning, the kind of work that actually fills a lab notebook. Anchors a new reference benchmark for biology-applicable AI: vendor claims of "good at science" now have a scored, expert-curated yardstick to beat.

Read the source →

Why it matters

Resets the floor for what counts as a credible biomedical AI claim — "trust us, it works on biology" stops being a pitch the moment competitors post LifeSciBench numbers.

 

Nº 02  —  OpenAI  —  Drug discovery · Computational

GPT-5.4 improves a drug-making reaction

Fig. II  OpenAI · Filed 18 Jun 2026.

GPT-5.4 improves a drug-making reaction

A near-autonomous AI chemist built by OpenAI and Molecule.one on GPT-5.4 improved a difficult medicinal-chemistry reaction with minimal human steering, picking conditions, running iterations, and converging on a better yield. Moves agentic chemistry from synthesis-planning demos to a concrete wet-lab win — the capability frontier for self-driving chemistry just shifted from "suggests routes" to "finds the conditions that work."

Read more →

 

Nº 03  —  bioRxiv  —  Field report

Orion drives a lab via the screen

Fig. III  bioRxiv · Filed 18 Jun 2026.

Orion drives a lab via the screen

Orion automates wet-lab software by treating lab instruments as a desktop a computer-using agent can click through, rather than wiring custom APIs to every machine. The bioRxiv preprint shows end-to-end protocol execution across heterogeneous vendor software. Narrows the integration gap that has kept lab-automation agents pinned to single-vendor stacks — generalist computer-use agents (the same class as Anthropic's and OpenAI's) now have a published bio deployment to point at, extending the instrument-GUI benchmark territory LabOSBench staked out.

Read more →

 

Also Filed  ·  Four Briefs from the queue

Nº 04  —  arXiv  —  Clinical AI · Evaluation

TxBench-PP grades preclinical pharma agents

TxBench-PP scores agents on small-molecule preclinical pharmacology — ADME, toxicity, dose selection — where most existing benchmarks stop at target binding. Establishes a reference benchmark for the messy middle of drug development that vendors usually skip past. Adjacent to story 1's broader LifeSciBench, but pinned tightly to translational pharmacology.

Read →

Nº 05  —  bioRxiv  —  Agents · Infrastructure

scIsoAgent reads isoforms autonomously

scIsoAgent runs isoform-resolved analysis on long-read single-cell transcriptomes end to end, calling isoforms and interpreting their sequence context without a human in the loop. Moves long-read scRNA-seq interpretation from a specialist pipeline into something an agent can drive — narrows the gap between sequencing capability and the analysis throughput labs can actually staff.

Read →

Nº 06  —  X  —  Field report

OpenAI rethinks what evals measure

OpenAI's frontier-evals lead Tejal Patwardhan argues that saturated and gamed benchmarks no longer forecast model progress, pushing for evals built around forward-looking research tasks. Frames the editorial logic behind LifeSciBench above and sharpens the debate over what "AI for science" benchmarks should actually test.

Read →

Nº 07  —  arXiv  —  Field report

AdsMind self-corrects catalyst predictions

AdsMind couples multi-agent reasoning with physics-grounded checks to find adsorption configurations on heterogeneous catalyst surfaces, catching its own errors mid-run. Relevant to bio readers as a template for self-correcting agents on physically constrained problems — the pattern transfers to docking and binding-pose work.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.