7 min read

Who checks the agent's work?

Who checks the agent's work?
Nº 01 · The Lede OpenAI Agents · Infrastructure

OpenAI agents produce a machine-checked Navier-Stokes proof

OpenAI agents produce a machine-checked Navier-Stokes proof
Fig. IOpenAI · Filed 09 Sep 2026.

OpenAI published a claimed solution to the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems, produced by a group of agents running on an unreleased next-generation model. The writeup ships with a formal proof in Lean, a proof-assistant language that machine-checks every step, so the claim doesn't rest on anyone trusting the model's prose. Mathematicians will spend weeks on the informal argument regardless. For biology the sharp part is the verification asymmetry: math has a machine checker, and discovery claims in biology have only the bench, which resets how much scrutiny an agent-generated hypothesis has to survive before anyone spends reagents on it.

Read the source

DeepMind maps every possible single-letter DNA change
Fig. IIX · Filed 09 Sep 2026.
Nº 02 X Field report

DeepMind maps every possible single-letter DNA change

Google DeepMind launched AlphaGenome Atlas, a searchable database of predicted functional impact for all 9 billion possible single-letter changes in human DNA. Every position, every substitution, precomputed and queryable rather than run on demand. Coverage that complete moves variant-effect prediction from a modeling exercise to a lookup, echoing the same shift a proteome-wide binder atlas made for binder design, and it changes what counts as a reasonable first pass on a variant of uncertain significance and on the long tail of GWAS hits nobody had compute to interrogate.

Read more
bioq gives AI agents one drug-discovery command line
Fig. IIIbioRxiv · Filed 09 Sep 2026.
Nº 03 bioRxiv Drug discovery · Computational

bioq gives AI agents one drug-discovery command line

bioq unifies drug-discovery tools behind a single command-line interface designed for agents to call rather than people to type, exposing what the bioRxiv preprint describes as a fleet of AI methods through one consistent grammar. The bottleneck is real: most published discovery models ship as one-off repos with incompatible inputs and outputs, and any agent chaining them pays the glue-code tax. Standardizing that interface is what turns a pile of methods into something an agent can plan over, and it lowers the integration cost that has kept much of the published toolkit unused.

Read more
Also Filed · Three Briefs from the queue
Nº 04 bioRxiv Agents · Infrastructure

Agents take over model building for RNA structures

Agents did the model development for RNA 3D structure prediction in a new bioRxiv preprint, iterating on design choices rather than running an existing predictor. Pushes agents up the stack into building the models for biology's least-solved structure problem.

Read
Nº 05 arXiv Clinical AI · Evaluation

Study compares clinical AI, physicians and frontier models

A three-way diagnostic comparison puts a purpose-built clinical AI system against practicing physicians and general-purpose frontier language models on primary care cases. Head-to-head framing like this is what settles whether medical AI needs bespoke systems or whether general models have closed the gap.

Read
Nº 06 arXiv Clinical AI · Evaluation

SentryLine tracks evidence across changing oncology records

SentryLine grounds oncology answers in the specific document version they came from, targeting question answering over patient records that keep changing across a course of care. Version-aware evidence separates a usable clinical answer from one that was true three notes ago.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 95  ·  09 Sep 2026

Editor's Note

A proof machine-checked in Lean, nine billion variants precomputed, and biology still validating the old way.

 

Nº 01 · The Lede  —  OpenAI  —  Agents · Infrastructure

OpenAI agents produce a machine-checked Navier-Stokes proof

OpenAI agents produce a machine-checked Navier-Stokes proof

Fig. I  OpenAI · Filed 09 Sep 2026.

OpenAI published a claimed solution to the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems, produced by a group of agents running on an unreleased next-generation model. The writeup ships with a formal proof in Lean, a proof-assistant language that machine-checks every step, so the claim doesn't rest on anyone trusting the model's prose. Mathematicians will spend weeks on the informal argument regardless. For biology the sharp part is the verification asymmetry: math has a machine checker, and discovery claims in biology have only the bench, which resets how much scrutiny an agent-generated hypothesis has to survive before anyone spends reagents on it.

Read the source →

Why it matters

Formal verification is what makes an agent-generated result checkable without human trust, and biology has no equivalent, so every agent-generated hypothesis still buys its credibility with wet-lab time.

The Bench NoteFrom Heureka Labs

Biology checks an agent's result the way it checks a colleague's: by reading how the work was done.

What lands with an analysis. ARC reads the data, writes the code, runs it and writes up the outcome into the project's Analysis folder, so the script sits beside the figure it produced.
Lineage, recorded live. Bench links animals, samples, reagents, experiments and datasets while the work happens, and a methods note with a provenance table comes out of any record's lineage.
Where the limit sits. Neither of those settles whether a finding is true; they shorten the distance to finding out.

What we’re watching: whether the field converges on what an agent should hand over with a result before anyone calls it a finding

 

Nº 02  —  X  —  Field report

DeepMind maps every possible single-letter DNA change

Fig. II  X · Filed 09 Sep 2026.

DeepMind maps every possible single-letter DNA change

Google DeepMind launched AlphaGenome Atlas, a searchable database of predicted functional impact for all 9 billion possible single-letter changes in human DNA. Every position, every substitution, precomputed and queryable rather than run on demand. Coverage that complete moves variant-effect prediction from a modeling exercise to a lookup, echoing the same shift a proteome-wide binder atlas made for binder design, and it changes what counts as a reasonable first pass on a variant of uncertain significance and on the long tail of GWAS hits nobody had compute to interrogate.

Read more →

 

Nº 03  —  bioRxiv  —  Drug discovery · Computational

bioq gives AI agents one drug-discovery command line

Fig. III  bioRxiv · Filed 09 Sep 2026.

bioq gives AI agents one drug-discovery command line

bioq unifies drug-discovery tools behind a single command-line interface designed for agents to call rather than people to type, exposing what the bioRxiv preprint describes as a fleet of AI methods through one consistent grammar. The bottleneck is real: most published discovery models ship as one-off repos with incompatible inputs and outputs, and any agent chaining them pays the glue-code tax. Standardizing that interface is what turns a pile of methods into something an agent can plan over, and it lowers the integration cost that has kept much of the published toolkit unused.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  bioRxiv  —  Agents · Infrastructure

Agents take over model building for RNA structures

Agents did the model development for RNA 3D structure prediction in a new bioRxiv preprint, iterating on design choices rather than running an existing predictor. Pushes agents up the stack into building the models for biology's least-solved structure problem.

Read →

Nº 05  —  arXiv  —  Clinical AI · Evaluation

Study compares clinical AI, physicians and frontier models

A three-way diagnostic comparison puts a purpose-built clinical AI system against practicing physicians and general-purpose frontier language models on primary care cases. Head-to-head framing like this is what settles whether medical AI needs bespoke systems or whether general models have closed the gap.

Read →

Nº 06  —  arXiv  —  Clinical AI · Evaluation

SentryLine tracks evidence across changing oncology records

SentryLine grounds oncology answers in the specific document version they came from, targeting question answering over patient records that keep changing across a course of care. Version-aware evidence separates a usable clinical answer from one that was true three notes ago.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.