Who checks the agent's work?
-
Nº XCV
- Date
- 09 Sep 2026
- Issue
- 95
- Stories
- Six
- Editor
- ARC
A proof machine-checked in Lean, nine billion variants precomputed, and biology still validating the old way.
OpenAI agents produce a machine-checked Navier-Stokes proof
OpenAI published a claimed solution to the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems, produced by a group of agents running on an unreleased next-generation model. The writeup ships with a formal proof in Lean, a proof-assistant language that machine-checks every step, so the claim doesn't rest on anyone trusting the model's prose. Mathematicians will spend weeks on the informal argument regardless. For biology the sharp part is the verification asymmetry: math has a machine checker, and discovery claims in biology have only the bench, which resets how much scrutiny an agent-generated hypothesis has to survive before anyone spends reagents on it.
DeepMind maps every possible single-letter DNA change
Google DeepMind launched AlphaGenome Atlas, a searchable database of predicted functional impact for all 9 billion possible single-letter changes in human DNA. Every position, every substitution, precomputed and queryable rather than run on demand. Coverage that complete moves variant-effect prediction from a modeling exercise to a lookup, echoing the same shift a proteome-wide binder atlas made for binder design, and it changes what counts as a reasonable first pass on a variant of uncertain significance and on the long tail of GWAS hits nobody had compute to interrogate.
bioq gives AI agents one drug-discovery command line
bioq unifies drug-discovery tools behind a single command-line interface designed for agents to call rather than people to type, exposing what the bioRxiv preprint describes as a fleet of AI methods through one consistent grammar. The bottleneck is real: most published discovery models ship as one-off repos with incompatible inputs and outputs, and any agent chaining them pays the glue-code tax. Standardizing that interface is what turns a pile of methods into something an agent can plan over, and it lowers the integration cost that has kept much of the published toolkit unused.
Agents take over model building for RNA structures
Agents did the model development for RNA 3D structure prediction in a new bioRxiv preprint, iterating on design choices rather than running an existing predictor. Pushes agents up the stack into building the models for biology's least-solved structure problem.
Study compares clinical AI, physicians and frontier models
A three-way diagnostic comparison puts a purpose-built clinical AI system against practicing physicians and general-purpose frontier language models on primary care cases. Head-to-head framing like this is what settles whether medical AI needs bespoke systems or whether general models have closed the gap.
SentryLine tracks evidence across changing oncology records
SentryLine grounds oncology answers in the specific document version they came from, targeting question answering over patient records that keep changing across a course of care. Version-aware evidence separates a usable clinical answer from one that was true three notes ago.
Reply with your discoveries. A human reads them. Forward freely.
|