Agents cross into original research
-
Nº XCVI
- Date
- 10 Sep 2026
- Issue
- 96
- Stories
- Seven
- Editor
- ARC
A Millennium Prize proof, a weekend genome browser, and target pairs that survived contact with a bench.
OpenAI says its agents proved a Millennium problem
OpenAI says agents solved Navier-Stokes existence and smoothness, one of the seven Millennium Prize problems, with a proof it credits to a group of agents running on an unreleased next-generation model. The proof is public; mathematicians haven't verified it yet, and the claim rests entirely on that review. What's notable for discovery science is the shape of the work: agents sustaining one problem long enough to produce an artifact another field can check line by line. Mathematics gives that structure an unusually clean scoreboard — a proof either holds or it doesn't, where an AI-generated biological hypothesis needs months of bench work to adjudicate. If it survives, the ceiling on what autonomous research systems are expected to produce moves in a single step.
AI-proposed lung cancer target pairs pass early tests
AI-proposed target pairs in lung squamous cell carcinoma held up in early experiments, per a bioRxiv preprint from the Emet AI Research Environment, a system built to generate and rank combination hypotheses rather than single targets. Discovery-stage assays back the proposals, which puts this among the few AI-generated oncology hypotheses arriving with wet-lab support instead of a benchmark score. Combination selection is where target discovery is thinnest, since the space of pairs is far too large to screen exhaustively.
A drug-property agent improves itself across long runs
ADMET-EvO keeps improving across a run instead of resetting at every task, an arXiv preprint describing a scientific agent that accumulates its own methods while working absorption, distribution, metabolism, excretion and toxicity problems. The heterogeneous setup is the real test: the agent has to carry what it learned on one endpoint into an unrelated one. Sustained self-improvement over long horizons is the missing piece between one-shot property prediction and a system that can be left running on a program for weeks.
A consumer DNA test becomes a 40-minute app
A consumer DNA array met GPT-6 Astra, OpenAI's newest reasoning model: raw Illumina data on roughly 660,000 variants in, an interactive interpretation tool out, built in 40 minutes. Variant interpretation is becoming a weekend build, clinical-grade standards notwithstanding.
New training method fixes agents' genomics tool choices
Enumerating tool choices beats sampling them, an arXiv preprint argues, training genomics agents by exact optimization over a finite tool set instead of random rollouts. Tool selection is where agent-run analyses quietly go wrong, so this raises the reliability floor without raising the compute bill.
ProteinSage adds structure rules to a protein model
ProteinSage writes structure in explicitly rather than hoping a protein language model absorbs it from sequence alone, a bioRxiv preprint reporting comparable modeling at lower cost. Cheaper protein models shift where structural priors belong: in the training objective, not only the data.
OpenAI's model calibrates qubits in an MIT lab
GPT-5.6 Sol runs experiments in an MIT quantum lab, per an OpenAI writeup, calibrating qubits and analyzing results through Codex, OpenAI's coding agent. Instrument-level autonomy is landing in physics first; that same closed loop is what bench automation still lacks.
Reply with your discoveries. A human reads them. Forward freely.
|