Weekly Arc: OpenAI's new biology tests only measure risk

Last week's issue noted that GPT-6 Astra arrived with best-published scores on research-level math, abstract reasoning and long command-line tasks, and nothing at all on biology. The tests measuring biology were being written by everyone except the labs shipping the largest models. That gap closed this week, in a direction worth sitting with. OpenAI shipped GPT-6 Astra with three new evaluations, all of them aimed at advanced biological capability, and folded them into Preparedness, the company's internal system for deciding whether a model is too dangerous to release. The frontier labs can write biology evaluations. They are writing them to measure what a model might help a bad actor do, not whether it can do a scientist's job.
Biology enters the frontier scorecard as a risk category
Preparedness is a threshold system. It asks whether a model crosses a line past which release requires extra safeguards, and the three new evaluations put biology squarely inside that question for the first time at this scale. Anthropic moved the same way, publishing its most detailed account yet of attempted misuse, covering cyberattacks, surveillance, influence operations and biology. Both are useful work and both are the right thing to measure.
The asymmetry is what matters for anyone reading model releases. A frontier lab now has a documented, internal, biology-specific evaluation capability. It points at harm. The question a working scientist actually has, which is whether this model is any good at the analysis sitting on their desk, still gets answered by third parties: expert-written rubrics like Karenina, which let specialists score an answer along several axes instead of collapsing quality into one accuracy number, or a benchmark for long multimodal research tasks spanning figures, tables and text. The labs grade the downside. The field grades the upside, on its own budget.
When the check can ship alongside the claim
The strongest thing OpenAI published this week was not a benchmark score. It was a claimed solution to the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems, produced by a group of agents running on an unreleased model and shipped with a formal machine-checked proof. A proof assistant re-derives every step mechanically. Accepting the argument does not require trusting the agents, the model, or the company that owns both.
That is not the end of the matter. By the time the story developed two days later, mathematicians had not yet signed off. A machine check confirms an argument is internally valid; it does not confirm that the formal statement proved is the statement the field cares about, and translating between the two is human work. Even with that caveat, the bar here sits far above anything biology can currently offer. Mathematics has a checker that runs in minutes and costs nothing to rerun.
Biology's biggest claim of the week came without one
Insilico Medicine reported that rentosertib shifted six proteomic aging clocks toward younger predicted biological age. This is an AI-designed molecule that has already reached patients, which is more than nearly any other molecule in this category can say, and the readout comes from real human samples. It was also the most-clicked story of the week by a wide margin, which is the right instinct.
It is also worth being precise about what the evidence is. An aging clock is itself a predictive model that estimates biological age from protein levels. So a molecule designed by models was assessed by models, and the result was reported by the company that owns the design pipeline. None of that makes the finding wrong. It means the claim and its verification are not going to arrive together, and the gap between them may run for years. Compare the shape of the answer in the clinic, where a three-way comparison put a purpose-built clinical AI system against practicing physicians and general-purpose frontier models on the same primary care cases. Head-to-head, one protocol, someone other than the vendor running it. That is expensive, slow, and the only thing that settles anything.
Before auditing an agent, measure how much it wobbles
The sharpest complication came from a quieter paper. An arXiv analysis argued that fairness audits of clinical agents need a noise floor: a multi-step agent produces different output on different runs of the same case, and unless that run-to-run instability is measured first, an apparent bias may just be variance. The finding generalizes past fairness. Any audit of an agent is measuring a distribution, not a result, and most published agent evaluations report a single run. A related preprint attacked the same instability at its source, training genomics agents by enumerating tool choices across a finite set rather than sampling random paths, on the grounds that tool selection is where agent-run analyses quietly go wrong.
So the principle this arc has been circling sharpens to something usable. The question to ask of any result this year is not whether a model produced it. It is who is positioned to check it, how much that check costs, and whether the checker has any stake in the answer. Where checking is cheap and mechanical, as in the Navier-Stokes proof, verification travels with the claim and skepticism is inexpensive. Where checking requires a wet lab, a second cohort, or a hundred reruns of a nondeterministic agent, the claim travels alone and the audit arrives late or not at all. A five-year retrospective on machine-learning-guided protein engineering is exactly that late audit, tallying which sequence-function methods actually improved campaigns after half a decade of confident announcements.
The labs have now demonstrated they will build biology evaluations when a question matters enough to them. So far the question that matters is misuse. Whether competence evaluations follow from the same source, or stay the unfunded work of everyone downstream, is the thing to watch next.
- Agents as bench scientists: the work moved up another level this week, with agents doing the model development for RNA 3D structure prediction rather than running an existing predictor, GPT-5.6 Sol calibrating qubits in an MIT lab, and AI-proposed lung cancer target pairs holding up in early experiments — watch whether a method built by an agent gets described in a paper as the agent's design choice or the authors'.
- The open stack and its liabilities: safety work shifted from retrospective log review to pre-release gates with OpenAI's advanced-biocapability evaluations and Anthropic's threat intelligence report — the open loop is whether any outside party ever gets to rerun one of those gates.
- The infrastructure nobody owns: July's AlphaFold shutdown thread got its answer as DeepMind shipped a one-petabyte genome atlas thirty times the size of the AlphaFold Database, while methods began arriving as things agents call rather than things labs install — an epigenomics toolkit shipped as a skill, bioq as one agent-facing command line, and OpenAI's managed Agents API.
Reply with what you're seeing. A human reads them. Forward freely.
|