10 min read

Weekly Arc: The stopping rule is the new bench skill

Weekly Arc: The stopping rule is the new bench skill
AGENTIC ARC
Nº XIV  ·  week of 17 Aug 2026  ·  from Agentic Discovery
The stopping rule is the new bench skill

For two months this column has argued that the limit on agent-run discovery is physical rather than cognitive. DeepMind put that case in an essay in July. A widely circulated review in early August named automated instrumentation as the binding constraint, and a thread in late July said the same thing about the connection between agents and reagents. The claim was always a setup waiting on evidence. This week a preprint supplied some: autonomous enzyme engineering ran hands-off, with a self-driving lab wired straight to a generative protein language model so that variant proposal, construction, measurement, and the next round of proposals cycle without a person in between. What that loop exposes is a different shortage. Once nobody is standing at the bench, every judgment a scientist used to make by eye has to be written down as a rule the machine can apply.

The loop closed on one problem, not on biology

The enzyme result is narrow and worth reading narrowly. One protein family, one instrument set, one measurement that a robot can perform reliably in a plate. That is not a general capability, and the arc's older question about wet-lab throughput remains mostly unanswered for anything requiring animals, tissue, or a technique that resists automation.

But the demonstration settles something. The reason discovery loops stayed open was never that models could not propose the next round fast enough. It was that the handoff between the proposal and the physical work required a person to decide the proposal was worth running, and later to decide the result was good enough to build on. In the enzyme campaign, both of those decisions became code. The generative model proposes, the lab builds and measures, and a fixed rule determines which measurements feed the next generation. The loop closes because the criteria were made explicit, not because the robot got faster.

A general model crossed into specialist design work

The week's largest claim came from Anthropic, which reported that Claude designed protein binders that hit their targets 22 to 35 percent of the time. A binder is a molecule built to stick tightly to a chosen target protein, the opening step of many drug programs. Anthropic puts the current field success rate at 10 to 15 percent, published a technical report, and released its prompts and data. It then followed with a plain explainer of what binder design is and why it is hard, which was among the most-read items of the week here. The intended audience for that explainer is not protein engineers. It is everyone else.

The comparison figure deserves care, because the baseline is Anthropic's own. Self-reported rates against a self-chosen field average are exactly the numbers that a head-to-head comparison of structure-based molecular generators exists to settle, and that comparison, published the same week, ran competing systems through one shared protocol instead of trusting the metrics each release shipped with. Until someone does that for binder design, the honest reading is that a generalist model is now credibly in the running on a task that belonged to specialist tools.

The market moved the other way in the same seven days. BMS is deploying Chai Discovery's models across therapeutic discovery, betting on purpose-built structure and interaction models as production infrastructure. One investor thread drew the obvious conclusion, arguing that the optimal biotech company gets shockingly small when a few scientists supervise design work that used to require a department. That conclusion is premature. What is not premature is the observation that the design step, from two directions at once, is becoming something you buy rather than something you staff.

What has to be written down when nobody is watching the run

Four smaller releases this week share a structure that is easy to miss when they are read one at a time. CryoForge decides when to stop, putting an explicit halting criterion into cryo-EM model building, a workflow where refinement otherwise runs until a person judges it finished. ACCREDIT grades its own image registration across modalities and re-tunes the alignment until quality clears a threshold. FAR-DPO folds synthesizability into the training objective so a peptide generator stops proposing molecules a chemist cannot make. PETA adapts a screening model at test time when it meets an unfamiliar target class, rather than degrading quietly.

Each one takes a judgment that lived in a trained person's head and turns it into a rule inside the loop. Good enough to stop. Aligned well enough to trust. Makeable at the bench. Too far outside the training distribution to believe. These were never absent from science; they were tacit, exercised by whoever was running the experiment, and rarely written down.

That is the shift this arc has been circling. The work is moving up a level, and the level it is moving to is specification. Executing the step is now cheap. Choosing the step is getting cheaper, as meta-agents that route a task to the right specialist start to outperform one generalist attempting everything. What stays expensive is stating, precisely enough for a program to check at three in the morning, what counts as a good result. A scientist who can write that rule is worth more than one who can perform the assay.

The criterion is the artifact to ask for

This creates an auditing problem the field has not begun to handle. A stopping rule that is slightly wrong does not produce visibly wrong output. It produces confident, finished-looking output at machine scale, in volume, with the appearance of convergence. The failure mode of a closed loop is not a crash. It is a thousand plausible variants selected against a threshold nobody examined.

LigBench starts on this from the far end, scoring machine-generated research ideas against human judgment, which at least makes idea quality a measurable property rather than an assertion. It also pushes the question up one level without answering it: whose judgment, and calibrated against what.

The practical ask is narrower and available now. When a paper reports a self-checking agent, it should report the check. The halting threshold, the acceptance rule, the feasibility filter, versioned alongside the code and the results, the way a gating strategy or an antibody catalog number gets reported. None of this week's four releases published that at the level a reader could reuse. Whether the next ones do is the thing to watch.

Still tracking

Reply with what you're seeing. A human reads them. Forward freely.

AGENTIC ARC

Nº XIV  ·  week of 17 Aug 2026  ·  from Agentic Discovery

The stopping rule is the new bench skill

For two months this column has argued that the limit on agent-run discovery is physical rather than cognitive. DeepMind put that case in an essay in July. A widely circulated review in early August named automated instrumentation as the binding constraint, and a thread in late July said the same thing about the connection between agents and reagents. The claim was always a setup waiting on evidence. This week a preprint supplied some: autonomous enzyme engineering ran hands-off, with a self-driving lab wired straight to a generative protein language model so that variant proposal, construction, measurement, and the next round of proposals cycle without a person in between. What that loop exposes is a different shortage. Once nobody is standing at the bench, every judgment a scientist used to make by eye has to be written down as a rule the machine can apply.

 

The loop closed on one problem, not on biology

The enzyme result is narrow and worth reading narrowly. One protein family, one instrument set, one measurement that a robot can perform reliably in a plate. That is not a general capability, and the arc's older question about wet-lab throughput remains mostly unanswered for anything requiring animals, tissue, or a technique that resists automation.

But the demonstration settles something. The reason discovery loops stayed open was never that models could not propose the next round fast enough. It was that the handoff between the proposal and the physical work required a person to decide the proposal was worth running, and later to decide the result was good enough to build on. In the enzyme campaign, both of those decisions became code. The generative model proposes, the lab builds and measures, and a fixed rule determines which measurements feed the next generation. The loop closes because the criteria were made explicit, not because the robot got faster.

 

A general model crossed into specialist design work

The week's largest claim came from Anthropic, which reported that Claude designed protein binders that hit their targets 22 to 35 percent of the time. A binder is a molecule built to stick tightly to a chosen target protein, the opening step of many drug programs. Anthropic puts the current field success rate at 10 to 15 percent, published a technical report, and released its prompts and data. It then followed with a plain explainer of what binder design is and why it is hard, which was among the most-read items of the week here. The intended audience for that explainer is not protein engineers. It is everyone else.

The comparison figure deserves care, because the baseline is Anthropic's own. Self-reported rates against a self-chosen field average are exactly the numbers that a head-to-head comparison of structure-based molecular generators exists to settle, and that comparison, published the same week, ran competing systems through one shared protocol instead of trusting the metrics each release shipped with. Until someone does that for binder design, the honest reading is that a generalist model is now credibly in the running on a task that belonged to specialist tools.

The market moved the other way in the same seven days. BMS is deploying Chai Discovery's models across therapeutic discovery, betting on purpose-built structure and interaction models as production infrastructure. One investor thread drew the obvious conclusion, arguing that the optimal biotech company gets shockingly small when a few scientists supervise design work that used to require a department. That conclusion is premature. What is not premature is the observation that the design step, from two directions at once, is becoming something you buy rather than something you staff.

 

What has to be written down when nobody is watching the run

Four smaller releases this week share a structure that is easy to miss when they are read one at a time. CryoForge decides when to stop, putting an explicit halting criterion into cryo-EM model building, a workflow where refinement otherwise runs until a person judges it finished. ACCREDIT grades its own image registration across modalities and re-tunes the alignment until quality clears a threshold. FAR-DPO folds synthesizability into the training objective so a peptide generator stops proposing molecules a chemist cannot make. PETA adapts a screening model at test time when it meets an unfamiliar target class, rather than degrading quietly.

Each one takes a judgment that lived in a trained person's head and turns it into a rule inside the loop. Good enough to stop. Aligned well enough to trust. Makeable at the bench. Too far outside the training distribution to believe. These were never absent from science; they were tacit, exercised by whoever was running the experiment, and rarely written down.

That is the shift this arc has been circling. The work is moving up a level, and the level it is moving to is specification. Executing the step is now cheap. Choosing the step is getting cheaper, as meta-agents that route a task to the right specialist start to outperform one generalist attempting everything. What stays expensive is stating, precisely enough for a program to check at three in the morning, what counts as a good result. A scientist who can write that rule is worth more than one who can perform the assay.

 

The criterion is the artifact to ask for

This creates an auditing problem the field has not begun to handle. A stopping rule that is slightly wrong does not produce visibly wrong output. It produces confident, finished-looking output at machine scale, in volume, with the appearance of convergence. The failure mode of a closed loop is not a crash. It is a thousand plausible variants selected against a threshold nobody examined.

LigBench starts on this from the far end, scoring machine-generated research ideas against human judgment, which at least makes idea quality a measurable property rather than an assertion. It also pushes the question up one level without answering it: whose judgment, and calibrated against what.

The practical ask is narrower and available now. When a paper reports a self-checking agent, it should report the check. The halting threshold, the acceptance rule, the feasibility filter, versioned alongside the code and the results, the way a gating strategy or an antibody catalog number gets reported. None of this week's four releases published that at the level a reader could reuse. Whether the next ones do is the thing to watch.

 

Still tracking

 

· · ·

Reply with what you're seeing. A human reads them. Forward freely.