7 min read

Agents plug into the lab bench

Agents plug into the lab bench
Nº 01 · The Lede X Agents · Infrastructure

Anthropic ships a standard for agents to run hardware

Anthropic ships a standard for agents to run hardware
Fig. IX · Filed 31 Aug 2026.

Anthropic shipped MHS , a standard for letting AI agents operate physical hardware, aimed at a problem every automated lab hits: wiring a model to an instrument takes days or weeks of bespoke integration today, with no common way for an agent to run equipment safely. Anthropic says MHS cuts that to hours or minutes and gives agents a defined interface to the machine instead of a one-off script per device. Analysis agents have matured fast while the pipetting robot stayed on its own island. A shared hardware layer is what moves closed-loop experimentation from single-site demos to something portable across benches and vendors.

Read the source

CIViC-Fact verifies cancer variant claims against sources
Fig. IIbioRxiv · Filed 31 Aug 2026.
Nº 02 bioRxiv Field report

CIViC-Fact verifies cancer variant claims against sources

CIViC-Fact verifies variant interpretations , a proof-of-concept framework that puts a language model to work checking whether curated claims about cancer variants hold up against the literature they cite. CIViC's knowledgebase is hand-curated and, like every clinical knowledgebase, drifts as evidence accumulates. Framing the job as verification rather than generation matters: the model audits existing entries instead of writing new ones. That reframes where AI first earns trust in clinical genomics, as a second reader on curated evidence long before it gets to curate.

Read more
Claude trains models to fix their own alignment failures
Fig. IIIAnthropic · Filed 31 Aug 2026.
Nº 03 Anthropic Field report

Claude trains models to fix their own alignment failures

Claude autonomously fixed alignment failures across all 10 categories covered by a set of public benchmarks, Anthropic reports, with the fixes improving each target score without degrading general capability. The setup is an agent running the full research loop itself: propose a change, train, evaluate, repeat. That is the same loop biology wants for binder design and screen analysis, tested here on a problem where the failure mode is legible and the scoring is fixed. It is the clearest evidence yet that autonomous research loops close on something harder than toy tasks.

Read more
Also Filed · Five Briefs from the queue
Nº 04 bioRxiv Structural biology · Protein design

RegimeFormer learns how proteins shift under perturbation

RegimeFormer models perturbation regimes across proteins rather than one variant at a time, in a new bioRxiv preprint. Shifting the unit of prediction from single mutations to whole response states is what would make large protein models useful for reading perturbation screens.

Read
Nº 05 Anthropic Field report

Anthropic expands its support program for scientists

Anthropic expanded its scientist program , widening the support it offers researchers using its models in their work. Frontier labs courting academic biology directly changes who sets the terms of access, and how fast bench work inherits each new model generation.

Read
Nº 06 arXiv Clinical AI · Evaluation

Concept-guided training makes clinical models auditable

Concept-guided fine-tuning binds a clinical language model's predictions to named clinical concepts, so a reviewer can see what drove each output. Building auditability into the training objective instead of a post-hoc explanation is the version regulators keep asking for.

Read
Nº 07 arXiv Agents · Infrastructure

BALMS tests AI agents on months of patient data

BALMS benchmarks agents on longitudinal mental health sensing, scoring how well LLM-driven systems read months of passive sensor streams instead of a single snapshot. Long-horizon patient data is where agent evaluation has been thinnest.

Read
Nº 08 X Field report

DeepMind pilots double-blind safety tests of its models

Google DeepMind is piloting double-blind safety evaluations, in which neither the test prompts nor the model weights are revealed, letting outside evaluators probe frontier models blind. Blinded external assessment is standard practice in clinical trials and nearly absent from AI capability claims.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 88  ·  31 Aug 2026

Editor's Note

Monday, and the news is plumbing: a standard for agents to drive instruments, plus a variant-curation auditor.

 

Nº 01 · The Lede  —  X  —  Agents · Infrastructure

Anthropic ships a standard for agents to run hardware

Anthropic ships a standard for agents to run hardware

Fig. I  X · Filed 31 Aug 2026.

Anthropic shipped MHS , a standard for letting AI agents operate physical hardware, aimed at a problem every automated lab hits: wiring a model to an instrument takes days or weeks of bespoke integration today, with no common way for an agent to run equipment safely. Anthropic says MHS cuts that to hours or minutes and gives agents a defined interface to the machine instead of a one-off script per device. Analysis agents have matured fast while the pipetting robot stayed on its own island. A shared hardware layer is what moves closed-loop experimentation from single-site demos to something portable across benches and vendors.

Read the source →

Why it matters

A common instrument interface moves the cost of closed-loop experimentation from a per-device engineering project to a configuration step, which is the gate that has kept self-driving labs to a handful of well-funded sites.

 

Nº 02  —  bioRxiv  —  Field report

CIViC-Fact verifies cancer variant claims against sources

Fig. II  bioRxiv · Filed 31 Aug 2026.

CIViC-Fact verifies cancer variant claims against sources

CIViC-Fact verifies variant interpretations , a proof-of-concept framework that puts a language model to work checking whether curated claims about cancer variants hold up against the literature they cite. CIViC's knowledgebase is hand-curated and, like every clinical knowledgebase, drifts as evidence accumulates. Framing the job as verification rather than generation matters: the model audits existing entries instead of writing new ones. That reframes where AI first earns trust in clinical genomics, as a second reader on curated evidence long before it gets to curate.

Read more →

 

Nº 03  —  Anthropic  —  Field report

Claude trains models to fix their own alignment failures

Fig. III  Anthropic · Filed 31 Aug 2026.

Claude trains models to fix their own alignment failures

Claude autonomously fixed alignment failures across all 10 categories covered by a set of public benchmarks, Anthropic reports, with the fixes improving each target score without degrading general capability. The setup is an agent running the full research loop itself: propose a change, train, evaluate, repeat. That is the same loop biology wants for binder design and screen analysis, tested here on a problem where the failure mode is legible and the scoring is fixed. It is the clearest evidence yet that autonomous research loops close on something harder than toy tasks.

Read more →

 

Also Filed  ·  Five Briefs from the queue

Nº 04  —  bioRxiv  —  Structural biology · Protein design

RegimeFormer learns how proteins shift under perturbation

RegimeFormer models perturbation regimes across proteins rather than one variant at a time, in a new bioRxiv preprint. Shifting the unit of prediction from single mutations to whole response states is what would make large protein models useful for reading perturbation screens.

Read →

Nº 05  —  Anthropic  —  Field report

Anthropic expands its support program for scientists

Anthropic expanded its scientist program , widening the support it offers researchers using its models in their work. Frontier labs courting academic biology directly changes who sets the terms of access, and how fast bench work inherits each new model generation.

Read →

Nº 06  —  arXiv  —  Clinical AI · Evaluation

Concept-guided training makes clinical models auditable

Concept-guided fine-tuning binds a clinical language model's predictions to named clinical concepts, so a reviewer can see what drove each output. Building auditability into the training objective instead of a post-hoc explanation is the version regulators keep asking for.

Read →

Nº 07  —  arXiv  —  Agents · Infrastructure

BALMS tests AI agents on months of patient data

BALMS benchmarks agents on longitudinal mental health sensing, scoring how well LLM-driven systems read months of passive sensor streams instead of a single snapshot. Long-horizon patient data is where agent evaluation has been thinnest.

Read →

Nº 08  —  X  —  Field report

DeepMind pilots double-blind safety tests of its models

Google DeepMind is piloting double-blind safety evaluations, in which neither the test prompts nor the model weights are revealed, letting outside evaluators probe frontier models blind. Blinded external assessment is standard practice in clinical trials and nearly absent from AI capability claims.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.