6 min read

Virtual cells fail a mechanism test

Virtual cells fail a mechanism test
Nº 01 · The Lede bioRxiv Drug discovery · Computational

Benchmark finds AI cell models miss drug mechanisms

Benchmark finds AI cell models miss drug mechanisms
Fig. IbioRxiv · Filed 25 Aug 2026.

A mechanism-annotated benchmark finds single-cell perturbation models show limited fidelity to real drug-response signatures, per a new bioRxiv preprint. Perturbation models, trained to predict how cells shift transcriptionally under a compound or a genetic knockout, are the core bet behind virtual-cell programs across the field. By tagging compounds with their known mechanism of action and testing whether predicted profiles recover the matching signature, the benchmark measures something aggregate expression-correlation scores don't. The models come up short. That resets the reference test for perturbation prediction and moves the debate from how closely predictions correlate to whether they carry mechanism at all.

Read the source

CodonMamba designs mRNA coding sequences to order
Fig. IIbioRxiv · Filed 25 Aug 2026.
Nº 02 bioRxiv Field report

CodonMamba designs mRNA coding sequences to order

CodonMamba tunes mRNA coding sequences as a design variable rather than a fixed readout of a protein sequence. A new bioRxiv preprint describes it as a foundation model (trained broadly across coding sequences rather than fitted to one gene) aimed at programmable design: specify the protein, get codon choices tuned toward the expression behavior you want. Codon optimization has run on heuristics and vendor black boxes for two decades, so moving it onto a general sequence model gives mRNA therapeutics and protein-expression work a shared, inspectable starting point.

Read more
A thread argues AI shrinks the optimal biotech company
Fig. IIIX · Filed 25 Aug 2026.
Nº 03 X Computational biology

A thread argues AI shrinks the optimal biotech company

A widely-shared X thread on Anthropic's Claude protein-design results argues the optimal biotech company gets "shockingly small" — a couple of great scientists supervising model-driven design instead of a full discovery organization. The claim is structural rather than technical: if binder generation and optimization stop consuming headcount, the constraint moves to wet-lab validation and judgment. Read it as a position, not a finding. It is the sharpest current version of an org-design debate we flagged now running alongside every AI-discovery capability claim.

Read more
Also Filed · Three Briefs from the queue
Nº 04 arXiv Agents · Infrastructure

Survey maps how molecular AI agents are built

A new survey maps molecular LLM agents from architectural design through what the authors call scientific autonomy, sorting how such systems chain tools, memory, and planning across chemistry tasks. It hands a subfield where every group ships its own stack a shared vocabulary for comparison.

Read
Nº 05 arXiv Field report

Study compares two LLM routes to O-RADS scoring

O-RADS ovarian risk scoring from free-text ultrasound reports gets a head-to-head test: one model handling the whole job versus a hybrid that splits report parsing from rule-based scoring. The comparison speaks to how much deterministic logic clinical decision support should keep outside the model.

Read
Nº 06 X Field report

Rig crosses two million downloads, credits the plumbing

Rig crossed two million downloads — the open-source agent framework's maintainer used the milestone to argue that reliable agents come from the infrastructure around the model, not the model itself. The plumbing under biology agents is consolidating faster than the models running on top of it.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 84  ·  25 Aug 2026

Editor's Note

A benchmark takes the shine off virtual cells, while codon design finally gets a general model.

 

Nº 01 · The Lede  —  bioRxiv  —  Drug discovery · Computational

Benchmark finds AI cell models miss drug mechanisms

Benchmark finds AI cell models miss drug mechanisms

Fig. I  bioRxiv · Filed 25 Aug 2026.

A mechanism-annotated benchmark finds single-cell perturbation models show limited fidelity to real drug-response signatures, per a new bioRxiv preprint. Perturbation models, trained to predict how cells shift transcriptionally under a compound or a genetic knockout, are the core bet behind virtual-cell programs across the field. By tagging compounds with their known mechanism of action and testing whether predicted profiles recover the matching signature, the benchmark measures something aggregate expression-correlation scores don't. The models come up short. That resets the reference test for perturbation prediction and moves the debate from how closely predictions correlate to whether they carry mechanism at all.

Read the source →

Why it matters

Mechanism recovery becomes the number a perturbation model has to post, and expression-correlation leaderboards lose their standing as evidence that virtual-cell prediction works.

 

Nº 02  —  bioRxiv  —  Field report

CodonMamba designs mRNA coding sequences to order

Fig. II  bioRxiv · Filed 25 Aug 2026.

CodonMamba designs mRNA coding sequences to order

CodonMamba tunes mRNA coding sequences as a design variable rather than a fixed readout of a protein sequence. A new bioRxiv preprint describes it as a foundation model (trained broadly across coding sequences rather than fitted to one gene) aimed at programmable design: specify the protein, get codon choices tuned toward the expression behavior you want. Codon optimization has run on heuristics and vendor black boxes for two decades, so moving it onto a general sequence model gives mRNA therapeutics and protein-expression work a shared, inspectable starting point.

Read more →

 

Nº 03  —  X  —  Computational biology

A thread argues AI shrinks the optimal biotech company

Fig. III  X · Filed 25 Aug 2026.

A thread argues AI shrinks the optimal biotech company

A widely-shared X thread on Anthropic's Claude protein-design results argues the optimal biotech company gets "shockingly small" — a couple of great scientists supervising model-driven design instead of a full discovery organization. The claim is structural rather than technical: if binder generation and optimization stop consuming headcount, the constraint moves to wet-lab validation and judgment. Read it as a position, not a finding. It is the sharpest current version of an org-design debate we flagged now running alongside every AI-discovery capability claim.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  arXiv  —  Agents · Infrastructure

Survey maps how molecular AI agents are built

A new survey maps molecular LLM agents from architectural design through what the authors call scientific autonomy, sorting how such systems chain tools, memory, and planning across chemistry tasks. It hands a subfield where every group ships its own stack a shared vocabulary for comparison.

Read →

Nº 05  —  arXiv  —  Field report

Study compares two LLM routes to O-RADS scoring

O-RADS ovarian risk scoring from free-text ultrasound reports gets a head-to-head test: one model handling the whole job versus a hybrid that splits report parsing from rule-based scoring. The comparison speaks to how much deterministic logic clinical decision support should keep outside the model.

Read →

Nº 06  —  X  —  Field report

Rig crosses two million downloads, credits the plumbing

Rig crossed two million downloads — the open-source agent framework's maintainer used the milestone to argue that reliable agents come from the infrastructure around the model, not the model itself. The plumbing under biology agents is consolidating faster than the models running on top of it.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.