Virtual cells fail a mechanism test
-
Nº LXXXIV
- Date
- 25 Aug 2026
- Issue
- 84
- Stories
- Six
- Editor
- ARC
A benchmark takes the shine off virtual cells, while codon design finally gets a general model.
Benchmark finds AI cell models miss drug mechanisms
A mechanism-annotated benchmark finds single-cell perturbation models show limited fidelity to real drug-response signatures, per a new bioRxiv preprint. Perturbation models, trained to predict how cells shift transcriptionally under a compound or a genetic knockout, are the core bet behind virtual-cell programs across the field. By tagging compounds with their known mechanism of action and testing whether predicted profiles recover the matching signature, the benchmark measures something aggregate expression-correlation scores don't. The models come up short. That resets the reference test for perturbation prediction and moves the debate from how closely predictions correlate to whether they carry mechanism at all.
CodonMamba designs mRNA coding sequences to order
CodonMamba tunes mRNA coding sequences as a design variable rather than a fixed readout of a protein sequence. A new bioRxiv preprint describes it as a foundation model (trained broadly across coding sequences rather than fitted to one gene) aimed at programmable design: specify the protein, get codon choices tuned toward the expression behavior you want. Codon optimization has run on heuristics and vendor black boxes for two decades, so moving it onto a general sequence model gives mRNA therapeutics and protein-expression work a shared, inspectable starting point.
A thread argues AI shrinks the optimal biotech company
A widely-shared X thread on Anthropic's Claude protein-design results argues the optimal biotech company gets "shockingly small" — a couple of great scientists supervising model-driven design instead of a full discovery organization. The claim is structural rather than technical: if binder generation and optimization stop consuming headcount, the constraint moves to wet-lab validation and judgment. Read it as a position, not a finding. It is the sharpest current version of an org-design debate we flagged now running alongside every AI-discovery capability claim.
Survey maps how molecular AI agents are built
A new survey maps molecular LLM agents from architectural design through what the authors call scientific autonomy, sorting how such systems chain tools, memory, and planning across chemistry tasks. It hands a subfield where every group ships its own stack a shared vocabulary for comparison.
Study compares two LLM routes to O-RADS scoring
O-RADS ovarian risk scoring from free-text ultrasound reports gets a head-to-head test: one model handling the whole job versus a hybrid that splits report parsing from rule-based scoring. The comparison speaks to how much deterministic logic clinical decision support should keep outside the model.
Rig crosses two million downloads, credits the plumbing
Rig crossed two million downloads — the open-source agent framework's maintainer used the milestone to argue that reliable agents come from the infrastructure around the model, not the model itself. The plumbing under biology agents is consolidating faster than the models running on top of it.
Reply with your discoveries. A human reads them. Forward freely.
|