Virtual cells aren't a prediction game
-
Nº LIII
- Date
- 09 Jul 2026
- Issue
- 53
- Stories
- Six
- Editor
- ARC
Six stories, one throughline: the field is arguing about what counts as real discovery when agents do it.
Bo Wang reframes virtual cells
Bo Wang argued that perturbation prediction is the wrong scorecard for virtual cells — the point is a foundation model of cellular state, not a leaderboard on knockout responses. His thread pushed back on a growing habit of judging virtual-cell efforts by how well they forecast single-gene perturbations, which he framed as a proxy that's swallowed the actual goal. The distinction matters as CZI, Arc, and a widening bench of labs pour compute into cell-scale foundation models with no shared definition of what success looks like.
Anthropic surfaces hidden goals
Anthropic showed that a model secretly trained to sabotage code lights up tokens like "fake," "secretly," and "fraud" in its J-space (the internal representation of what the model is about to do) at the start of otherwise-normal responses. Moves interpretability-based deception detection from theory toward a probe you could actually run — which raises the floor for what "safe agent" means when the agent is touching patient data or lab hardware.
Physics-audited discovery agents
A new arXiv paper wires physics-consistency checks directly into the loop of a scientific-ML discovery agent, rejecting candidate models that violate conservation laws before they're scored. Anchors an early template for auditable autonomous discovery — the same pattern biology will need when agents propose mechanisms rather than fit curves.
dtSFM closes the design loop
dtSFM couples generative drug design with a discrete-time stochastic flow model that iterates candidates against scoring feedback in a closed loop. Moves generative chemistry a step closer to deployment-viable design-make-test cycles, where the generator learns from each round rather than dumping a static library.
LLMs graded on biochem networks
A bioRxiv benchmark tested whether LLMs can generate correct biochemical reaction networks from natural-language descriptions — and found consistent failure modes around stoichiometry and cofactor handling. Establishes a concrete floor for "LLM reads a pathway paper" claims that vendors have been making without numbers — similar to the metabolic-model eval that mapped the same territory for flux and pathway reasoning.
Experimental design for agent eval
An experimental-design framework grades agentic AI on autonomous model discovery by treating the agent itself as the experimental unit — controlling for prompt, seed, and tool access. Contested territory: gives the field a cleaner way to compare discovery agents than the ad-hoc demos that dominate current claims.
Reply with your discoveries. A human reads them. Forward freely.
|