Past the hypothesis, into the lab
-
Nº XCI
- Date
- 03 Sep 2026
- Issue
- 91
- Stories
- Six
- Editor
- ARC
Hypothesis machines look for a second act, and single-cell models get handed their reproducibility bill.
Google's Co-Scientist moves past generating hypotheses
Google DeepMind extends Co-Scientist beyond hypothesis generation in biomedicine, Vivek Natarajan announced, describing work with outside collaborators built on Gemini. Co-Scientist was first shown as a multi-agent system (several models proposing and critiquing each other's ideas) that reads literature and returns testable hypotheses, the part of research cheapest to automate and hardest to trust. Pushing past that means the system takes on more of the arc after the idea rather than handing a ranked list to a human. Specifics are still thin. But the frontier question has already shifted from whether an agent can propose good biomedical hypotheses to whether it can carry one through to something a lab can act on.
A compression trick runs Evo 2 on one GPU
Evo 2's full million-token context now fits on a single GPU, via a compression method that needs no calibration dataset to set up. Evo 2 is a genome foundation model, trained across DNA from many species, and its enormous read-at-once window is the whole point, since regulatory signal often sits far from the sequence it controls. Getting that onto one card drops genome-scale inference from cluster budgets to hardware a single group already owns.
Redeployment tests map where single-cell models stop working
Single-cell foundation models hit boundaries that only appear when someone rebuilds them from scratch. The bioRxiv work redeploys published single-cell foundation models (general models trained on millions of cells and meant to transfer across tissues and tasks) in a standardized, reproducible setup, then reports where their practical limits fall. Reproducible redeployment as the test method raises the floor for evidence in the virtual-cell debate, where reported benchmark scores have run well ahead of anyone's ability to reproduce them.
Google ships two Gemini models for agent work
Google DeepMind shipped two Gemini models aimed at scaling agents and securing code, including 3.8 Flash, claimed to gain on 3.7 Flash across software engineering and multi-step agent tasks. Fast, cheap models are what long-running analysis agents burn through, so the cost floor for autonomous biological work moves with each Flash release.
NS-Copilot runs neuroscience analyses without a human driver
NS-Copilot automates neuroscience analysis end to end, an agent system described in a new arXiv preprint that plans and executes the analysis itself. Field-specific analysis agents keep landing one discipline at a time, and neuroscience data handling is among the least standardized to hand over.
Four sites tune a chest X-ray model without pooling data
Federated fine-tuning adapts BiomedCLIP across four international chest X-ray cohorts, updating a shared medical image-text model with LoRA (small trainable weight patches) instead of moving patient scans between sites. Keeps multi-site imaging models trainable in the many settings where data-sharing agreements never close.
Reply with your discoveries. A human reads them. Forward freely.
|