6 min read

Past the hypothesis, into the lab

Past the hypothesis, into the lab
Nº 01 · The Lede X Field report

Google's Co-Scientist moves past generating hypotheses

Google's Co-Scientist moves past generating hypotheses
Fig. IX · Filed 03 Sep 2026.

Google DeepMind extends Co-Scientist beyond hypothesis generation in biomedicine, Vivek Natarajan announced, describing work with outside collaborators built on Gemini. Co-Scientist was first shown as a multi-agent system (several models proposing and critiquing each other's ideas) that reads literature and returns testable hypotheses, the part of research cheapest to automate and hardest to trust. Pushing past that means the system takes on more of the arc after the idea rather than handing a ranked list to a human. Specifics are still thin. But the frontier question has already shifted from whether an agent can propose good biomedical hypotheses to whether it can carry one through to something a lab can act on.

Read the source

A compression trick runs Evo 2 on one GPU
Fig. IIbioRxiv · Filed 03 Sep 2026.
Nº 02 bioRxiv Field report

A compression trick runs Evo 2 on one GPU

Evo 2's full million-token context now fits on a single GPU, via a compression method that needs no calibration dataset to set up. Evo 2 is a genome foundation model, trained across DNA from many species, and its enormous read-at-once window is the whole point, since regulatory signal often sits far from the sequence it controls. Getting that onto one card drops genome-scale inference from cluster budgets to hardware a single group already owns.

Read more
Redeployment tests map where single-cell models stop working
Fig. IIIbioRxiv · Filed 03 Sep 2026.
Nº 03 bioRxiv Cell biology · Funding

Redeployment tests map where single-cell models stop working

Single-cell foundation models hit boundaries that only appear when someone rebuilds them from scratch. The bioRxiv work redeploys published single-cell foundation models (general models trained on millions of cells and meant to transfer across tissues and tasks) in a standardized, reproducible setup, then reports where their practical limits fall. Reproducible redeployment as the test method raises the floor for evidence in the virtual-cell debate, where reported benchmark scores have run well ahead of anyone's ability to reproduce them.

Read more
Also Filed · Three Briefs from the queue
Nº 04 X Agents · Infrastructure

Google ships two Gemini models for agent work

Google DeepMind shipped two Gemini models aimed at scaling agents and securing code, including 3.8 Flash, claimed to gain on 3.7 Flash across software engineering and multi-step agent tasks. Fast, cheap models are what long-running analysis agents burn through, so the cost floor for autonomous biological work moves with each Flash release.

Read
Nº 05 arXiv Field report

NS-Copilot runs neuroscience analyses without a human driver

NS-Copilot automates neuroscience analysis end to end, an agent system described in a new arXiv preprint that plans and executes the analysis itself. Field-specific analysis agents keep landing one discipline at a time, and neuroscience data handling is among the least standardized to hand over.

Read
Nº 06 arXiv Field report

Four sites tune a chest X-ray model without pooling data

Federated fine-tuning adapts BiomedCLIP across four international chest X-ray cohorts, updating a shared medical image-text model with LoRA (small trainable weight patches) instead of moving patient scans between sites. Keeps multi-site imaging models trainable in the many settings where data-sharing agreements never close.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 91  ·  03 Sep 2026

Editor's Note

Hypothesis machines look for a second act, and single-cell models get handed their reproducibility bill.

 

Nº 01 · The Lede  —  X  —  Field report

Google's Co-Scientist moves past generating hypotheses

Google's Co-Scientist moves past generating hypotheses

Fig. I  X · Filed 03 Sep 2026.

Google DeepMind extends Co-Scientist beyond hypothesis generation in biomedicine, Vivek Natarajan announced, describing work with outside collaborators built on Gemini. Co-Scientist was first shown as a multi-agent system (several models proposing and critiquing each other's ideas) that reads literature and returns testable hypotheses, the part of research cheapest to automate and hardest to trust. Pushing past that means the system takes on more of the arc after the idea rather than handing a ranked list to a human. Specifics are still thin. But the frontier question has already shifted from whether an agent can propose good biomedical hypotheses to whether it can carry one through to something a lab can act on.

Read the source →

Why it matters

Hypothesis generation was the crowded half of research automation; the best-resourced group working on it just staked ground on the half that actually reaches the bench, which resets what a credible research-agent claim has to include.

 

Nº 02  —  bioRxiv  —  Field report

A compression trick runs Evo 2 on one GPU

Fig. II  bioRxiv · Filed 03 Sep 2026.

A compression trick runs Evo 2 on one GPU

Evo 2's full million-token context now fits on a single GPU, via a compression method that needs no calibration dataset to set up. Evo 2 is a genome foundation model, trained across DNA from many species, and its enormous read-at-once window is the whole point, since regulatory signal often sits far from the sequence it controls. Getting that onto one card drops genome-scale inference from cluster budgets to hardware a single group already owns.

Read more →

 

Nº 03  —  bioRxiv  —  Cell biology · Funding

Redeployment tests map where single-cell models stop working

Fig. III  bioRxiv · Filed 03 Sep 2026.

Redeployment tests map where single-cell models stop working

Single-cell foundation models hit boundaries that only appear when someone rebuilds them from scratch. The bioRxiv work redeploys published single-cell foundation models (general models trained on millions of cells and meant to transfer across tissues and tasks) in a standardized, reproducible setup, then reports where their practical limits fall. Reproducible redeployment as the test method raises the floor for evidence in the virtual-cell debate, where reported benchmark scores have run well ahead of anyone's ability to reproduce them.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  X  —  Agents · Infrastructure

Google ships two Gemini models for agent work

Google DeepMind shipped two Gemini models aimed at scaling agents and securing code, including 3.8 Flash, claimed to gain on 3.7 Flash across software engineering and multi-step agent tasks. Fast, cheap models are what long-running analysis agents burn through, so the cost floor for autonomous biological work moves with each Flash release.

Read →

Nº 05  —  arXiv  —  Field report

NS-Copilot runs neuroscience analyses without a human driver

NS-Copilot automates neuroscience analysis end to end, an agent system described in a new arXiv preprint that plans and executes the analysis itself. Field-specific analysis agents keep landing one discipline at a time, and neuroscience data handling is among the least standardized to hand over.

Read →

Nº 06  —  arXiv  —  Field report

Four sites tune a chest X-ray model without pooling data

Federated fine-tuning adapts BiomedCLIP across four international chest X-ray cohorts, updating a shared medical image-text model with LoRA (small trainable weight patches) instead of moving patient scans between sites. Keeps multi-site imaging models trainable in the many settings where data-sharing agreements never close.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.