7 min read

GPT-6 lands with no biology score

GPT-6 lands with no biology score
Nº 01 · The Lede X Benchmarks · Evaluation

OpenAI ships GPT-6 Astra with science benchmark claims

OpenAI ships GPT-6 Astra with science benchmark claims
Fig. IX · Filed 04 Sep 2026.

OpenAI launched GPT-6 Astra with claimed state-of-the-art results on FrontierMath Tier 4 (research-level math problems), ARC-AGI 3 (abstract reasoning puzzles), and TerminalBench-4.0 (long-running command-line tasks), pitching the model as a major advance for scientific discovery. Rollout starts with a limited set of organizations and widens over the coming days to ChatGPT Plus, Pro, Business, and Enterprise tiers plus the API. All three named benchmarks are OpenAI's own reported numbers, and none of them is a biology task. Terminal competence is the one that touches research agents most directly, since most computational biology is a shell session with extra steps. The ceiling that biology agents inherit for free just moved again.

Read the source

Preprint encodes the physical lab for checkable workflows
Fig. IIarXiv · Filed 04 Sep 2026.
Nº 02 arXiv Field report

Preprint encodes the physical lab for checkable workflows

A machine-readable model of the lab anchors a new arXiv preprint, which proposes encoding benches, instruments, containers, and reagents in a formal representation software can reason over. The point is verification: a protocol gets checked against the actual state of a room before anything moves, instead of failing at execution. Automation has mostly meant scripting individual instruments one at a time. Formalizing the room itself is what makes autonomous experiment planning auditable rather than hopeful.

Read more
Bio-Babel rebuilds bio software so agents can run it
Fig. IIIbioRxiv · Filed 04 Sep 2026.
Nº 03 bioRxiv Agents · Infrastructure

Bio-Babel rebuilds bio software so agents can run it

Bio-Babel rebuilds computational biology tools across programming languages without human supervision, per a new bioRxiv preprint. The target is software sprawl: decades of packages written in different languages with incompatible interfaces, most of which an autonomous system cannot call. Reconstructing them into agent-ready form attacks the plumbing problem rather than the reasoning one, and plumbing is where autonomous analysis usually breaks. If it holds up, the callable surface for research agents widens from a curated handful of tools toward the corpus that already exists.

Read more
Also Filed · Three Briefs from the queue
Nº 04 bioRxiv Field report

Language models add annotation layers to scRNA-seq data

Pre-trained language models enrich scRNA-seq with additional layers of information in a bioRxiv preprint, pulling signal past what the count matrix carries on its own. Model-derived annotation as a routine step shifts what counts as a finished single-cell dataset.

Read
Nº 05 arXiv Field report

New analysis tests how prompts shape toxicity predictions

Prompt engineering gets measured against drug toxicity prediction in a new arXiv analysis, testing how phrasing choices change what general-purpose models report about compound safety. Quantifying that sensitivity sets a condition on whether model-based toxicity screens can be reported as results rather than demos.

Read
Nº 06 Hacker News Field report

Y Combinator backs personalized peptide and GLP-1 startup

RonanRX launched out of Y Combinator's S26 batch with personalized peptides and GLP-1s, posted to Hacker News. Personalizing an entire metabolic drug class raises the evidence bar for what the word means, in a category where personalization claims have so far traveled well ahead of trial data.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 92  ·  04 Sep 2026

Editor's Note

A frontier model shipped with math and terminal scores, and not one biology number in sight.

 

Nº 01 · The Lede  —  X  —  Benchmarks · Evaluation

OpenAI ships GPT-6 Astra with science benchmark claims

OpenAI ships GPT-6 Astra with science benchmark claims

Fig. I  X · Filed 04 Sep 2026.

OpenAI launched GPT-6 Astra with claimed state-of-the-art results on FrontierMath Tier 4 (research-level math problems), ARC-AGI 3 (abstract reasoning puzzles), and TerminalBench-4.0 (long-running command-line tasks), pitching the model as a major advance for scientific discovery. Rollout starts with a limited set of organizations and widens over the coming days to ChatGPT Plus, Pro, Business, and Enterprise tiers plus the API. All three named benchmarks are OpenAI's own reported numbers, and none of them is a biology task. Terminal competence is the one that touches research agents most directly, since most computational biology is a shell session with extra steps. The ceiling that biology agents inherit for free just moved again.

Read the source →

Why it matters

Frontier reasoning gains arrive inside biology agents by default, but no bio-native evaluation of Astra exists yet, so vendor benchmark claims are the only thing the field can price against.

The Bench NoteFrom Heureka Labs

General reasoning improvements reach biologists as better answers inside software they already run, rather than as a paper to read.

What a credit buys. Credits pay for thinking and computing, priced by the work a job takes: a substantial dataset analysis typically runs about 750 credits, with the middle half landing between 475 and 1,200.
How hard it thinks. The depth selector in the ARC panel sets Standard or Deep for a session, and a Deep session delegates its sub-tasks at Deep as well.
Where results land. Run Analysis writes the figures, the tables and the script that produced them into the project folder on your own disk, so a number and its code sit together.

What we’re watching: whether cost per finished analysis moves as general reasoning improves, or only the size of question worth attempting

 

Nº 02  —  arXiv  —  Field report

Preprint encodes the physical lab for checkable workflows

Fig. II  arXiv · Filed 04 Sep 2026.

Preprint encodes the physical lab for checkable workflows

A machine-readable model of the lab anchors a new arXiv preprint, which proposes encoding benches, instruments, containers, and reagents in a formal representation software can reason over. The point is verification: a protocol gets checked against the actual state of a room before anything moves, instead of failing at execution. Automation has mostly meant scripting individual instruments one at a time. Formalizing the room itself is what makes autonomous experiment planning auditable rather than hopeful.

Read more →

 

Nº 03  —  bioRxiv  —  Agents · Infrastructure

Bio-Babel rebuilds bio software so agents can run it

Fig. III  bioRxiv · Filed 04 Sep 2026.

Bio-Babel rebuilds bio software so agents can run it

Bio-Babel rebuilds computational biology tools across programming languages without human supervision, per a new bioRxiv preprint. The target is software sprawl: decades of packages written in different languages with incompatible interfaces, most of which an autonomous system cannot call. Reconstructing them into agent-ready form attacks the plumbing problem rather than the reasoning one, and plumbing is where autonomous analysis usually breaks. If it holds up, the callable surface for research agents widens from a curated handful of tools toward the corpus that already exists.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  bioRxiv  —  Field report

Language models add annotation layers to scRNA-seq data

Pre-trained language models enrich scRNA-seq with additional layers of information in a bioRxiv preprint, pulling signal past what the count matrix carries on its own. Model-derived annotation as a routine step shifts what counts as a finished single-cell dataset.

Read →

Nº 05  —  arXiv  —  Field report

New analysis tests how prompts shape toxicity predictions

Prompt engineering gets measured against drug toxicity prediction in a new arXiv analysis, testing how phrasing choices change what general-purpose models report about compound safety. Quantifying that sensitivity sets a condition on whether model-based toxicity screens can be reported as results rather than demos.

Read →

Nº 06  —  Hacker News  —  Field report

Y Combinator backs personalized peptide and GLP-1 startup

RonanRX launched out of Y Combinator's S26 batch with personalized peptides and GLP-1s, posted to Hacker News. Personalizing an entire metabolic drug class raises the evidence bar for what the word means, in a category where personalization claims have so far traveled well ahead of trial data.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.