GPT-6 lands with no biology score
-
Nº XCII
- Date
- 04 Sep 2026
- Issue
- 92
- Stories
- Six
- Editor
- ARC
A frontier model shipped with math and terminal scores, and not one biology number in sight.
OpenAI ships GPT-6 Astra with science benchmark claims
OpenAI launched GPT-6 Astra with claimed state-of-the-art results on FrontierMath Tier 4 (research-level math problems), ARC-AGI 3 (abstract reasoning puzzles), and TerminalBench-4.0 (long-running command-line tasks), pitching the model as a major advance for scientific discovery. Rollout starts with a limited set of organizations and widens over the coming days to ChatGPT Plus, Pro, Business, and Enterprise tiers plus the API. All three named benchmarks are OpenAI's own reported numbers, and none of them is a biology task. Terminal competence is the one that touches research agents most directly, since most computational biology is a shell session with extra steps. The ceiling that biology agents inherit for free just moved again.
Preprint encodes the physical lab for checkable workflows
A machine-readable model of the lab anchors a new arXiv preprint, which proposes encoding benches, instruments, containers, and reagents in a formal representation software can reason over. The point is verification: a protocol gets checked against the actual state of a room before anything moves, instead of failing at execution. Automation has mostly meant scripting individual instruments one at a time. Formalizing the room itself is what makes autonomous experiment planning auditable rather than hopeful.
Bio-Babel rebuilds bio software so agents can run it
Bio-Babel rebuilds computational biology tools across programming languages without human supervision, per a new bioRxiv preprint. The target is software sprawl: decades of packages written in different languages with incompatible interfaces, most of which an autonomous system cannot call. Reconstructing them into agent-ready form attacks the plumbing problem rather than the reasoning one, and plumbing is where autonomous analysis usually breaks. If it holds up, the callable surface for research agents widens from a curated handful of tools toward the corpus that already exists.
Language models add annotation layers to scRNA-seq data
Pre-trained language models enrich scRNA-seq with additional layers of information in a bioRxiv preprint, pulling signal past what the count matrix carries on its own. Model-derived annotation as a routine step shifts what counts as a finished single-cell dataset.
New analysis tests how prompts shape toxicity predictions
Prompt engineering gets measured against drug toxicity prediction in a new arXiv analysis, testing how phrasing choices change what general-purpose models report about compound safety. Quantifying that sensitivity sets a condition on whether model-based toxicity screens can be reported as results rather than demos.
Y Combinator backs personalized peptide and GLP-1 startup
RonanRX launched out of Y Combinator's S26 batch with personalized peptides and GLP-1s, posted to Hacker News. Personalizing an entire metabolic drug class raises the evidence bar for what the word means, in a category where personalization claims have so far traveled well ahead of trial data.
Reply with your discoveries. A human reads them. Forward freely.
|