6 min read

GPT-5.6 rewrites the agent scoreboard

GPT-5.6 rewrites the agent scoreboard
Nº 01 · The Lede X Agents · Infrastructure

GPT-5.6 Sol tops Agents' Last Exam

GPT-5.6 Sol tops Agents' Last Exam
Fig. IX · Filed 15 Jul 2026.

GPT-5.6 Sol posts 53.6 on Agents' Last Exam, a benchmark measuring long-horizon tool use, beating Claude Fable 5 by 13.1 points and doing it at roughly one-quarter the compute cost at medium reasoning. On the Artificial Analysis Coding Agent Index — a separate industry scoreboard — the same model hits 80.0, 2.8 points above Fable 5, while using under half the output tokens and half the wall time. Two independent benchmarks moving together resets the reference frontier for agent capability and, more consequentially, the cost curve underneath it.

Read the source

Multi-agent tumor boards for breast cancer
Fig. IIarXiv · Filed 15 Jul 2026.
Nº 02 arXiv Agents · Infrastructure

Multi-agent tumor boards for breast cancer

Agentic tumor boards drafted in a new arXiv paper, with specialist agents deliberating over breast cancer treatment plans and reconciling guideline conflicts before returning a recommendation. Moves oncology decision-support from single-model Q&A toward structured multi-agent deliberation — the shape clinical adoption debates will now argue over, since a deliberating panel is auditable in a way a single LLM answer isn't.

Read more
Symptom detection without fine-tuning
Fig. IIIarXiv · Filed 15 Jul 2026.
Nº 03 arXiv Field report

Symptom detection without fine-tuning

A multi-agent system extracts clinical symptoms from notes without any task-specific fine-tuning, per a new validation study. Collapses one of the standing cost items in clinical NLP: teams needed labeled corpora and a training run for every new symptom ontology. If the accuracy holds under external validation, fine-tuning-free becomes the default expectation for clinical extraction.

Read more
Also Filed · Four Briefs from the queue
Nº 04 bioRxiv Agents · Infrastructure

Transplant-Agents audits rejection biomarkers

Transplant-Agents reruns published post-transplant risk-prediction and rejection-biomarker studies to check whether their results reproduce, per a bioRxiv preprint. Establishes automated reproducibility assessment as a viable capability, not just a manifesto — a new pressure point on how transplant biomarker papers get evaluated after publication.

Read
Nº 05 bioRxiv Field report

Frozen embeddings rank antibody affinity

Frozen protein foundation-model embeddings rank antibody-antigen binding-affinity changes competitively without any task-specific training, per a bioRxiv result. Narrows the case for expensive fine-tuning in antibody engineering — pretrained embeddings are becoming a stronger default baseline than earlier antibody affinity work here suggested.

Read
Nº 06 Hacker News Field report

Skillscript proposes a sandboxed tool language

Skillscript debuted on HN as a declarative, sandboxed language for orchestrating agent tool calls — an early attempt to make tool-use programs auditable and safely re-executable. Adds to the growing debate over whether agent tool orchestration should live in freeform code or in a constrained DSL, a question that will shape how biology agents get certified for regulated workflows.

Read
Nº 07 OpenAI Agents · Infrastructure

OpenAI's playbook for agentic AI budgets

OpenAI published guidance on managing enterprise AI spend in the agentic era, centered on measuring useful work per dollar rather than seat licenses. Signals the metric enterprise buyers — including biopharma IT — will be asked to defend budgets against next cycle.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 56  ·  15 Jul 2026

Editor's Note

New OpenAI numbers land while three clinical multi-agent papers quietly show what the field will actually do with them.

 

Nº 01 · The Lede  —  X  —  Agents · Infrastructure

GPT-5.6 Sol tops Agents' Last Exam

GPT-5.6 Sol tops Agents' Last Exam

Fig. I  X · Filed 15 Jul 2026.

GPT-5.6 Sol posts 53.6 on Agents' Last Exam, a benchmark measuring long-horizon tool use, beating Claude Fable 5 by 13.1 points and doing it at roughly one-quarter the compute cost at medium reasoning. On the Artificial Analysis Coding Agent Index — a separate industry scoreboard — the same model hits 80.0, 2.8 points above Fable 5, while using under half the output tokens and half the wall time. Two independent benchmarks moving together resets the reference frontier for agent capability and, more consequentially, the cost curve underneath it.

Read the source →

Why it matters

The cost-per-agent-task ceiling just dropped by roughly 4x at the top of the leaderboard — biology workflows that were priced out of full-agent execution six months ago now pencil, and every downstream vendor pitch resets against the new number.

 

Nº 02  —  arXiv  —  Agents · Infrastructure

Multi-agent tumor boards for breast cancer

Fig. II  arXiv · Filed 15 Jul 2026.

Multi-agent tumor boards for breast cancer

Agentic tumor boards drafted in a new arXiv paper, with specialist agents deliberating over breast cancer treatment plans and reconciling guideline conflicts before returning a recommendation. Moves oncology decision-support from single-model Q&A toward structured multi-agent deliberation — the shape clinical adoption debates will now argue over, since a deliberating panel is auditable in a way a single LLM answer isn't.

Read more →

 

Nº 03  —  arXiv  —  Field report

Symptom detection without fine-tuning

Fig. III  arXiv · Filed 15 Jul 2026.

Symptom detection without fine-tuning

A multi-agent system extracts clinical symptoms from notes without any task-specific fine-tuning, per a new validation study. Collapses one of the standing cost items in clinical NLP: teams needed labeled corpora and a training run for every new symptom ontology. If the accuracy holds under external validation, fine-tuning-free becomes the default expectation for clinical extraction.

Read more →

 

Also Filed  ·  Four Briefs from the queue

Nº 04  —  bioRxiv  —  Agents · Infrastructure

Transplant-Agents audits rejection biomarkers

Transplant-Agents reruns published post-transplant risk-prediction and rejection-biomarker studies to check whether their results reproduce, per a bioRxiv preprint. Establishes automated reproducibility assessment as a viable capability, not just a manifesto — a new pressure point on how transplant biomarker papers get evaluated after publication.

Read →

Nº 05  —  bioRxiv  —  Field report

Frozen embeddings rank antibody affinity

Frozen protein foundation-model embeddings rank antibody-antigen binding-affinity changes competitively without any task-specific training, per a bioRxiv result. Narrows the case for expensive fine-tuning in antibody engineering — pretrained embeddings are becoming a stronger default baseline than earlier antibody affinity work here suggested.

Read →

Nº 06  —  Hacker News  —  Field report

Skillscript proposes a sandboxed tool language

Skillscript debuted on HN as a declarative, sandboxed language for orchestrating agent tool calls — an early attempt to make tool-use programs auditable and safely re-executable. Adds to the growing debate over whether agent tool orchestration should live in freeform code or in a constrained DSL, a question that will shape how biology agents get certified for regulated workflows.

Read →

Nº 07  —  OpenAI  —  Agents · Infrastructure

OpenAI's playbook for agentic AI budgets

OpenAI published guidance on managing enterprise AI spend in the agentic era, centered on measuring useful work per dollar rather than seat licenses. Signals the metric enterprise buyers — including biopharma IT — will be asked to defend budgets against next cycle.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.