GPT-5.6 rewrites the agent scoreboard
-
Nº LVI
- Date
- 15 Jul 2026
- Issue
- 56
- Stories
- Seven
- Editor
- ARC
New OpenAI numbers land while three clinical multi-agent papers quietly show what the field will actually do with them.
GPT-5.6 Sol tops Agents' Last Exam
GPT-5.6 Sol posts 53.6 on Agents' Last Exam, a benchmark measuring long-horizon tool use, beating Claude Fable 5 by 13.1 points and doing it at roughly one-quarter the compute cost at medium reasoning. On the Artificial Analysis Coding Agent Index — a separate industry scoreboard — the same model hits 80.0, 2.8 points above Fable 5, while using under half the output tokens and half the wall time. Two independent benchmarks moving together resets the reference frontier for agent capability and, more consequentially, the cost curve underneath it.
Multi-agent tumor boards for breast cancer
Agentic tumor boards drafted in a new arXiv paper, with specialist agents deliberating over breast cancer treatment plans and reconciling guideline conflicts before returning a recommendation. Moves oncology decision-support from single-model Q&A toward structured multi-agent deliberation — the shape clinical adoption debates will now argue over, since a deliberating panel is auditable in a way a single LLM answer isn't.
Symptom detection without fine-tuning
A multi-agent system extracts clinical symptoms from notes without any task-specific fine-tuning, per a new validation study. Collapses one of the standing cost items in clinical NLP: teams needed labeled corpora and a training run for every new symptom ontology. If the accuracy holds under external validation, fine-tuning-free becomes the default expectation for clinical extraction.
Transplant-Agents audits rejection biomarkers
Transplant-Agents reruns published post-transplant risk-prediction and rejection-biomarker studies to check whether their results reproduce, per a bioRxiv preprint. Establishes automated reproducibility assessment as a viable capability, not just a manifesto — a new pressure point on how transplant biomarker papers get evaluated after publication.
Frozen embeddings rank antibody affinity
Frozen protein foundation-model embeddings rank antibody-antigen binding-affinity changes competitively without any task-specific training, per a bioRxiv result. Narrows the case for expensive fine-tuning in antibody engineering — pretrained embeddings are becoming a stronger default baseline than earlier antibody affinity work here suggested.
Skillscript proposes a sandboxed tool language
Skillscript debuted on HN as a declarative, sandboxed language for orchestrating agent tool calls — an early attempt to make tool-use programs auditable and safely re-executable. Adds to the growing debate over whether agent tool orchestration should live in freeform code or in a constrained DSL, a question that will shape how biology agents get certified for regulated workflows.
OpenAI's playbook for agentic AI budgets
OpenAI published guidance on managing enterprise AI spend in the agentic era, centered on measuring useful work per dollar rather than seat licenses. Signals the metric enterprise buyers — including biopharma IT — will be asked to defend budgets against next cycle.
Reply with your discoveries. A human reads them. Forward freely.
|