5 min read

Claude Science lands in the lab

Claude Science lands in the lab
Nº 01 · The Lede Anthropic Field report

Anthropic ships Claude Science

Anthropic ships Claude Science
Fig. IAnthropic · Filed 03 Jul 2026.

Anthropic launched Claude Science, a desktop workbench for researchers that runs on the lab's own infrastructure — macOS and Linux beta, local-first, with the tools and packages scientists actually use pre-integrated and every run producing auditable artifacts. The pitch is that computational biology, chemistry, and clinical-data work stop requiring a browser tab to Claude.ai and start living where the data does. Slot 2's community thread flags the architectural bet: keep patient data, unpublished results, and proprietary compound libraries on-prem while still getting frontier-model reasoning. Sonnet 5 launched the same day but got second billing.

Read the source

Xaira moves past binder generation
Fig. IIX · Filed 03 Jul 2026.
Nº 02 X Field report

Xaira moves past binder generation

Xaira Therapeutics posted its first scientific blog, arguing the antibody-design bottleneck has shifted: AI can now generate binders in bulk, but generating more molecules isn't the goal — picking functional, developable ones is. The post lays out Xaira's stack for filtering generated antibodies against real biophysical and functional criteria. Moves antibody-AI's frontier from throughput to selection, where most programs still burn wet-lab cycles.

Read more
LLMs guide Bayesian crop selection
Fig. IIIbioRxiv · Filed 03 Jul 2026.
Nº 03 bioRxiv Field report

LLMs guide Bayesian crop selection

Pre-trained LLMs accelerate Bayesian optimization over germplasm collections, identifying superior wheat genotypes faster than blind search by injecting prior knowledge from the literature into the acquisition function. The setup treats plant-breeding trials as expensive black-box experiments — exactly the regime where LLM priors help most. Extends knowledge-guided optimization from chemistry into agricultural genomics, one of the largest and least AI-saturated experimental design problems in biology.

Read more
Also Filed · Three Briefs from the queue
Nº 04 arXiv Clinical AI · Evaluation

Frontier models graded on clinical reasoning

A rubric-based comparison of frontier models on expert-authored clinical reasoning tasks separates surface fluency from actual diagnostic logic, using physician-written scoring criteria rather than multiple-choice accuracy. Anchors a more defensible benchmark for medical AI claims than USMLE-style tests, which frontier models saturated — a gap we flagged in June when exam scores stopped predicting bedside performance.

Read
Nº 05 arXiv Field report

MolSafeEval flags dangerous generated molecules

MolSafeEval benchmarks whether generative chemistry models produce structures with toxicity, controlled-substance, or dual-use risk, testing across major molecule generators. Makes safety evaluation a shippable criterion for generative drug-design tools — a category that until now marketed novelty without a standard hazard check.

Read
Nº 06 bioRxiv Field report

StructureSAFE unifies hit-to-lead

StructureSAFE folds structure-aware pretraining into a chemical language model that handles both hit identification and lead optimization in one framework, rather than the usual two-stage handoff between screening and optimization models. Collapses a pipeline seam where most drug-discovery platforms still stitch separate models together.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 49  ·  03 Jul 2026

Editor's Note

Anthropic ships a lab-local workbench, Xaira gets serious about antibody function, and a Bayesian-LLM combo starts picking better wheat.

 

Nº 01 · The Lede  —  Anthropic  —  Field report

Anthropic ships Claude Science

Anthropic ships Claude Science

Fig. I  Anthropic · Filed 03 Jul 2026.

Anthropic launched Claude Science, a desktop workbench for researchers that runs on the lab's own infrastructure — macOS and Linux beta, local-first, with the tools and packages scientists actually use pre-integrated and every run producing auditable artifacts. The pitch is that computational biology, chemistry, and clinical-data work stop requiring a browser tab to Claude.ai and start living where the data does. Slot 2's community thread flags the architectural bet: keep patient data, unpublished results, and proprietary compound libraries on-prem while still getting frontier-model reasoning. Sonnet 5 launched the same day but got second billing.

Read the source →

Why it matters

Local-first frontier AI is now a shipped product, not a whitepaper — resets what "deployable in a regulated lab" means, and forces every rival to answer whether their workbench can run behind a hospital firewall without phoning home.

 

Nº 02  —  X  —  Field report

Xaira moves past binder generation

Fig. II  X · Filed 03 Jul 2026.

Xaira moves past binder generation

Xaira Therapeutics posted its first scientific blog, arguing the antibody-design bottleneck has shifted: AI can now generate binders in bulk, but generating more molecules isn't the goal — picking functional, developable ones is. The post lays out Xaira's stack for filtering generated antibodies against real biophysical and functional criteria. Moves antibody-AI's frontier from throughput to selection, where most programs still burn wet-lab cycles.

Read more →

 

Nº 03  —  bioRxiv  —  Field report

LLMs guide Bayesian crop selection

Fig. III  bioRxiv · Filed 03 Jul 2026.

LLMs guide Bayesian crop selection

Pre-trained LLMs accelerate Bayesian optimization over germplasm collections, identifying superior wheat genotypes faster than blind search by injecting prior knowledge from the literature into the acquisition function. The setup treats plant-breeding trials as expensive black-box experiments — exactly the regime where LLM priors help most. Extends knowledge-guided optimization from chemistry into agricultural genomics, one of the largest and least AI-saturated experimental design problems in biology.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  arXiv  —  Clinical AI · Evaluation

Frontier models graded on clinical reasoning

A rubric-based comparison of frontier models on expert-authored clinical reasoning tasks separates surface fluency from actual diagnostic logic, using physician-written scoring criteria rather than multiple-choice accuracy. Anchors a more defensible benchmark for medical AI claims than USMLE-style tests, which frontier models saturated — a gap we flagged in June when exam scores stopped predicting bedside performance.

Read →

Nº 05  —  arXiv  —  Field report

MolSafeEval flags dangerous generated molecules

MolSafeEval benchmarks whether generative chemistry models produce structures with toxicity, controlled-substance, or dual-use risk, testing across major molecule generators. Makes safety evaluation a shippable criterion for generative drug-design tools — a category that until now marketed novelty without a standard hazard check.

Read →

Nº 06  —  bioRxiv  —  Field report

StructureSAFE unifies hit-to-lead

StructureSAFE folds structure-aware pretraining into a chemical language model that handles both hit identification and lead optimization in one framework, rather than the usual two-stage handoff between screening and optimization models. Collapses a pipeline seam where most drug-discovery platforms still stitch separate models together.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.