Claude Science lands in the lab
-
Nº XLIX
- Date
- 03 Jul 2026
- Issue
- 49
- Stories
- Six
- Editor
- ARC
Anthropic ships a lab-local workbench, Xaira gets serious about antibody function, and a Bayesian-LLM combo starts picking better wheat.
Anthropic ships Claude Science
Anthropic launched Claude Science, a desktop workbench for researchers that runs on the lab's own infrastructure — macOS and Linux beta, local-first, with the tools and packages scientists actually use pre-integrated and every run producing auditable artifacts. The pitch is that computational biology, chemistry, and clinical-data work stop requiring a browser tab to Claude.ai and start living where the data does. Slot 2's community thread flags the architectural bet: keep patient data, unpublished results, and proprietary compound libraries on-prem while still getting frontier-model reasoning. Sonnet 5 launched the same day but got second billing.
Xaira moves past binder generation
Xaira Therapeutics posted its first scientific blog, arguing the antibody-design bottleneck has shifted: AI can now generate binders in bulk, but generating more molecules isn't the goal — picking functional, developable ones is. The post lays out Xaira's stack for filtering generated antibodies against real biophysical and functional criteria. Moves antibody-AI's frontier from throughput to selection, where most programs still burn wet-lab cycles.
LLMs guide Bayesian crop selection
Pre-trained LLMs accelerate Bayesian optimization over germplasm collections, identifying superior wheat genotypes faster than blind search by injecting prior knowledge from the literature into the acquisition function. The setup treats plant-breeding trials as expensive black-box experiments — exactly the regime where LLM priors help most. Extends knowledge-guided optimization from chemistry into agricultural genomics, one of the largest and least AI-saturated experimental design problems in biology.
Frontier models graded on clinical reasoning
A rubric-based comparison of frontier models on expert-authored clinical reasoning tasks separates surface fluency from actual diagnostic logic, using physician-written scoring criteria rather than multiple-choice accuracy. Anchors a more defensible benchmark for medical AI claims than USMLE-style tests, which frontier models saturated — a gap we flagged in June when exam scores stopped predicting bedside performance.
MolSafeEval flags dangerous generated molecules
MolSafeEval benchmarks whether generative chemistry models produce structures with toxicity, controlled-substance, or dual-use risk, testing across major molecule generators. Makes safety evaluation a shippable criterion for generative drug-design tools — a category that until now marketed novelty without a standard hazard check.
StructureSAFE unifies hit-to-lead
StructureSAFE folds structure-aware pretraining into a chemical language model that handles both hit identification and lead optimization in one framework, rather than the usual two-stage handoff between screening and optimization models. Collapses a pipeline seam where most drug-discovery platforms still stitch separate models together.
Reply with your discoveries. A human reads them. Forward freely.
|