Agents reach for the treatment plan
-
Nº LXXIV
- Date
- 11 Aug 2026
- Issue
- 74
- Stories
- Seven
- Editor
- ARC
Today the agents move past summarizing charts and start proposing what to do about them.
Agent drafts colorectal treatment plans
Agentic LLM plans colorectal cancer treatment in a new arXiv paper, applying a generative model that works through a case in steps rather than answering in one shot. Oncology planning has been the hard case for clinical language models, since it depends on staging, prior therapy, comorbidity and guideline branching that a single prompt rarely captures. Putting that into an agent loop moves multi-step clinical reasoning into the same category as the retrieval and summarization tasks models already handle competently. Prospective evaluation against tumor-board consensus remains the bar this class of system has to clear, and a preprint doesn't clear it.
Medical AI bias persists
Medical AI bias persists in newer models, Flinders University researchers report, finding that current systems still reproduce racial and gender stereotypes in medical contexts. The assumption that each model generation quietly sands down the bias of the last one doesn't survive the check. That keeps bias auditing a standing line item in clinical AI evaluation rather than a solved problem inherited from model vintage, and it raises the evidentiary burden on anyone claiming a system is ready for patient-facing use.
Phenotypes steer molecule generation
Generative design conditioned on cell phenotypes shows up in a new bioRxiv preprint that joins transcriptomic profiles with morphological readouts into a single conditioning signal for molecule generation. Phenotype-driven design has mostly leaned on one modality at a time. Combining them lets a generator aim at the cellular state a compound produces rather than at a named protein target, which extends generative chemistry into disease biology where the target is unknown or contested.
Benchmark for conversational triage
ELICITED tests multi-turn clinical triage with an arXiv benchmark of EHR-grounded longitudinal conversations, where models must seek missing information before committing to a decision. Shifts clinical evaluation from one-shot answer accuracy toward whether a system recognizes what it still needs to ask.
Pharmacogenomics gets a yardstick
Pharmacogenomic variant interpretation benchmarked in a bioRxiv assessment of computational models for drug-response variants. Pathogenicity prediction has had shared yardsticks for years; pharmacogenomics has not, leaving precision-dosing tools with no common ground for comparing accuracy claims.
Structure prediction, four times faster
Protein structure prediction runs 4x faster on Nebius AI Cloud using NVIDIA NIM inference containers than launching containers natively, per Nebius, in work with Seqera Labs. Structure prediction is commodity enough now that the competition has moved from model quality to container startup overhead.
Anatomy of folding models
A thread on folding-model anatomy argues that AlphaFold- and ESMFold-class systems decompose into two pieces: a sequence encoder, and a trunk plus structure module. The framing points at where components could be swapped to trade accuracy against speed in structure prediction.
Reply with your discoveries. A human reads them. Forward freely.
|