5 min read

Self-evolving agents top drug discovery

Self-evolving agents top drug discovery
Nº 01 · The Lede X Drug discovery · Computational

DrugSAGE tops drug-discovery agents

DrugSAGE tops drug-discovery agents
Fig. IX · Filed 29 Jun 2026.

DrugSAGE ranks first among nine state-of-the-art drug-discovery agents by accumulating and reusing experience across tasks, according to a new release from Wengong Jin's group. The self-evolving LLM agent (one that updates its own playbook from prior runs rather than relying on a frozen prompt) transfers learned strategies between assays, hit-finding, and property prediction. Cross-task transfer has been the missing piece in agent-based discovery — most systems forget everything between projects.

Read the source

Real science breaks AI frameworks
Fig. IIX · Filed 29 Jun 2026.
Nº 02 X Field report

Real science breaks AI frameworks

Advanced AI frameworks stumble when tested against published studies in uncertainty quantification, Therapeutic Data Commons ML, and agent-based modeling — not curated benchmarks. The bioRxiv evaluation finds the gap between leaderboard scores and reproducing actual papers remains wide. Anchors a counterargument to benchmark-driven capability claims and forces vendors to show published-study replication, not just TDC numbers.

Read more
Agents reviewing agents
Fig. IIIbioRxiv · Filed 29 Jun 2026.
Nº 03 bioRxiv Agents · Infrastructure

Agents reviewing agents

An AI agent spends three weeks analyzing protein data and flags a drug target; a second agent decides whether the finding is strong enough to act on. The X thread sparked 30 replies debating where human judgment enters the loop. Moves the agent-oversight debate from theoretical to operational — multi-agent review is becoming the default architecture before anyone has agreed on accountability.

Read more
Also Filed · Three Briefs from the queue
Nº 04 bioRxiv Cell biology · Funding

scBench-Long stresses single-cell agents

scBench-Long benchmarks long-horizon single-cell biology tasks with verifiable answers, exposing where agents drift over multi-step scRNA-seq workflows. Establishes a reference benchmark for sustained reasoning in cell biology, where most existing evals stop at single-turn questions.

Read
Nº 05 arXiv Field report

Variant calling goes client-server

Client-server interfaces enable agent-driven variant calling, per a new bioRxiv preprint pushing heavy compute off the agent and onto dedicated callers. Narrows the gap between conversational agents and production genomics infrastructure — variant calling stops being a wall the agent hits.

Read
Nº 06 arXiv Clinical AI · Evaluation

Synthetic longitudinal clinical notes

A new pipeline generates longitudinal synthetic clinical notes with LLMs, giving researchers patient trajectories without de-identification headaches. Lowers the data-access ceiling for clinical NLP work that has been gated on real-EHR agreements.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 45  ·  29 Jun 2026

Editor's Note

Monday opens with agents that learn from their own past runs, and a sobering reminder that benchmarks aren't real science.

 

Nº 01 · The Lede  —  X  —  Drug discovery · Computational

DrugSAGE tops drug-discovery agents

DrugSAGE tops drug-discovery agents

Fig. I  X · Filed 29 Jun 2026.

DrugSAGE ranks first among nine state-of-the-art drug-discovery agents by accumulating and reusing experience across tasks, according to a new release from Wengong Jin's group. The self-evolving LLM agent (one that updates its own playbook from prior runs rather than relying on a frozen prompt) transfers learned strategies between assays, hit-finding, and property prediction. Cross-task transfer has been the missing piece in agent-based discovery — most systems forget everything between projects.

Read the source →

Why it matters

Persistent experience across tasks resets what counts as competitive for a discovery agent — single-task SOTA stops being the right scoreboard, and any platform that can't carry learning between campaigns falls behind the new reference.

 

Nº 02  —  X  —  Field report

Real science breaks AI frameworks

Fig. II  X · Filed 29 Jun 2026.

Real science breaks AI frameworks

Advanced AI frameworks stumble when tested against published studies in uncertainty quantification, Therapeutic Data Commons ML, and agent-based modeling — not curated benchmarks. The bioRxiv evaluation finds the gap between leaderboard scores and reproducing actual papers remains wide. Anchors a counterargument to benchmark-driven capability claims and forces vendors to show published-study replication, not just TDC numbers.

Read more →

 

Nº 03  —  bioRxiv  —  Agents · Infrastructure

Agents reviewing agents

Fig. III  bioRxiv · Filed 29 Jun 2026.

Agents reviewing agents

An AI agent spends three weeks analyzing protein data and flags a drug target; a second agent decides whether the finding is strong enough to act on. The X thread sparked 30 replies debating where human judgment enters the loop. Moves the agent-oversight debate from theoretical to operational — multi-agent review is becoming the default architecture before anyone has agreed on accountability.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  bioRxiv  —  Cell biology · Funding

scBench-Long stresses single-cell agents

scBench-Long benchmarks long-horizon single-cell biology tasks with verifiable answers, exposing where agents drift over multi-step scRNA-seq workflows. Establishes a reference benchmark for sustained reasoning in cell biology, where most existing evals stop at single-turn questions.

Read →

Nº 05  —  arXiv  —  Field report

Variant calling goes client-server

Client-server interfaces enable agent-driven variant calling, per a new bioRxiv preprint pushing heavy compute off the agent and onto dedicated callers. Narrows the gap between conversational agents and production genomics infrastructure — variant calling stops being a wall the agent hits.

Read →

Nº 06  —  arXiv  —  Clinical AI · Evaluation

Synthetic longitudinal clinical notes

A new pipeline generates longitudinal synthetic clinical notes with LLMs, giving researchers patient trajectories without de-identification headaches. Lowers the data-access ceiling for clinical NLP work that has been gated on real-EHR agreements.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.