7 min read

AI's biology risk stops being hypothetical

AI's biology risk stops being hypothetical
Nº 01 · The Lede X Field report

Anthropic details how people tried to misuse Claude

Anthropic details how people tried to misuse Claude
Fig. IX · Filed 14 Sep 2026.

Anthropic published its most detailed threat intelligence report yet, cataloguing how people have tried to misuse Claude for cyberattacks, influence operations, surveillance, biology, and weapons development. Biology appears as a named misuse category, which is itself the news: a frontier model vendor is now publishing what it observes people attempting on the bio side rather than asserting that safeguards work. The disclosure lands amid wider alarm about AI-designed viruses and the thin machinery available to check any of it, with Axios reporting this week that there is still no reliable way to know how these systems will behave. Documented attempts give the biosecurity guardrail debate something concrete to argue over instead of scenarios.

Read the source

Reinforcement learning optimizes a protein for brightness
Fig. IIX · Filed 14 Sep 2026.
Nº 02 X Structural biology · Protein design

Reinforcement learning optimizes a protein for brightness

Reinforcement learning pushes a protein past what evolution ever optimized for. CreiLOV is a light-sensing protein, and brightness was never the trait selection scored it on, so a protein language model trained on natural sequences has no reason to favor bright variants. Adding reinforcement learning (training against a reward for a chosen property, instead of copying existing examples) reorients the model toward the trait a designer wants, the thread argues. That moves protein language models from imitating evolutionary statistics to optimizing objectives nature never ran, which is where most therapeutic and reagent design actually lives.

Read more
GEM-GPT designs therapies cell type by cell type
Fig. IIIbioRxiv · Filed 14 Sep 2026.
Nº 03 bioRxiv Cell biology · Funding

GEM-GPT designs therapies cell type by cell type

GEM-GPT resolves therapeutic design down to individual cell types, a bioRxiv preprint reports, pointing generative drug design at systems pharmacology rather than single targets. Interventions are proposed against a patient-level picture of which cell types are doing what. Single-cell atlases made that resolution visible years ago, while drug design has mostly kept predicting at tissue average. Closing that gap on the design side is the claim worth watching here.

Read more
Also Filed · Three Briefs from the queue
Nº 04 bioRxiv Field report

ShEPhERD-2 designs molecules from interaction patterns

ShEPhERD-2 generates molecules from interaction profiles, the pattern of contacts a ligand makes with its target, rather than from scaffolds or fingerprints. Using contacts as a shared representation lets generative design carry across targets instead of retraining for each one.

Read
Nº 05 arXiv Drug discovery · Computational

scDEFT predicts drug effects and tests counterfactuals

scDEFT predicts drug effects and answers counterfactuals: what a cell would have done under a different treatment. Adding that second question moves single-cell drug modeling from response prediction toward the causal comparisons trial design actually rests on.

Read
Nº 06 arXiv Field report

General AI forecasters get tested on glucose data

Time-series foundation models get tested on continuous glucose monitoring, with dietary logs added as context. Benchmarking general pretrained forecasters (trained broadly, not built for this task) against a physiological signal sets a reference point for whether generic models earn a place in metabolic monitoring.

Read

Reply with your discoveries. A human reads them. Forward freely.

Agentic Discovery  ·  Nº 98  ·  14 Sep 2026

Editor's Note

Monday: a threat report, a protein tuned past evolution, and three models arguing about what drugs do to cells.

 

Nº 01 · The Lede  —  X  —  Field report

Anthropic details how people tried to misuse Claude

Anthropic details how people tried to misuse Claude

Fig. I  X · Filed 14 Sep 2026.

Anthropic published its most detailed threat intelligence report yet, cataloguing how people have tried to misuse Claude for cyberattacks, influence operations, surveillance, biology, and weapons development. Biology appears as a named misuse category, which is itself the news: a frontier model vendor is now publishing what it observes people attempting on the bio side rather than asserting that safeguards work. The disclosure lands amid wider alarm about AI-designed viruses and the thin machinery available to check any of it, with Axios reporting this week that there is still no reliable way to know how these systems will behave. Documented attempts give the biosecurity guardrail debate something concrete to argue over instead of scenarios.

Read the source →

Why it matters

Misuse telemetry from a frontier vendor gives the biosecurity guardrail debate its first shared evidence base, and raises the disclosure bar every other model provider will be measured against.

The Bench NoteFrom Heureka Labs

A biosafety committee reviewing AI use in a lab can now point to observed behaviour rather than speculation.

What a project lock blocks. Bench's Privacy Mode switches off web access, literature and database lookups, compound lookups, cloud jobs and sync for a single project, enforced at the tool level rather than promised in a settings page.
The paperwork it leaves. Each run appends a plain summary of network activity — counts and reasons, nothing sensitive — to a log inside the project, which is the sort of thing an IRB or an industrial partner asks to see.
Overnight work counts. Scheduled and background tasks in a private project follow the same rules, so the boundary holds while nobody is watching.

What we’re watching: whether IRBs and biosafety reviews start asking for a log of that kind as standard paperwork

 

Nº 02  —  X  —  Structural biology · Protein design

Reinforcement learning optimizes a protein for brightness

Fig. II  X · Filed 14 Sep 2026.

Reinforcement learning optimizes a protein for brightness

Reinforcement learning pushes a protein past what evolution ever optimized for. CreiLOV is a light-sensing protein, and brightness was never the trait selection scored it on, so a protein language model trained on natural sequences has no reason to favor bright variants. Adding reinforcement learning (training against a reward for a chosen property, instead of copying existing examples) reorients the model toward the trait a designer wants, the thread argues. That moves protein language models from imitating evolutionary statistics to optimizing objectives nature never ran, which is where most therapeutic and reagent design actually lives.

Read more →

 

Nº 03  —  bioRxiv  —  Cell biology · Funding

GEM-GPT designs therapies cell type by cell type

Fig. III  bioRxiv · Filed 14 Sep 2026.

GEM-GPT designs therapies cell type by cell type

GEM-GPT resolves therapeutic design down to individual cell types, a bioRxiv preprint reports, pointing generative drug design at systems pharmacology rather than single targets. Interventions are proposed against a patient-level picture of which cell types are doing what. Single-cell atlases made that resolution visible years ago, while drug design has mostly kept predicting at tissue average. Closing that gap on the design side is the claim worth watching here.

Read more →

 

Also Filed  ·  Three Briefs from the queue

Nº 04  —  bioRxiv  —  Field report

ShEPhERD-2 designs molecules from interaction patterns

ShEPhERD-2 generates molecules from interaction profiles, the pattern of contacts a ligand makes with its target, rather than from scaffolds or fingerprints. Using contacts as a shared representation lets generative design carry across targets instead of retraining for each one.

Read →

Nº 05  —  arXiv  —  Drug discovery · Computational

scDEFT predicts drug effects and tests counterfactuals

scDEFT predicts drug effects and answers counterfactuals: what a cell would have done under a different treatment. Adding that second question moves single-cell drug modeling from response prediction toward the causal comparisons trial design actually rests on.

Read →

Nº 06  —  arXiv  —  Field report

General AI forecasters get tested on glucose data

Time-series foundation models get tested on continuous glucose monitoring, with dietary logs added as context. Benchmarking general pretrained forecasters (trained broadly, not built for this task) against a physiological signal sets a reference point for whether generic models earn a place in metabolic monitoring.

Read →

 

· · ·

Reply with your discoveries. A human reads them. Forward freely.