DrugTargetWorld
Budgeted Experimentation and Verifiable Reward for AI Scientists
In DrugTargetWorld an AI scientist must decide which experiments are worth their cost. Virtual knockdowns and cell perturbations re-run the hidden causal model, so every experiment returns a true counterfactual and every claim can be rewarded against it.
01 Question
Why must AI scientists decide which experiments are worth conducting?
Laboratory experiments are expensive, so a useful scientist must judge when an experiment is worth its cost and which uncertainty it should resolve. Some questions cannot be settled from observational data at all: when two proteins share a genetic instrument, or when a causal effect appears only over years of follow-up, genetics alone cannot say which protein is causal.
DrugTargetWorld evaluates each agent under three budgets: observational-only ($0), limited ($450,000), enough for one knockdown or three cell perturbations, and expanded ($2 million), enough for up to five knockdowns. The amounts are an ordinal scale of experimental access, not estimates of real costs. Analysis of the released data is always free.
02 Question
What virtual experiments can an agent run?
- Cell perturbation ($150,000): a smaller virtual experiment with measurement noise that measures the effect at the first visit.
- In vivo knockdown ($400,000): a larger virtual experiment that returns the effect over five years of simulated follow-up, without additional noise.
Either applies to any protein and returns the change in cavity radius, wall thickness, ejection fraction and five-year survival relative to the unperturbed state. Because the generator is the ground truth, perturbing a protein re-runs the downstream structural equations: the result is a counterfactual, not a lookup, and a protein that is not a driver returns no effect. The cheaper experiment cannot detect a delayed effect, so choosing between them is part of the task.
03 Question
How does budgeted experimentation provide a verifiable reward?
The agent's work is a sequence of decisions: at ~ π(a | st, Bt), with the budget falling by each experiment's cost, Bt+1 = Bt − c(at). The causal-confidence score (25 points) ties claims to evidence on pairs that observational data cannot separate: naming the causal protein after testing it experimentally earns 25 points, acknowledging that the pair is unresolved earns up to 17.5, and an unsupported choice earns nothing. Without a budget, 17.5 is therefore the maximum per pair.
Because each world's causal structure is known to the scorer, this reward can train agents across the whole workflow, from phenotype to experiment to final claim. See scientific worlds.
04 Question
Do AI agents allocate experimental resources well?
Of the 360 episodes with a budget, 143 received at least one experiment, and 413 experiments were delivered, 76.5% of them by Opus 5, GPT-5.6 Sol and Sonnet 5. Of those experiments, 171 (41.4%) targeted a causal driver, 52 (12.6%) the harmful surrogate and 49 (11.9%) the reverse-causation protein. Where a pair shared a genetic instrument, the leading agents tested the causal member in 14 of 24 eligible episodes, and every driver they tested and nominated carried the direction its knockdown implied. No experiment targeted any of the 15 delayed-effect pairs.
Agents usually read returned results (124 of 143 episodes) and 84 changed their conclusion; 16 of the 19 that did not had submitted in the same turn as their request. Yet more budget changed little for the leading agents: +0.13 points for Opus 5 and +0.07 for GPT-5.6 Sol under the expanded budget, despite spending 93% and 95% of it. See findings.

In the paper
- 3.4 Action space and experiments
- 3.5 Scoring: causal confidence
- 4.4 Experiments and the expanded budget
- Supplementary Methods S1.5: intervention model
Related
- What Frontier AI Agents Can and Cannot Do in End-to-End Drug Target Discovery
- Procedurally Generated Scientific Worlds for Training AI Scientists
- Causal Traps for Evaluating AI Scientists
- Benchmarking End-to-End Scientific Discovery
- A Multimodal Synthetic Biobank with Known Causal Ground Truth
- The End-to-End Drug Target Discovery Task