DrugTargetWorld
Benchmarking End-to-End Scientific Discovery
Most scientific-agent benchmarks specify the question and the analysis. DrugTargetWorld leaves the whole research strategy to the AI scientist and scores the result against causal truth.
01 Question
Why evaluate complete scientific workflows rather than isolated tasks?
Recent benchmarks have moved from testing static scientific knowledge toward testing whether agents can execute research workflows. In most of them, though, the research question and often the analysis methods are specified. In real drug target discovery the hard decisions come earlier and between steps: how to measure the disease, which analyses to run, which candidates deserve an expensive experiment, and when the evidence is enough.
Evaluating the whole workflow matters particularly here because experimental validation is costly in both money and time, so an agent must prioritise which candidates merit testing.
02 Question
What makes DrugTargetWorld an open-strategy benchmark?
The agent receives one world's released data and no analysis plan. It chooses the phenotype, the hypotheses, the analyses, the experiments and the stopping point. The harness is deliberately minimal: it runs the Python programs the agent writes, returns their output, and gives no feedback on scientific correctness, so differences in score reflect the agents' own research strategies.
| Benchmark | Task | Method choice | Sequential | Simulated | Ground truth | Biobank | Trainable |
|---|---|---|---|---|---|---|---|
| ScienceAgentBench | Scientific analysis | Defined | Limited | No | No | No | No |
| DiscoveryBench | Scientific discovery | Defined | Limited | Variable | Variable | No | No |
| BixBench | Bioinformatics | Defined | Yes | No | No | No | No |
| BixBench3 | Full study execution | Guided | Yes | No | No | No | No |
| GeneBench-Pro | Genetic analysis | Open analysis | Yes | Yes | Yes | No | No |
| Aviary | Scientific tasks | Task specific | Yes | Variable | Variable | No | Yes |
| ResidencyRL | Clinical care | Scenario defined | Yes | Yes | Scenario | No | Yes |
| Causal benchmarks | Causal discovery | Variable | Variable | Yes | Yes | No | Variable |
| DrugTargetWorld | Drug target discovery | Open strategy | Yes | Yes | Yes | Yes | Yes |
From Table 1 of the paper. "Open strategy" means the agent chooses the phenotype, hypotheses, analyses, experiments and stopping point; "Trainable" means the environment can supply a verifiable reward for training.
03 Question
How is scientific discovery framed as sequential decision-making?
An episode is a sequence of up to 30 turns. At each turn the agent chooses its next analysis or experiment from the evidence gathered so far and the budget that remains, at ~ π(a | st, Bt), and experiments reduce the budget, Bt+1 = Bt − c(at), while analyses of the released data are free. The same formulation defines the policy that training in this environment would optimise. See budgeted experimentation.
04 Question
Where do autonomous AI scientists break down in open-ended research?
Failures arose at several steps. Phenotype construction earned zero credit in 77.0% of episodes. Agents generated evidence without using it: Opus 5 and GPT-5.6 Sol read genotypes in all 60 episodes, but instrumental-variable analyses appeared in only 41 and 29. Code execution failed in smaller models, and Devstral Small hit the 30-turn limit in 53 of 60 episodes by repeatedly re-deriving information.
The number of core workflow milestones an episode reached correlated with its score (Spearman ρ = 0.649), a descriptive relationship confounded by model identity. And correct answers did not always reflect supporting analysis. See findings.

In the paper
- 2. Related work: scientific agents, environments and causal benchmarks
- Table 1: comparison with other benchmarks
- 3.3 Agent environment and harness
- 4.5 Failures across the workflow
Related
- What Frontier AI Agents Can and Cannot Do in End-to-End Drug Target Discovery
- Procedurally Generated Scientific Worlds for Training AI Scientists
- Causal Traps for Evaluating AI Scientists
- A Multimodal Synthetic Biobank with Known Causal Ground Truth
- Budgeted Experimentation and Verifiable Reward for AI Scientists
- The End-to-End Drug Target Discovery Task