DrugTargetWorld
What Frontier AI Agents Can and Cannot Do in End-to-End Drug Target Discovery
Findings from evaluating nine AI agents in 540 episodes of DrugTargetWorld: agents recover many causal drivers but struggle with the integrative judgments of end-to-end scientific discovery.
01 Question
Frontier AI agents recover causal drivers but struggle with end-to-end scientific judgment
Nine language-model agents were evaluated in 540 episodes across 20 synthetic biobank worlds and three experimental budgets, each choosing its own research strategy with no prescribed pipeline. The two leading agents, Opus 5 and GPT-5.6 Sol, recovered 64% of causal drivers on average (target recall 0.64 for both), yet achieved mean overall scores of only 39.98 and 35.38 of 100.
Across all 540 episodes the mean total score was 14.1, and fewer than half of the episodes (258, 47.8%) scored above 0. Sonnet 5 and Haiku 4.5 averaged 21.33 and 12.92, and the five open-weight models 0.81 to 7.41, each with a median of 0. Only six episodes (1.1%) scored above 80; all six nominated exactly the two planted causal drivers with no false positives.
The agents could carry out the individual analyses of a biobank-based target study. What limited them were the judgments that connect those analyses: how to measure the disease, which evidence separates a causal driver from a non-causal protein, and when a claim is sufficiently supported.
02 Question
AI agents fail to reliably distinguish causal from misleading non-causal proteins
The leading agents based most final nominations on cis Mendelian randomization (57 of 60 Opus 5 episodes, 50 of 60 GPT-5.6 Sol) and checked instrument strength in nearly all episodes. Even so, no agent averaged more than 2.87 of 15 points for bias identification, which requires rejecting planted non-causal proteins with the correct source of bias.
The reverse-causation protein, the one rejected most often, was rejected without being nominated in 121 of 378 eligible episodes (32.0%), and only 54 (14.3%) also named the correct source of bias. The harmful surrogate, which improves imaging but increases mortality, was nominated in 44 of 60 Opus 5 and 25 of 60 GPT-5.6 Sol episodes. These agents usually flagged it as harmful; every surrogate nomination by Haiku 4.5 and GPT-OSS-20B claimed it was beneficial and incurred the safety penalty. See causal traps.
03 Question
Phenotype construction is a bottleneck in autonomous biobank research
Cardiac MRI was released only as raw images, so agents had to extract imaging features and build their own disease phenotype. The difference between the two leading agents lay mainly here: Opus 5 averaged 7.2 of 10 phenotype points and GPT-5.6 Sol 3.2, and Opus 5 earned positive phenotype credit in 56 of 60 episodes against 34 for GPT-5.6 Sol. Mean correlations with latent disease severity were 0.785, 0.637 and 0.502 for Opus 5, GPT-5.6 Sol and Sonnet 5.
Phenotype construction received zero credit in 416 of 540 episodes (77.0%). Phenotypes checked against clinical endpoints before submission correlated more strongly with disease (mean 0.663 against 0.402). Without phenotype credit, the gap between the two leading agents fell from 4.6 to 0.9 points and the ranking of all nine agents was unchanged. See the synthetic biobank.
04 Question
Additional experimental budget does not necessarily improve scientific-agent performance
Pooled mean scores were 13.11, 12.77 and 16.28 under the observational-only, limited ($450,000) and expanded ($2 million) budgets. For the leading agents the expanded budget changed little: +0.13 points for Opus 5 and +0.07 for GPT-5.6 Sol, even though they bought experiments in 19 of 20 expanded-budget episodes each and spent 93% and 95% of the funds. Their recall stayed at 0.65 and 0.60 against 0.64 without experiments.
Where two candidates shared a genetic instrument, the leading agents tested the causal member in 14 of 24 eligible episodes, but no experiment targeted any of the 15 delayed-effect pairs. See budgeted experimentation.
05 Question
Correct answers did not always reflect supporting analysis
Three selected episodes scored 75 and recovered the exact planted set of causal drivers. One cited knockdown responses as evidence; the other two reached the same set through correlation-based screening and hard-coded outcome-alignment labels, yet all three received full causal-confidence and direction credit because the score evaluates final claims. A training reward built on this benchmark should therefore be paired with trajectory annotation and exploit audits.

In the paper
- 4.1 Leading agents recovered many causal drivers but remained unreliable
- 4.2 Phenotype construction separated the two leading agents
- 4.3 No agent reliably rejected non-causal proteins
- 4.4 Experiments and the expanded budget
- 4.5 Failures across the workflow
- Leaderboard with all nine agents
Related
- Procedurally Generated Scientific Worlds for Training AI Scientists
- Causal Traps for Evaluating AI Scientists
- Benchmarking End-to-End Scientific Discovery
- A Multimodal Synthetic Biobank with Known Causal Ground Truth
- Budgeted Experimentation and Verifiable Reward for AI Scientists
- The End-to-End Drug Target Discovery Task