# DrugTargetWorld > An open benchmark for training and evaluating AI agents on end-to-end, biobank-based drug target discovery. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. Samuel Margolis, Paul Schmiedmayer, Alan Huang, Ethan Chen, Ishan Bhattacharjee, Atman Shah, Ben Viggiano, Fang Cao, Shriya Reddy, Roger Xia, Jack O'Sullivan, Daniel Katz, Matthew Wheeler, Euan Ashley, Bruna Gomes. arXiv:2610.09558 (2026). DOI 10.48550/arXiv.2610.09558. ## Key facts - DrugTargetWorld is a benchmark for training and evaluating AI scientists on end-to-end drug target discovery. - It procedurally generates synthetic biobanks whose causal ground truth is known to the evaluator but hidden from the agent. - Each world contains genotypes, proteomics, transcriptomics, metabolomics, health records, outcomes and synthetic cardiac MRI for 54,000 participants. - Agents construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and can run virtual knockdown experiments under a budget. - Initial evaluation: 9 agents, 540 episodes, 20 cardiovascular worlds, 3 experimental budgets, rubric v0.9. - Opus 5 and GPT-5.6 Sol scored highest, 39.98 and 35.38 of 100; both recovered 64% of causal drivers on average, and no agent reliably distinguished misleading non-causal proteins. ## Methodological contributions DrugTargetWorld uses procedurally generated scientific worlds defined by hidden structural causal models, realistic causal traps and budgeted virtual experiments to formulate open-ended scientific discovery as a sequential decision problem with verifiable reward. - Procedural scientific worlds: a known but concealed causal truth allows verifiable, contamination-resistant evaluation and the generation of new training and test environments (https://drugtargetworld.vercel.app/scientific-worlds). - Open-strategy discovery: agents choose the phenotype, analyses, experiments and stopping point rather than execute a prescribed workflow (https://drugtargetworld.vercel.app/open-ended-science). - Causal traps: confounding, reverse causation, selection (collider) bias, non-identifiability, pleiotropy, surrogate-outcome discordance and measurement artifacts are built into each world (https://drugtargetworld.vercel.app/causal-traps). - Budgeted experimentation: agents allocate a finite experimental budget across virtual knockdowns and cell perturbations according to accumulated evidence (https://drugtargetworld.vercel.app/experimentation). - Multimodal synthetic biobanks: each world jointly generates genetics, molecular measurements, clinical records, longitudinal outcomes and raw imaging (https://drugtargetworld.vercel.app/synthetic-biobank). Findings: https://drugtargetworld.vercel.app/findings ## When to use DrugTargetWorld Use it when the job is one of these: - Evaluating an AI agent or AI scientist on end-to-end drug target discovery: building a disease phenotype from raw data, separating causal proteins from confounded ones, choosing a therapeutic direction, and spending an experimental budget. - Comparing agents or models on causal inference from multimodal biobank data (genotypes, proteomics, health records, cardiac MRI) where the right answer is known. - Training an agent with verifiable reward: each world's causal structure is known to the scorer, so a score is a check against truth, not a judgment. - Answering questions about the DrugTargetWorld paper (arXiv:2610.09558), its leaderboard, or its methods. Do not use it for real therapeutic targets: every world is synthetic, and no result here is evidence about a real protein or disease. ## How an agent runs it DrugTargetWorld runs in Harbor. The world data is gated behind the ACDC terms on Hugging Face, so a human must accept them once and provide a token. ```bash uv tool install harbor export HF_TOKEN=hf_... # from an account that accepted the terms at https://huggingface.co/datasets/sammargolis/drugtargetworld-assets harbor run -d drugtargetbench/drugtargetbench@v1.1 -a -m ``` Each of the 60 tasks pairs one of 20 worlds with one of 3 experimental budgets. The agent writes `/app/results/submission.json` (ranked causal drivers with direction, optional rejected decoys and abstentions) and `/app/results/phenotype.csv`; a separate scorer returns rubric v0.9 on 0 to 100. Full steps: https://drugtargetworld.vercel.app/run Every page on this site is also available as Markdown: request it with `Accept: text/markdown`. ## Pages - [Home and leaderboard](https://drugtargetworld.vercel.app/): scores for every agent, disease states, method - [Paper, full text](https://drugtargetworld.vercel.app/paper): the main text as HTML, with figures and Table 1 - [References](https://drugtargetworld.vercel.app/paper/references): the paper's reference list - [Supplementary information](https://drugtargetworld.vercel.app/paper/supplementary): supplementary figures S1-S2, tables S1-S11, methods and the verbatim agent prompt - [Supplementary notes](https://drugtargetworld.vercel.app/paper/supplementary/notes): turn-by-turn traces of two human-guided episodes and additional results - [What Frontier AI Agents Can and Cannot Do in End-to-End Drug Target Discovery](https://drugtargetworld.vercel.app/findings) - [Procedurally Generated Scientific Worlds for Training AI Scientists](https://drugtargetworld.vercel.app/scientific-worlds) - [Causal Traps for Evaluating AI Scientists](https://drugtargetworld.vercel.app/causal-traps) - [Benchmarking End-to-End Scientific Discovery](https://drugtargetworld.vercel.app/open-ended-science) - [A Multimodal Synthetic Biobank with Known Causal Ground Truth](https://drugtargetworld.vercel.app/synthetic-biobank) - [Budgeted Experimentation and Verifiable Reward for AI Scientists](https://drugtargetworld.vercel.app/experimentation) - [The End-to-End Drug Target Discovery Task](https://drugtargetworld.vercel.app/drug-target-discovery) - [Run](https://drugtargetworld.vercel.app/run): running the benchmark in Harbor - [About](https://drugtargetworld.vercel.app/about), [Contact](https://drugtargetworld.vercel.app/contact), [Privacy](https://drugtargetworld.vercel.app/privacy) ## Links - [arXiv](https://arxiv.org/abs/2610.09558) - [PDF](https://drugtargetworld.vercel.app/paper/DrugTargetWorld_v1.pdf) - [Code](https://github.com/sammargolis/DrugTargetWorld/) - [Dataset](https://huggingface.co/datasets/sammargolis/drugtargetworld-assets) ## Citation ```bibtex @article{margolis2026drugtargetworld, title = {{DrugTargetWorld}: A Synthetic Biobank for Training and Benchmarking {AI} Scientists}, author = {Margolis, Samuel and Schmiedmayer, Paul and Huang, Alan and Chen, Ethan and Bhattacharjee, Ishan and Shah, Atman and Viggiano, Ben and Cao, Fang and Reddy, Shriya and Xia, Roger and O'Sullivan, Jack and Katz, Daniel and Wheeler, Matthew and Ashley, Euan and Gomes, Bruna}, journal = {arXiv preprint arXiv:2610.09558}, year = {2026}, doi = {10.48550/arXiv.2610.09558} } ```