Skip to content

DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

References

Abifadel, M., Varret, M., Rabès, J.-P., Allard, D., Ouguerram, K., Devillers, M., Cruaud, C., Benjannet, S., Wickham, L., Erlich, D., Derré, A., Villéger, L., Farnier, M., Beucler, I., Bruckert, E., Chambaz, J., Chanu, B., Lecerf, J.-M., Luc, G., … Boileau, C. (2003). Mutations in PCSK9 cause autosomal dominant hypercholesterolemia. Nature Genetics , 34 (2), 154–156. https://doi.org/10.1038/ng1161

An, U., Pazokitoroudi, A., Alvarez, M., Huang, L., Bacanu, S., Schork, A. J., Kendler, K., Pajukanta, P., Flint, J., Zaitlen, N., Cai, N., Dahl, A., & Sankararaman, S. (2023). Deep learning-based phenotype imputation on population-scale biobank data increases genetic discoveries. Nature Genetics , 55 (12), 2269–2276. https://doi.org/10.1038/s41588-023-01558-w

Aung, N., Lopes, L. R., Van Duijvenboden, S., Harper, A. R., Goel, A., Grace, C., Ho, C. Y., Weintraub, W. S., Kramer, C. M., Neubauer, S., Watkins, H. C., Petersen, S. E., & Munroe, P. B. (2023). Genome-Wide Analysis of Left Ventricular Maximum Wall Thickness in the UK Biobank Cohort Reveals a Shared Genetic Background With Hypertrophic Cardiomyopathy. Circulation: Genomic and Precision Medicine , 16 (1). https://doi.org/10.1161/CIRCGEN.122.003716

Aung, N., Vargas, J. D., Yang, C., Fung, K., Sanghvi, M. M., Piechnik, S. K., Neubauer, S., Manichaikul, A., Rotter, J. I., Taylor, K. D., Lima, J. A. C., Bluemke, D. A., Kawut, S. M., Petersen, S. E., & Munroe, P. B. (2022). Genome-wide association analysis reveals insights into the genetic architecture of right ventricular structure and function. Nature Genetics , 54 (6), 783–791. https://doi.org/10.1038/s41588-022-01083-2

Bycroft, C., Freeman, C., Petkova, D., Band, G., Elliott, L. T., Sharp, K., Motyer, A., Vukcevic, D., Delaneau, O., O’Connell, J., Cortes, A., Welsh, S., Young, A., Effingham, M., McVean, G., Leslie, S., Allen, N., Donnelly, P., & Marchini, J. (2018). The UK Biobank resource with deep phenotyping and genomic data. Nature , 562 (7726), 203–209. https://doi.org/10.1038/s41586-018-0579-z

Chen, Z. (2025). ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery . International Conference on Learning Representations. https://arxiv.org/abs/2410.05080

Chen, Z., Chen, Y., Liu, C., Yu, J., Song, X., Li, Z., Li, J., Torr, P., Han, B., & Zhang, K. (2026). CausalGame: Benchmarking Causal Thinking of LLM Agents in Games . International Conference on Machine Learning. https://doi.org/10.48550/arXiv.2607.04293

Cobbe, K., Hesse, C., Hilton, J., & Schulman, J. (2020). Leveraging Procedural Generation to Benchmark Reinforcement Learning . 119 , 2048–2056.

Cohen, J. C., Boerwinkle, E., Mosley, T. H., & Hobbs, H. H. (2006). Sequence Variations in PCSK9, Low LDL, and Protection against Coronary Heart Disease. New England Journal of Medicine , 354 (12), 1264–1272. https://doi.org/10.1056/NEJMoa054013

Duffy, Á., Petrazzini, B. O., Stein, D., Park, J. K., Forrest, I. S., Gibson, K., Vy, H. M., Chen, R., Márquez-Luna, C., Mort, M., Verbanck, M., Schlessinger, A., Itan, Y., Cooper, D. N., Rocheleau, G., Jordan, D. M., & Do, R. (2024). Development of a human genetics-guided priority score for 19,365 genes and 399 drug indications. Nature Genetics , 56 (1), 51–59. https://doi.org/10.1038/s41588-023-01609-2

Ference, B. A., Robinson, J. G., Brook, R. D., Catapano, A. L., Chapman, M. J., Neff, D. R., Voros, S., Giugliano, R. P., Davey Smith, G., Fazio, S., & Sabatine, M. S. (2016). Variation in PCSK9 and HMGCR and Risk of Cardiovascular Disease and Diabetes. New England Journal of Medicine , 375 (22), 2144–2153. https://doi.org/10.1056/NEJMoa1604304

Fortin, J.-P., Cullen, N., Sheline, Y. I., Taylor, W. D., Aselcioglu, I., Cook, P. A., Adams, P., Cooper, C., Fava, M., McGrath, P. J., McInnis, M., Phillips, M. L., Trivedi, M. H., Weissman, M. M., & Shinohara, R. T. (2018). Harmonization of cortical thickness measurements across scanners and sites. NeuroImage , 167 , 104–120. https://doi.org/10.1016/j.neuroimage.2017.11.024 Gomes, B., Singh, A., O’Sullivan, J. W., Schnurr, T. M., Goddard, P. C., Loong, S., Amar, D., Hughes, J. W., Kostur, M., Haddad, F., Salerno, M., Foo, R., Montgomery, S. B., Parikh, V. N., Meder, B., & Ashley, E. A. (2024). Genetic architecture of cardiac dynamic flow volumes. Nature Genetics , 56 (2), 245–257. https://doi.org/10.1038/s41588-023-01587-5

Gong, W., Bai, S., Zheng, Y.-Q., Smith, S. M., & Beckmann, C. F. (2023). Supervised Phenotype Discovery From Multimodal Brain Imaging. IEEE Transactions on Medical Imaging , 42 (3), 834–849. https://doi.org/10.1109/TMI.2022.3218720

Gong, W., Beckmann, C. F., & Smith, S. M. (2021). Phenotype discovery from population brain imaging. Medical Image Analysis , 71 , 102050. https://doi.org/10.1016/j.media.2021.102050

Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., … Zhang, Z. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature , 645 (8081), 633–638. https://doi.org/10.1038/s41586-025-09422-z

Heckbert, S. R., Post, W., Pearson, G. D. N., Arnett, D. K., Gomes, A. S., Jerosch-Herold, M., Hundley, W. G., Lima, J. A., & Bluemke, D. A. (2006). Traditional Cardiovascular Risk Factors in Relation to Left Ventricular Mass, Volume, and Systolic Function by Cardiac Magnetic Resonance Imaging. Journal of the American College of Cardiology , 48 (11), 2285–2292. https://doi.org/10.1016/j.jacc.2006.03.072

Henry, A., Gordillo-Marañón, M., Finan, C., Schmidt, A. F., Ferreira, J. P., Karra, R., Sundström, J., Lind, L., Ärnlöv, J., Zannad, F., Mälarstig, A., Hingorani, A. D., Lumbers, R. T., & HERMES and SCALLOP Consortia. (2022). Therapeutic Targets for Heart Failure Identified Using Proteomics and Mendelian Randomization. Circulation , 145 (16), 1205–1217. https://doi.org/10.1161/CIRCULATIONAHA.121.056663

Huang, K., Zhang, S., Wang, H., Qu, Y., Lu, Y., Li, R., Roohani, Y., Qiu, L., Cao, S., Li, G., Zhang, J., Yin, D., Wierenga, R., Kavi, D., Liu, S., She, T., Marwaha, S., Carter, J. N., Zhou, X., … Leskovec, J. (2026). Autonomous biomedical research with an artificial intelligence agent. Science , 393 (6813), eadz4351. https://doi.org/10.1126/science.adz4351

Hubert, T., Mehta, R., Sartran, L., Horváth, M. Z., Žužić, G., Wieser, E., Huang, A., Schrittwieser, J., Schroecker, Y., Masoom, H., Bertolli, O., Zahavy, T., Mandhane, A., Yung, J., Beloshapka, I., Ibarz, B., Veeriah, V., Yu, L., Nash, O., … Silver, D. (2026). Olympiad-level formal mathematical reasoning with reinforcement learning. Nature , 651 (8106), 607–613. https://doi.org/10.1038/s41586-025-09833-y

Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv:2310.06770). arXiv. https://doi.org/10.48550/arXiv.2310.06770

Jin, Z., Chen, Y., Leeb, F., Gresele, L., Kamal, O., Lyu, Z., Blin, K., Gonzalez Adauto, F., Kleiman-Weiner, M., Sachan, M., & Schölkopf, B. (2023). CLadder: Assessing Causal Reasoning in Language Models . 36 , 31038–31065. https://doi.org/10.48550/arXiv.2312.04350

King, E. A., Davis, J. W., & Degner, J. F. (2019). Are drug targets with genetic support twice as likely to be approved? Revised estimates of the impact of genetic support for drug mechanisms on the probability of drug approval. PLOS Genetics , 15 (12), e1008489. https://doi.org/10.1371/journal.pgen.1008489

Koch, Z. (2026). BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks . https://doi.org/10.48550/arXiv.2608.25286

Leban, A., & Sun, Y. (2026). CausalDS: Benchmarking Causal Reasoning in Data-Science Agents. arXiv . https://doi.org/10.48550/arXiv.2607.08093

Leek, J. T., Scharpf, R. B., Bravo, H. C., Simcha, D., Langmead, B., Johnson, W. E., Geman, D., Baggerly, K., & Irizarry, R. A. (2010). Tackling the widespread and critical impact of batch effects in high-throughput data. Nature Reviews Genetics , 11 (10), 733–739. https://doi.org/10.1038/nrg2825

Li, J. H., & Ho, A. J. (2026). GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine. bioRxiv . https://doi.org/10.64898/2026.06.29.735386

Liévin, V., Schmidgall, S., Strother, T., Bijamov, A., Goel, A., Palepu, A., Park, C., Balazadeh, V., Sun, M. W., Guerard, M., Chen, J., Steiner, D., Dhillon, V., Azar, I., Mehta, A., Spetsieris, N., Shah, S., Abdelrahim, M., Dahiya, A., … Yang, L. (2026). ResidencyRL: Reinforcement Learning in Simulated Clinical Environments (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2608.07418

Lind, L., Mazidi, M., Clarke, R., Bennett, D. A., & Zheng, R. (2024). Measured and genetically predicted protein levels and cardiovascular diseases in UK Biobank and China Kadoorie Biobank. Nature Cardiovascular Research , 3 (10), 1189–1198. https://doi.org/10.1038/s44161-024-00545-6

Liu, C.-Y., Liu, Y.-C., Wu, C., Armstrong, A., Volpe, G. J., Van Der Geest, R. J., Liu, Y., Hundley, W. G., Gomes, A. S., Liu, S., Nacif, M., Bluemke, D. A., & Lima, J. A. C. (2013). Evaluation of Age-Related Interstitial Myocardial Fibrosis With Cardiac Magnetic Resonance Contrast-Enhanced T1 Mapping. Journal of the American College of Cardiology , 62 (14), 1280–1287. https://doi.org/10.1016/j.jacc.2013.05.078

Majumder, B. P., Surana, H., Agarwal, D., Mishra, B. D., Meena, A., Prakhar, A., Vora, T., Khot, T., Sabharwal, A., & Clark, P. (2024). DiscoveryBench: Towards Data-Driven Discovery with Large Language Models (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2407.01725

Marques, M. D., Weinberg, R., Kapoor, S., Ostovaneh, M. R., Kato, Y., Liu, C. Y., Shea, S., McClelland, R. L., Post, W. S., Bluemke, D. A., Lima, J. A. C., & Ambale-Venkatesh, B. (2022). Myocardial fibrosis by T1 mapping magnetic resonance imaging predicts incident cardiovascular events and all-cause mortality: The Multi-Ethnic Study of Atherosclerosis. European Heart Journal - Cardiovascular Imaging , 23 (10), 1407–1416. https://doi.org/10.1093/ehjci/jeac010

Minikel, E. V., Painter, J. L., Dong, C. C., & Nelson, M. R. (2024). Refining the impact of genetic evidence on clinical success. Nature , 629 (8012), 624–629. https://doi.org/10.1038/s41586-024-07316-0

Mitchener, L., Laurent, J. M., Tenmann, B., Narayanan, S., Wellawatte, G. P., White, A., Sani, L., & Rodriques, S. G. (2025). BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology . https://doi.org/10.48550/arXiv.2503.00096

Mooij, J. M., Peters, J., Janzing, D., Zscheischler, J., & Schölkopf, B. (2016). Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks. Journal of Machine Learning Research , 17 (32), 1–102.

Munafò, M. R., Tilling, K., Taylor, A. E., Evans, D. M., & Davey Smith, G. (2018). Collider scope: When selection bias can substantially influence observed associations. International Journal of Epidemiology , 47 (1), 226–235. https://doi.org/10.1093/ije/dyx206

Narayanan, S., Braza, J. D., Griffiths, R.-R., Ponnapati, M., Bou, A., Laurent, J., Kabeli, O., Wellawatte, G., Cox, S., Rodriques, S. G., & White, A. D. (2024). Aviary: Training language agents on challenging scientific tasks. arXiv . https://doi.org/10.48550/arXiv.2412.21154

Nelson, M. R., Tipney, H., Painter, J. L., Shen, J., Nicoletti, P., Shen, Y., Floratos, A., Sham, P. C., Li, M. J., Wang, J., Cardon, L. R., Whittaker, J. C., & Sanseau, P. (2015). The support of human genetic evidence for approved drug indications. Nature Genetics , 47 (8), 856–860. https://doi.org/10.1038/ng.3314

Ochoa, D., Hercules, A., Carmona, M., Suveges, D., Baker, J., Malangone, C., Lopez, I., Miranda, A., Cruz-Castillo, C., Fumis, L., Bernal-Llinares, M., Tsukanov, K., Cornu, H., Tsirigos, K., Razuvayevskaya, O., Buniello, A., Schwartzentruber, J., Karim, M., Ariano, B., … McDonagh, E. M. (2023). The next-generation Open Targets Platform: Reimagined, redesigned, rebuilt. Nucleic Acids Research , 51 (D1), D1353–D1359. https://doi.org/10.1093/nar/gkac1046

OpenAI. (2026). On the Navier–Stokes Millennium Prize Problem . https://cdn.openai.com/pdf/32d9f210-8b73-45e0-91bc-82a30aef8a9a/navier-stokes.pdf

Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2025). Training Software Engineering Agents and Verifiers with SWE-Gym . 267 , 47717–47737.

Pun, F. W., Podolskiy, D., Izumchenko, E., Mortlock, A., Oprea, T. I., Scheibye-Knudsen, M., Fortney, K., Morgen, E., Ren, F., & Zhavoronkov, A. (2026). Target identification and assessment in the era of AI. Nature Reviews Drug Discovery , 25 (7), 534–552. https://doi.org/10.1038/s41573-026-01412-8

Rasooly, D., Giambartolomei, C., Peloso, G. M., Dashti, H., Ferolito, B. R., Golden, D., Horimoto, A. R. V. R., Pietzner, M., Farber-Eger, E. H., Wells, Q. S., Bini, G., Proietti, G., Tartaglia, G. G., Kosik, N. M., Wilson, P. W. F., Phillips, L. S., Munroe, P. B., Petersen, S. E., Cho, K., … Joseph, J. (2025). Large-scale multi-omics identifies drug targets for heart failure with reduced and preserved ejection fraction. Nature Cardiovascular Research , 4 (3), 293–311. https://doi.org/10.1038/s44161-025-00609-1

Rasooly, D., Peloso, G. M., Pereira, A. C., Dashti, H., Giambartolomei, C., Wheeler, E., Aung, N., Ferolito, B. R., Pietzner, M., Farber-Eger, E. H., Wells, Q. S., Kosik, N. M., Gaziano, L., Posner, D. C., Bento, A. P., Hui, Q., Liu, C., Aragam, K., Wang, Z., … Casas, J. P. (2023). Genome-wide association analysis and Mendelian randomization proteomics identify drug targets for heart failure. Nature Communications , 14 (1), 3826. https://doi.org/10.1038/s41467-023-39253-3

Reddy, S. G., Cao, F., Xia, R., Loong, S., Chen, E., Steffner, K., O’Sullivan, J. W., Haddad, F., Foo, R., Parikh, V. N., Wheeler, M. T., Ashley, E. A., & Gomes, B. (2025). Deep learning representations and proteome-wide Mendelian randomization identify causal mediators of myocardial fibrosis . Cardiovascular Medicine. https://doi.org/10.64898/2025.12.13.25342200

Replogle, J. M., Saunders, R. A., Pogson, A. N., Hussmann, J. A., Lenail, A., Guna, A., Mascibroda, L., Wagner, E. J., Adelman, K., Lithwick-Yanai, G., Iremadze, N., Oberstrass, F., Lipson, D., Bonnar, J. L., Jost, M., Norman, T. M., & Weissman, J. S. (2022). Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell , 185 (14), 2559-2575.e28. https://doi.org/10.1016/j.cell.2022.05.013

Sanderson, E., Glymour, M. M., Holmes, M. V., Kang, H., Morrison, J., Munafò, M. R., Palmer, T., Schooling, C. M., Wallace, C., Zhao, Q., & Davey Smith, G. (2022). Mendelian randomization. Nature Reviews Methods Primers , 2 (1), 6. https://doi.org/10.1038/s43586-021-00092-5

Smith, S. M., Douaud, G., Chen, W., Hanayik, T., Alfaro-Almagro, F., Sharp, K., & Elliott, L. T. (2021). An expanded set of genome-wide association studies of brain imaging phenotypes in UK Biobank. Nature Neuroscience , 24 (5), 737–745. https://doi.org/10.1038/s41593-021-00826-4

Sun, B. B., Chiou, J., Traylor, M., Benner, C., Hsu, Y.-H., Richardson, T. G., Surendran, P., Mahajan, A., Robins, C., Vasquez-Grinnell, S. G., Hou, L., Kvikstad, E. M., Burren, O. S., Davitte, J., Ferber, K. L., Gillies, C. E., Hedman, Å. K., Hu, S., Lin, T., … Whelan, C. D. (2023). Plasma proteomic associations with genetics and health in the UK Biobank. Nature , 622 (7982), 329–338. https://doi.org/10.1038/s41586-023-06592-6

Wang, R., Jansen, P., Côté, M.-A., & Ammanabrolu, P. (2022). ScienceWorld: Is your Agent Smarter than a 5th Grader? Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 11279–11298. https://doi.org/10.18653/v1/2022.emnlp-main.775

Wen, X., Liu, Z., Zheng, S., Ye, S., Wu, Z., Wang, Y., Xu, Z., Liang, X., Li, J., Miao, Z., Bian, J., & Yang, M. (2025). Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2506.14245

Xi, Z., Ding, Y., Chen, W., Hong, B., Guo, H., Wang, J., Guo, X., Yang, D., Liao, C., He, W., Gao, S., Chen, L., Zheng, R., Zou, Y., Gui, T., Zhang, Q., Qiu, X., Huang, X., Wu, Z., & Jiang, Y.-G. (2025). AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 27914–27961. https://doi.org/10.18653/v1/2025.acl-long.1355

Xu, W., Luo, G., Meng, W., Zhai, X., Zheng, K., Wu, J., Li, Y., Xing, A., Li, J., Li, Z., Zheng, K., & Li, K. (2025). MRAgent: An LLM-based automated agent for causal knowledge discovery in disease via Mendelian randomization. Briefings in Bioinformatics , 26 (2), bbaf140. https://doi.org/10.1093/bib/bbaf140

Yang, J., Zhang, D., Song, X., Dai, Q., Liu, X., Chen, Y., Vashishtha, A., Shi, J., Tan, C., & Peng, H.

(2026). CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists. arXiv . https://doi.org/10.48550/arXiv.2605.26029

Yang, L., Wang, S., & Altman, R. B. (2023). POPDx: An automated framework for patient phenotyping across 392 246 individuals in the UK Biobank study. Journal of the American Medical Informatics Association , 30 (2), 245–255. https://doi.org/10.1093/jamia/ocac226

Yang, Z., Song, Z., Zabad, S., Legault, M.-A., & Li, Y. (2026). PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank. Briefings in Bioinformatics , 27 (1), bbag030. https://doi.org/10.1093/bib/bbag030

Zhang, H. G., Eckmann, P., Miao, J., Mahon, A. B., & Zou, J. (2026). The Virtual Biotech: A multi-agent AI framework for therapeutic discovery and development. Science , eaeg6779. https://doi.org/10.1126/science.aeg6779

Zheng, J., Haberland, V., Baird, D., Walker, V., Haycock, P. C., Hurle, M. R., Gutteridge, A., Erola, P., Liu, Y., Luo, S., Robinson, J., Richardson, T. G., Staley, J. R., Elsworth, B., Burgess, S., Sun, B. B., Danesh, J., Runz, H., Maranville, J. C., … Gaunt, T. R. (2020). Phenome-wide Mendelian randomization mapping the influence of the plasma proteome on complex diseases. Nature Genetics , 52 (10), 1122–1131. https://doi.org/10.1038/s41588-020-0682-6

Zhou, Y., Wu, X., Huang, B., Wu, J., Feng, L., & Tan, K. C. (2024). CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs. arXiv . https://doi.org/10.48550/arXiv.2404.06349

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI