Skip to content

About

Who builds DrugTargetWorld, what it is for, and how it is released.

DrugTargetWorld is an open benchmark for training and evaluating AI scientists on end-to-end drug target discovery. It procedurally generates synthetic biobanks, or worlds, each governed by a hidden causal model, so every causal claim an agent makes can be scored against a known truth.

Each world holds genotypes, proteomics, transcriptomics, metabolomics, health records, outcomes and synthetic cardiac MRI for 54,000 participants. Agents build a disease phenotype, identify the proteins that causally drive disease, choose a therapeutic direction for each, and can buy virtual experiments under a fixed budget. See synthetic biobank and drug target discovery.

DrugTargetWorld is research from the Department of Biomedical Data Science and the Department of Medicine at Stanford University, with collaborators at Brown University. It is described in the paper DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists by 15 authors; the corresponding author is Bruna Gomes. The full author list, affiliations and citation are on the paper page.

The code is released under the MIT License on GitHub. The benchmark runs in Harbor as drugtargetbench/drugtargetbench@v1.1; the name predates the project's rename and is kept so published runs stay reproducible.

The world data is on Hugging Face. The generated omics, health-record and outcome data are the project's own. The participants are simulated and carry no identifiers. The cine-MRI is warped from real anatomy in the ACDC dataset, so the dataset is gated behind the ACDC terms of use.