Skip to content

Procedurally Generated Scientific Worlds for Training AI Scientists

DrugTargetWorld generates scientific worlds from hidden structural causal models, so AI scientists can be evaluated and trained against a causal ground truth that is known to the evaluator but concealed from the agent.

From DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists · arXiv:2610.09558

Why do AI scientists need environments with verifiable ground truth?

Much of the recent progress of AI systems in mathematics and software engineering relies on a verifiable reward: an automatic check of whether an output is correct, which can both evaluate and train an agent. Drug target discovery in real biobanks lacks one. The causal links between proteins and disease are only partially known, so an agent's nomination of a causal driver requires experimental validation, and even compelling observational evidence can reflect confounding, reverse causation, selection bias, pleiotropy or measurement artifacts.

Real data also resists the repeated, large-scale interaction that training requires: participant-level biobank data is released under controlled-access agreements and analysed inside approved computing environments.

How does DrugTargetWorld procedurally generate a scientific world?

Each world is generated from a structural causal model that specifies genetic variation, molecular measurements, latent disease severity, observed disease phenotypes, longitudinal outcomes and responses to experimental intervention. A procedural seed determines the world-level parameters: the causal drivers, their effect sizes and directions, the disease archetype and the planted biases. They are sampled once and held fixed for every agent evaluated in that world, so score differences within a world reflect the agents' analyses rather than differences in the data.

Distributions are calibrated to aggregate statistics from the Multi-Ethnic Study of Atherosclerosis (MESA) and TOPMed-linked MESA data; only cohort-level summaries were used, and the causal structure was designed rather than estimated. The version 1 panel contains 20 worlds: six each of dilated cardiomyopathy, hypertrophic cardiomyopathy and heart failure with preserved ejection fraction, and two ischemic worlds. One null world has no causal driver; the others have between two and six.

Why does a known but concealed ground truth matter?

Because the causal model is known to the evaluator but hidden from the agent, every stage of the research process can be scored against an exact answer: the phenotype the agent builds, the drivers it nominates, the direction it recommends, the biases it rejects and the confidence it claims. Varying the causal structure across worlds produces progressively harder discovery problems while keeping that exact answer.

Proteins carry anonymous labels with no mapping to real genes, so an agent cannot reach an answer through published knowledge, and new worlds can be generated whenever a test set might have been seen. This makes the evaluation contamination-resistant by construction.

Can procedurally generated worlds be used to train scientific strategies?

Freshly generated worlds, each with its own ground truth, can supply a verifiable reward for training agents across the entire workflow, an approach that has already improved agents in other interactive environments. Strategies learned this way can then be frozen and tested on unseen worlds and in real cohorts.

The evaluation also shows the limits of an outcome-only reward: three episodes recovered the exact set of drivers, two of them without supporting analysis, and still earned full credit. A training reward should therefore be paired with trajectory annotation and exploit audits, and smaller open-weight models, which scored zero in 78% of episodes, will likely need denser intermediate rewards or a curriculum of easier worlds. See findings.

DrugTargetWorld generates multiple causal biobank worlds and scores agent trajectories against hidden ground truth
Figure 1 | DrugTargetWorld generates multiple causal biobank worlds and evaluates complete agent research trajectories against hidden ground truth. Each world comes from its own hidden causal model; repeating the procedure with different seeds produces a collection of distinct worlds rather than a single dataset.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI