Skip to content

Budgeted Experimentation and Verifiable Reward for AI Scientists

In DrugTargetWorld an AI scientist must decide which experiments are worth their cost. Virtual knockdowns and cell perturbations re-run the hidden causal model, so every experiment returns a true counterfactual and every claim can be rewarded against it.

From DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists · arXiv:2610.09558

Why must AI scientists decide which experiments are worth conducting?

Laboratory experiments are expensive, so a useful scientist must judge when an experiment is worth its cost and which uncertainty it should resolve. Some questions cannot be settled from observational data at all: when two proteins share a genetic instrument, or when a causal effect appears only over years of follow-up, genetics alone cannot say which protein is causal.

DrugTargetWorld evaluates each agent under three budgets: observational-only ($0), limited ($450,000), enough for one knockdown or three cell perturbations, and expanded ($2 million), enough for up to five knockdowns. The amounts are an ordinal scale of experimental access, not estimates of real costs. Analysis of the released data is always free.

What virtual experiments can an agent run?

  • Cell perturbation ($150,000): a smaller virtual experiment with measurement noise that measures the effect at the first visit.
  • In vivo knockdown ($400,000): a larger virtual experiment that returns the effect over five years of simulated follow-up, without additional noise.

Either applies to any protein and returns the change in cavity radius, wall thickness, ejection fraction and five-year survival relative to the unperturbed state. Because the generator is the ground truth, perturbing a protein re-runs the downstream structural equations: the result is a counterfactual, not a lookup, and a protein that is not a driver returns no effect. The cheaper experiment cannot detect a delayed effect, so choosing between them is part of the task.

How does budgeted experimentation provide a verifiable reward?

The agent's work is a sequence of decisions: at ~ π(a | st, Bt), with the budget falling by each experiment's cost, Bt+1 = Bt − c(at). The causal-confidence score (25 points) ties claims to evidence on pairs that observational data cannot separate: naming the causal protein after testing it experimentally earns 25 points, acknowledging that the pair is unresolved earns up to 17.5, and an unsupported choice earns nothing. Without a budget, 17.5 is therefore the maximum per pair.

Because each world's causal structure is known to the scorer, this reward can train agents across the whole workflow, from phenotype to experiment to final claim. See scientific worlds.

Do AI agents allocate experimental resources well?

Of the 360 episodes with a budget, 143 received at least one experiment, and 413 experiments were delivered, 76.5% of them by Opus 5, GPT-5.6 Sol and Sonnet 5. Of those experiments, 171 (41.4%) targeted a causal driver, 52 (12.6%) the harmful surrogate and 49 (11.9%) the reverse-causation protein. Where a pair shared a genetic instrument, the leading agents tested the causal member in 14 of 24 eligible episodes, and every driver they tested and nominated carried the direction its knockdown implied. No experiment targeted any of the 15 delayed-effect pairs.

Agents usually read returned results (124 of 143 episodes) and 84 changed their conclusion; 16 of the 19 that did not had submitted in the same turn as their request. Yet more budget changed little for the leading agents: +0.13 points for Opus 5 and +0.07 for GPT-5.6 Sol under the expanded budget, despite spending 93% and 95% of it. See findings.

What agents did with the planted biases and the targets of their 413 experiments
Figure 5 | What agents did with the planted biases. Panel c shows the targets of the 413 delivered experiments: true drivers, the harmful surrogate, the reverse-causation protein, or other proteins.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI