Skip to content

Benchmarking End-to-End Scientific Discovery

Most scientific-agent benchmarks specify the question and the analysis. DrugTargetWorld leaves the whole research strategy to the AI scientist and scores the result against causal truth.

From DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists · arXiv:2610.09558

Why evaluate complete scientific workflows rather than isolated tasks?

Recent benchmarks have moved from testing static scientific knowledge toward testing whether agents can execute research workflows. In most of them, though, the research question and often the analysis methods are specified. In real drug target discovery the hard decisions come earlier and between steps: how to measure the disease, which analyses to run, which candidates deserve an expensive experiment, and when the evidence is enough.

Evaluating the whole workflow matters particularly here because experimental validation is costly in both money and time, so an agent must prioritise which candidates merit testing.

What makes DrugTargetWorld an open-strategy benchmark?

The agent receives one world's released data and no analysis plan. It chooses the phenotype, the hypotheses, the analyses, the experiments and the stopping point. The harness is deliberately minimal: it runs the Python programs the agent writes, returns their output, and gives no feedback on scientific correctness, so differences in score reflect the agents' own research strategies.

BenchmarkTaskMethod choiceSequentialSimulatedGround truthBiobankTrainable
ScienceAgentBenchScientific analysisDefinedLimitedNoNoNoNo
DiscoveryBenchScientific discoveryDefinedLimitedVariableVariableNoNo
BixBenchBioinformaticsDefinedYesNoNoNoNo
BixBench3Full study executionGuidedYesNoNoNoNo
GeneBench-ProGenetic analysisOpen analysisYesYesYesNoNo
AviaryScientific tasksTask specificYesVariableVariableNoYes
ResidencyRLClinical careScenario definedYesYesScenarioNoYes
Causal benchmarksCausal discoveryVariableVariableYesYesNoVariable
DrugTargetWorldDrug target discoveryOpen strategyYesYesYesYesYes

From Table 1 of the paper. "Open strategy" means the agent chooses the phenotype, hypotheses, analyses, experiments and stopping point; "Trainable" means the environment can supply a verifiable reward for training.

How is scientific discovery framed as sequential decision-making?

An episode is a sequence of up to 30 turns. At each turn the agent chooses its next analysis or experiment from the evidence gathered so far and the budget that remains, at ~ π(a | st, Bt), and experiments reduce the budget, Bt+1 = Bt − c(at), while analyses of the released data are free. The same formulation defines the policy that training in this environment would optimise. See budgeted experimentation.

Where do autonomous AI scientists break down in open-ended research?

Failures arose at several steps. Phenotype construction earned zero credit in 77.0% of episodes. Agents generated evidence without using it: Opus 5 and GPT-5.6 Sol read genotypes in all 60 episodes, but instrumental-variable analyses appeared in only 41 and 29. Code execution failed in smaller models, and Devstral Small hit the 30-turn limit in 53 of 60 episodes by repeatedly re-deriving information.

The number of core workflow milestones an episode reached correlated with its score (Spearman ρ = 0.649), a descriptive relationship confounded by model identity. And correct answers did not always reflect supporting analysis. See findings.

Research workflow milestones and failure modes of autonomous agents
Figure 6 | Research workflow and failure modes of autonomous agents. Phenotype outcomes, workflow milestones reached, scores by milestones reached, and failure modes such as hitting the turn limit.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI