Skip to content

Causal Traps for Evaluating AI Scientists

DrugTargetWorld plants realistic causal traps, such as confounding, reverse causation, collider bias, pleiotropy and surrogate-outcome discordance, so that observational association alone leads an AI scientist to the wrong target.

From DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists · arXiv:2610.09558

Why aren't observational associations sufficient for evaluating AI scientists?

In a real biobank, a protein can be strongly associated with disease without causing it. Confounding, reverse causation, selection into the measured subcohort, pleiotropy and measurement artifacts all produce associations that look like causal evidence. An AI scientist that ranks proteins by association will nominate them, and in real data there is no answer key to show that it was wrong.

DrugTargetWorld builds these failure modes into the data-generating process. Each bias alters the causal or measurement process rather than adding random noise, so the proteins it affects can appear scientifically plausible, and the evaluator knows exactly which ones are decoys.

Which causal traps does DrugTargetWorld plant?

The generator can plant nine predefined biases, T1 to T9. Each world carries a subset; the count below is the number of worlds in the 20-world version 1 panel that carry it.

CodeTrapHow it misleadsWorlds
T1ConfoundingBMI raises both a non-causal protein and latent disease severity.18
T2Reverse causationDisease severity raises the protein, not the other way round.14
T3Selection (collider) biasDisease severity and the protein both raise the chance of being imaged, inducing an association among imaged participants.12
T4Causal non-identifiabilityOne variant drives a causal driver and a non-causal partner, so genetic evidence alone cannot separate them.6
T5Imaging batch effectScanner site correlates with ancestry and biases image intensity, so ancestry-related variants can appear to affect imaging.20
T6Benign remodelingExercise changes cardiac volumes without changing disease severity.20
T7Instrument pleiotropyThe protein's cis variant also affects disease directly, violating the exclusion restriction.8
T8Surrogate-outcome discordanceA protein improves imaging phenotypes but increases cardiac mortality and is not a driver.20
T9Assay unit mixingAt one site a disease-null protein is reported on a different scale.5

In addition, 15 worlds contain a delayed-effect driver: a causal driver whose effect appears mainly over five years of follow-up, paired with a non-causal partner that tracks the later disease. Together with T4, this creates ambiguous pairs that observational data cannot separate; only an experiment can.

How does DrugTargetWorld evaluate causal reasoning against these traps?

Because the structural causal model is known to the evaluator, each claim can be checked against truth. Bias identification (15 points) rewards rejecting a planted non-causal protein together with its correct source of bias, multiplying recall by precision so that indiscriminate rejection earns little. Causal confidence (25 points) rewards matching claims to evidence on ambiguous pairs: an experimentally tested causal choice earns full credit, an explicit statement that the pair cannot yet be resolved earns partial credit, and an unsupported pick earns none.

The most consequential error is penalised directly. Advancing the T8 harmful surrogate as beneficial costs 30 points, the full weight of target identification.

How did AI agents handle causal traps?

No agent averaged more than 2.87 of 15 points for bias identification. The reverse-causation protein, rejected most often, was rejected in 32.0% of eligible episodes and with the correct mechanism in 14.3%. All models nominated the confounding (T1) and selection (T3) proteins in fewer than 9% of eligible episodes, but the reverse-causation protein was nominated in half of the Haiku 4.5 and GPT-OSS-20B episodes.

The harmful surrogate was nominated in 44 of 60 Opus 5 and 25 of 60 GPT-5.6 Sol episodes. These agents usually labelled it harmful to outcomes (30 of 44 and 16 of 25), avoiding the penalty; every surrogate nomination by Haiku 4.5 and GPT-OSS-20B claimed it was beneficial. Instrument strength, which the leading agents checked almost always, does not establish the exclusion restriction that pleiotropy (T7) violates. See findings.

Cards showing how each planted bias is built into a world and the misleading signal it creates
Figure 3 | Planted biases in DrugTargetWorld. Each card shows how a bias is built into a world's causal or measurement model and the misleading signal it creates.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI