Skip to content

DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

Supplementary Figures:

Model performance across DrugTargetWorld
Figure S1 | Model performance across DrugTargetWorld. Mean score across the three budget regimes for each of nine agent systems in each of 20 worlds. Worlds are labeled by cardiac archetype and design class. Scores range from 0 to 100, with higher scores indicating better benchmark performance.
Cross-benchmark performance of evaluated models
Figure S2 | Cross-benchmark performance of evaluated models. Mean DrugTargetWorld score compared with reported performance on Terminal-Bench and SWE-bench Pro for the same model families.

Table S1 | Model roster and run provenance

Arm Model identifier Serving Hardware Dates tested (UTC)
opus-5 claude-opus-5 Anthropic API vendor 2026-09-07 18:37 to 2026-09-08 00:07
gpt-5.6-sol gpt-5.6-sol OpenAI API vendor 2026-09-05 22:29 to 2026-09-06 01:34
sonnet-5 claude-sonnet-5 Anthropic API vendor 2026-09-04 22:29 to 2026-09-05 00:59
haiku-4-5 claude-haiku-4-5 Anthropic API vendor 2026-09-04 15:25 to 17:28
gpt-oss-20b openai/gpt-oss-20b self-hosted vLLM 0.10.2 1 x H100 80GB 2026-09-04 15:13 to 16:44
qwen3-coder-30b Qwen/Qwen3-Coder-30 B-A3B-Instruct self-hosted vLLM 0.10.2, TP=2 2 x H100 80GB 2026-09-04 18:03 to 18:37
glm-4-32b zai-org/GLM-4-32B-0414 self-hosted vLLM 0.10.2, TP=2 2 x H100 80GB 2026-09-04 18:02 to 20:20
qwen3-8b Qwen/Qwen3-8B self-hosted vLLM 0.10.2, TP=1 1 x H100 80GB 2026-09-04 15:44 to 20:57
devstral-small mistralai/Devstral-Small-2507 self-hosted vLLM 0.10.2, TP=2 2 x H100 80GB 2026-09-04 17:59 to 20:22
human:bruna human player cardioseek-play, harness ced47324 n/a 2026-09-07 (standard-03; observational only)

Supplementary Table S2 | Inference cost and token use

Arm Provider Episodes Input tokens Output tokens API cost($) GPU-h Imputed GPU cost ($) Cost / episode ($)
opus-5 Anthropic 60 39,353,164 2,056,198 248.17 - - 4.14
gpt-5.6-sol OpenAI 60 26,602,045 739,149 121.19 - - 2.02
sonnet-5 Anthropic 60 21,247,607 1,616,687 87.99 - - 1.47
haiku-4-5 Anthropic 60 10,050,083 1,060,240 15.35 - - 0.26
devstral-small self-hosted 60 18,781,760 392,556 0 4.76 11.90 -
glm-4-32b self-hosted 60 16,047,705 598,681 0 4.72 11.80 -
qwen3-8b self-hosted 60 10,833,367 2,136,841 0 5.29 13.22 -
qwen3-coder-30b self-hosted 60 3,249,958 469,101 0 1.28 3.20 -
gpt-oss-20b self-hosted 60 3,086,488 468,774 0 1.56 3.90 -
Total - 540 149,252,177 9,538,227 472.70 17.61 44.02 -

Supplementary Table S3 | Evaluation-world design and phenotype-scoring baselines.

Panel A. World design

World ID Split Archetype Generation suite Hard Messy Null Causal drivers
hard-01DevelopmentDCMcardioseek-v1-hard-aYesNoNo5
hard-02DevelopmentHCMcardioseek-v1-hard-bYesNoNo3
hard-03DevelopmentHFpEFcardioseek-v1-hard-cYesNoNo6
hard-04DevelopmentDCMcardioseek-v1-hard-aYesNoNo4
hard-messy-01DevelopmentHCMcardioseek-v1-hard-messyYesYesNo5
heldout-dcm-01Held-outDCMcardioseek-v1-standard-aNoNoNo5
heldout-dcm-02Held-outDCMcardioseek-v1-hard-messyYesYesNo5
heldout-hcm-01Held-outHCMcardioseek-v1-standard-bNoNoNo2
heldout-hcm-02Held-outHCMcardioseek-v1-hard-bYesNoNo2
heldout-hfpef-01Held-outHFpEFcardioseek-v1-messyNoYesNo5
heldout-hfpef-02Held-outHFpEFcardioseek-v1-standard-aNoNoNo2
heldout-ischemic-01Held-outIschemiccardioseek-v1-standard-bNoNoNo4
heldout-ischemic-02Held-outIschemiccardioseek-v1-hard-messyYesYesNo4
high-concordance-01DevelopmentHCMcardioseek-v1-high-concordanceNoNoNo2
low-concordance-01DevelopmentDCMcardioseek-v1-low-concordanceNoNoNo3
messy-01DevelopmentHFpEFcardioseek-v1-messyNoYesNo5
null-01DevelopmentHFpEFcardioseek-v1-nullNoNoYes0
standard-01DevelopmentDCMcardioseek-v1-standard-aNoNoNo3
standard-02DevelopmentHCMcardioseek-v1-standard-bNoNoNo6
standard-03DevelopmentHFpEFcardioseek-v1-standard-aNoNoNo4

Panel B. Dimensions, causal non-identifiability, delayed effects, and phenotype baseline

World ID Participants MRI participants Proteins Variants T4 causal non-identifiability Delayed-effect setting Phenotype baseline, b
hard-0154,00010,8002,9418,192YesYes0.4010
hard-0254,00010,8002,9418,192NoYes0.6748
hard-0354,00010,8002,9418,192NoYes0.7473
hard-0454,00010,8002,9418,192NoYes0.4237
hard-messy-0154,00010,8002,9418,192NoYes0.6674
heldout-dcm-0154,00010,8002,9418,192NoYes0.4317
heldout-dcm-0254,00010,8002,9418,192NoYes0.4419
heldout-hcm-0154,00010,8002,9418,192YesNo0.6745
heldout-hcm-0254,00010,8002,9418,192NoNo0.6839
heldout-hfpef-0154,00010,8002,9418,192NoYes0.7527
heldout-hfpef-0254,00010,8002,9418,192YesNo0.7525
heldout-ischemic-0154,00010,8002,9418,192NoYes0.7172
heldout-ischemic-0254,00010,8002,9418,192NoYes0.7089
high-concordance-0154,00010,8002,9418,192NoNo0.6792
low-concordance-0154,00010,8002,9418,192YesYes0.4329
messy-0154,00010,8002,9418,192YesYes0.7465
null-0154,00010,8002,9418,192NoNo0.7505
standard-0154,00010,8002,9418,192YesYes0.4219
standard-0254,00010,8002,9418,192NoYes0.6743
standard-0354,00010,8002,9418,192NoYes0.7645

Notes. The fixed evaluation panel contained 20 worlds, each with 54,000 participants, 10,800 participants with MRI, 2,941 proteins, and 8,192 variants. T4 denotes the planted causal non-identifiability bias; in these worlds a causal driver and a noncausal alternative cannot be separated from the observational genetic evidence alone. Delayed-effect setting denotes worlds in which causal consequences emerge primarily over longer follow-up. The phenotype baseline b is the world-specific built-in absolute correlation used in the phenotype scoring function described in Supplementary Methods S2.5. DCM, dilated cardiomyopathy; HCM, hypertrophic cardiomyopathy; HFpEF, heart failure with preserved ejection fraction; MRI, magnetic resonance imaging.

Supplementary Table S4 | Benchmark performance by model.

Panel A. Total-score distribution

Model Episodes Mean SD Median IQR Score >0, n (%) Score >=50, n (%) Score >=80,n (%)
Opus 56039.9820.6541.2129.40-49.7456 (93.3%)15 (25.0%)2 (3.3%)
GPT-5.6 Sol6035.3820.6536.5920.00-48.6655 (91.7%)15 (25.0%)2 (3.3%)
Sonnet 56021.3322.0112.500.00-36.0843 (71.7%)6 (10.0%)2 (3.3%)
Haiku 4.56012.9217.018.300.00-19.6338 (63.3%)3 (5.0%)0 (0.0%)
GPT-OSS-20B607.4115.330.000.00-10.0024 (40.0%)2 (3.3%)0 (0.0%)
Qwen3-Coder-30B-A3B605.9012.300.000.00-7.5023 (38.3%)1 (1.7%)0 (0.0%)
GLM-4-32B601.464.790.000.00-0.006 (10.0%)0 (0.0%)0 (0.0%)
Qwen3-8B601.314.620.000.00-0.009 (15.0%)0 (0.0%)0 (0.0%)
Devstral Small600.813.320.000.00-0.004 (6.7%)0 (0.0%)0 (0.0%)

Panel B. Target recovery and scoring components

Model Target precision Target recall Target identification /30 Causal confidence/25 Bias identification /15 Direction /20 Phenotype /10 Safety penalties,n
Opus 50.6760.63614.804.262.8713.057.196
GPT-5.6 Sol0.6990.63915.814.191.0213.003.165
Sonnet 50.6000.34810.241.981.346.961.163
Haiku 4.50.3040.3395.271.610.445.910.215
GPT-OSS-20B0.1670.2002.940.920.003.430.683
Qwen3-Coder-30B-A3B0.1930.1703.260.560.001.940.140
GLM-4-32B0.0250.0171.120.000.000.250.080
Qwen3-8B0.0340.0240.700.000.000.220.390
Devstral Small0.0000.0000.750.000.000.000.060

Notes. Each model was evaluated in 60 episodes: one episode in each of 20 worlds under each of three experimental-budget conditions. Scores are reported on the 0-100 benchmark scale. SD describes variation across worlds and budget conditions, not run-to-run replicate variability. IQR is the first to third quartile. Target precision and recall are episode-level means. Safety penalties are the number of episodes in which the 30-point harmful-surrogate penalty was applied.

Supplementary Table S5 | Pairwise performance comparisons among API-served models.

Model A Model B Mean A MeanB Mean difference(A-B) 95% CI Paired-tP Worlds A>B/ A<B / tied Exact sign-testP
Opus 5GPT-5.6 Sol39.9835.384.60-0.26 to 9.450.06216/4/00.012
Opus 5Sonnet 539.9821.3318.6512.88 to 24.41<0.00119/1/0<0.001
Opus 5Haiku 4.539.9812.9227.0619.57 to 34.55<0.00119/1/0<0.001
GPT-5.6 SolSonnet 535.3821.3314.057.57 to 20.53<0.00118/2/0<0.001
GPT-5.6 SolHaiku 4.535.3812.9222.4715.54 to 29.40<0.00120/0/0<0.001
Sonnet 5Haiku 4.521.3312.928.411.86 to 14.970.01517/3/00.003

Notes. For each model and world, scores were first averaged across the three budget conditions. Pairwise differences were then calculated across the 20 worlds. Confidence intervals are t-based 95% confidence intervals for the paired world-level differences. Paired-t P values are exploratory two-sided tests. Exact sign tests are two-sided binomial tests with tied worlds excluded. P values are not adjusted for multiplicity.

Supplementary Table S6 | Performance by experimental-budget condition.

Model Observational mean Limited($450k) mean Expanded($2M) mean Limited -observational Expanded -observational 95% CI forexpanded contrast Paired-t P Exact sign-test P
Opus 541.1037.6141.23-3.490.13-8.71 to 8.980.9750.503
GPT-5.6 Sol35.7234.6435.79-1.090.07-10.08 to 10.220.9890.115
Sonnet 522.0510.0231.91-12.039.86-2.40 to 22.120.1090.238
Haiku 4.513.8112.2412.70-1.56-1.11-7.69 to 5.470.7280.143
GPT-OSS-20B0.929.6511.658.7310.741.95 to 19.530.0190.003
Qwen3-Coder-30B-A3B1.207.129.385.928.193.02 to 13.350.0040.092
GLM-4-32B2.000.991.38-1.01-0.62-1.93 to 0.680.3301.000
Qwen3-8B0.461.911.571.451.11-0.86 to 3.080.2540.453
Devstral Small0.750.750.940.000.19-0.21 to 0.600.3301.000

Notes. Budget contrasts were paired by world (n = 20 worlds per model). The limited condition provided a $450,000 experimental budget and the expanded condition a $2,000,000 budget; observational episodes had no experimental budget. The confidence interval and P values shown are for the expanded-minus-observational contrast. Experimental budgets and costs are simulated benchmark resource constraints.

Supplementary Table S7 | Phenotype construction and internal validation.

Panel A. Phenotype performance by model

Model Phenotypeat scored path, n Scorable, n Positive credit, n Full credit, n Mean phenotype score /10 Mean |r| SD |r| Internally validated, n
Opus 5606056177.190.7850.10554
GPT-5.6 Sol60603433.160.6370.19140
Sonnet 553531201.160.5020.23544
Haiku 4.55742400.210.3520.2214
GPT-OSS-20B5647710.680.3720.2410
Qwen3-Coder-30B-A3B3737300.140.3760.1590
GLM-4-32B1010100.080.2740.2510
Qwen3-8B4835600.390.3950.2060
Devstral Small55100.060.3680.2580

Panel B. Performance by phenotype-construction strategy

Construction strategy Scorable, n Mean |r| SD |r| Median |r| IQR Mean phenotype score /10
Weighted composite1760.5280.2420.5750.331-0.7262.23
Single imaging measure950.3600.2180.3180.274-0.5330.48
Outcome-supervised420.7510.1100.7710.637-0.8386.10
Other140.2950.2360.2880.111-0.3270.71
PCA/SVD110.5540.2420.6300.397-0.7202.21
Covariate-residualized imaging60.6240.2960.7280.573-0.8243.21
Factor model40.8440.0240.8480.832-0.8599.16
None10.483-0.4830.483-0.4830.00

Panel C. Association of internal validation with phenotype performance

Validated before submission Scorable, n Mean |r| SD |r| Median |r| Mean phenotype score /10
Yes1390.6630.2110.7054.55
No2100.4020.2280.3960.73

Panel D. Clinical or imaging endpoints used for internal validation

Validation endpoint Episodes
Mortality104
MACE97
Coronary CT calcium or stenosis99
ECG features76
EHR diagnoses26
Repeat imaging visit9

Notes. Scorable phenotypes are submissions for which a phenotype-latent-disease correlation could be evaluated. |r| is the absolute correlation between the submitted phenotype and the hidden latent disease severity. Positive credit indicates a phenotype score >0; full credit indicates a phenotype score of 10. Strategy labels are trajectory annotations of the submitted phenotype construction. Panel C is restricted to the 349 episodes with scorable phenotype correlations. Panel D counts validation checks across all annotated episodes that performed internal validation; an episode could use more than one endpoint, so counts are not mutually exclusive. CT, computed tomography; ECG, electrocardiogram; EHR, electronic health record; MACE, major adverse cardiovascular events; PCA, principal component analysis; SVD, singular value decomposition.

Supplementary Table S8 | Research strategy and process annotations by model.

Panel A. Causal-analysis strategy

Model Cis MR primary route, n/60 Used geneticinstruments,n/60 Checked instrument strength, n/60 Checked exclusion restriction, n/60 Mean hypothesis revisions Evidence-driven revision,n/60
Opus 5576059332.0560
GPT-5.6 Sol505957371.6059
Sonnet 523292211.1556
Haiku 4.5020301.1049
GPT-OSS-20B17000.100
Qwen3-Coder-30B-A3B01000.152
GLM-4-32B00000.022
Qwen3-8B08000.200
Devstral Small01000.000

Panel B. Experimental planning and process

Model Budget shaped plan, n/60 Experiment planned before purchase, n/60 Recognized unresolvable pair, n/60 Read data dictionary, n/60 Ended deliberately,n/60 Ended with analysis unfinished (trace annotation), n/60
Opus 5195660143
GPT-5.6 Sol1030560471
Sonnet 576516064
Haiku 4.516171443462
GPT-OSS-20B002101920
Qwen3-Coder-30B-A3B10011285
GLM-4-32B00835247
Qwen3-8B201314924
Devstral Small00033154

Notes. Counts derive from the trajectory strategy-annotation archive and describe recorded research behavior; they do not establish that an analysis was statistically valid or correctly interpreted. "Cis MR primary route" denotes the annotated strategy supporting final target prioritization and is distinct from the stricter operation-based MR/IV calculation audit reported in the main text. Actual harness turn-limit terminations are reported in Table S10. MR, Mendelian randomization; IV, instrumental variable.

Supplementary Table S9 | Experimental selection and use.

Panel A. Experimental use by model

Model Eligible episodes Requested>=1, n Received >=1, n Cell perturbations delivered Knockdowns delivered Refused requests Total simulated spend
Opus 540393981100$45,200,000
GPT-5.6 Sol40393901150$46,000,000
Sonnet 54034345782$31,950,000
Haiku 4.540181812504$21,800,000
GPT-OSS-20B4011010$400,000
Qwen3-Coder-30B-A3B4033312$850,000
GLM-4-32B4022210$700,000
Qwen3-8B408661626$7,300,000
Devstral Small4011055$2,000,000

Panel B. Targets of delivered experiments

Target category Experiments Percent of delivered
True causal driver17141.4%
T8 harmful surrogate5212.6%
Reverse-causation (T2) protein4911.9%
Confounding (T1) protein10.2%
Selection/collider (T3) protein10.2%
Other protein13933.7%

Panel C. Use of returned experimental results

Group Purchasing episodes Referenced returned result in code Percent referenced Changed conclusion, n (%)
All purchasing episodes14312486.7%84 (58.7%)
Opus 5393794.34%-
GPT-5.6 Sol393794.34%-
Sonnet 53434100%-
Haiku 4.5181161.1%-

Notes. There were 360 budget-eligible episodes. Across these episodes, 145 requested an experiment and 143 received at least one; 413 experiments were delivered, comprising 36 cell perturbations and 377 in vivo knockdowns. Thirty-nine requests were refused in budget-eligible episodes. Eleven additional requests were made and refused in observational episodes, yielding 50 refused requests overall. Total spend is the sum of simulated experimental charges across a model's 40 budget-eligible episodes. Panel C uses the source-linked code audit for returned-result reference; the changed-conclusion count is reported for all purchasing episodes. Target categories are mutually exclusive.

Supplementary Table S10 | Workflow failures and execution attrition by model.

Panel A. Phenotype and execution outcomes

Model No phenotype at scored path, n Phenotype present butzero credit, n Positive phenotypecredit, n No parseablesubmission at exit, n Reached 30-turn limit, n >=1 no-code reply, n Mean failed code turns
Opus 504560033.15
GPT-5.6 Sol026340004.48
Sonnet 57411211491.67
Haiku 4.535340002.97
GPT-OSS-20B44972263.45
Qwen3-Coder-30B-A3B233430012.82
GLM-4-32B50914038716.80
Qwen3-8B1242614141012.42
Devstral Small554154531215.80

Panel B. Additional failure and attrition indicators

Model Genotype accessed without genetic-instrument strategy, n Nomination cap engaged, n Safety penalty, n Context compaction, n Truncated completion, n
Opus 500630
GPT-5.6 Sol105140
Sonnet 5110300
Haiku 4.5151510
GPT-OSS-20B03310
Qwen3-Coder-30B-A3B250000
GLM-4-32B2010271
Qwen3-8B1330134
Devstral Small80030

Notes. All counts use 60 episodes per model unless otherwise indicated. Phenotype-file categories use the required scored submission path. "No parseable submission at exit" includes episodes for which no valid final submission was available when the episode was scored. A no-code reply consumed a turn without executable code. The genotype-use indicator combines harness-recorded genotype access with the strategy annotation for genetic-instrument use. The nomination cap is the scorer's 25-nomination limit. Context compaction and truncated completion are harness-level events.

Supplementary Table S11 | Sensitivity analyses.

Panel A. Model performance after removing phenotype credit

Model Original mean /100 No-phenotype mean /90 Original rank No-phenotype rank
Opus 539.9833.2411
GPT-5.6 Sol35.3832.3322
Sonnet 521.3320.2433
Haiku 4.512.9212.7144
GPT-OSS-20B7.416.7655
Qwen3-Coder-30B-A3B5.905.7666
GLM-4-32B1.461.3877
Qwen3-8B1.310.9288
Devstral Small0.810.7599

Panel B. Opus 5 versus GPT-5.6 Sol with and without phenotype credit

Analysis Worlds Mean difference (Opus-GPT) 95% CI Paired-tP Worlds Opus>GPT / GPT>Opus / tied Exact sign-test P Bootstrap 95% CI
Primary score204.60-0.26 to 9.450.06216/4/00.012-
No-phenotype score200.91-3.64 to 5.460.68010/10/01.000-3.08 to 5.23

Panel C. Fixed-label sensitivity for the two leading models

Model Original direction /20 Fixed-inhibit direction Original causal confidence /25 Original total /100 Fixed-label total Original safety penalties Always-aligned safety penalties
Opus 513.055.82-6.484.2639.9816.19-16.86644
GPT-5.6 Sol13.005.77-6.434.1935.3821.57-22.24525
All models-----2289

Panel D. Sensitivity to the follow-up MRI generation error

Analysis Leading-model cells Worlds Follow-up MRIread episodes (all 540) Opus-GPT mean difference 95% CI
Primary, all worlds4020394.60-0.26 to 9.45
Exclude leading model-world cells with any follow-up MRI read3717396.280.60 to 11.97

Notes. Panel A removes the 10-point phenotype component, recomputes the remaining components and signed safety penalty, and clips the score to 0-90 without rescaling. Panel B averages the three budgets within each world before comparing Opus 5 with GPT-5.6 Sol; the bootstrap interval is a 20,000-resample paired-world percentile interval and is shown only for the no-phenotype analysis. Panel C holds target selection, experiments, abstentions, rejections, and phenotype scores fixed while replacing every direction label with "inhibit" and every outcome-alignment label with "aligned". The displayed fixed-label ranges are identification bounds, not confidence intervals. Panel D reports the main leading-model comparison and the prespecified sensitivity that excludes leading-model world cells with any detected follow-up MRI read; the latter leaves 37 cells across 17 worlds. MRI, magnetic resonance imaging.

Supplementary Methods

S1. World generation

Each benchmark world is generated from a structural causal model (SCM) defining the joint distribution of participant characteristics, genetic variation, molecular measurements, latent disease state, observed phenotypes, longitudinal outcomes, and responses to intervention. A procedural seed determines the world level parameter set (Θ★), including the causal driver set ( 𝒟 ), driver effect sizes (wj), disease archetype (a), genetic instrument strengths, identities and parameters of the planted biases, measurement properties, and intervention effects. These parameters are sampled once and held fixed for all agents evaluated within the same world. Participant level observations are then sampled conditionally from the resulting SCM.

Unless otherwise stated, (i) indexes participants, (j) indexes molecular features, (k) indexes observed phenotypes, and z(·) denotes standardization to zero mean and unit variance within the generated population.

S1.1 Population and molecular generation

Calibration provenance and implementation. The 20-world panel was generated at revision 75e3a4bc55373a91df0bb04d11bfb28e405407c3 using the mesa-topmed-ukbscale profile. Its demographic, MRI, genetic and outcome target blocks are identical to the mesa-topmed profile; only generator dimensions and associated metadata differ. The evaluated profile specifies 54,000 participants, 10,800 imaged participants (20%), 20% second-visit sampling, 2,941 proteins and 8,192 variants. These dimensions are benchmark design choices inspired by UK Biobank scale, not estimates fitted to UK Biobank participant data.

The calibration JSON attributes its aggregate summaries to TOPMed Freeze 10b WGS-linked MESA. Age uses the recorded mean, standard deviation and truncation bounds; sex uses the male proportion. BMI uses the empirical center with designed age, sex and group effects and residual standard deviation 5.05. Selected MRI morphology variables are generated as empirical mean plus empirical standard deviation times a latent standardized value, with selected LV/RV correlations imposed; derived volumes, ejection fraction and physiological clipping can alter exact moments. Genotypes use synthetic frequencies anchored to common-variant summaries and a Gaussian-copula haplotype model. Protein missingness uses the recorded aggregate rate (0.004146). The cardiovascular-event proportion supplies a prevalence-scale anchor rather than a validated match to the synthetic five-year endpoint. Causal graphs, molecular effects, disease loadings and planted biases remain specified by design; the full joint biological distribution is not learned from MESA.

Empirical extraction. Calibration summaries were computed in a controlled MESA/TOPMed environment from 3,054 WGS-linked Exam 1 participants after one-to-one linkage of the clinical and ancillary RV tables. BMI was derived from measured height and weight; cardiac summaries excluded values outside prespecified broad physiologic bounds. Common-variant frequencies and adjacent-variant LD were summarized from the 8,192-variant TOPMed-derived panel after participant deduplication, reversal of stored dosage normalization, and sorting by chromosome and position. A scaled-beta allele-frequency distribution was fitted by least squares to five empirical quantiles. Protein missingness and the cardiovascular-event prevalence anchor came from separately recorded aggregate preprocessing summaries.

Provenance verification. The original source commands and recorded aggregate outputs were recovered from the controlled-environment analysis. Eighty-five checks against the committed calibration profile passed at its rounding precision, including means, standard deviations, counts and reference covariate R² values for all 13 cardiac traits. This audit verifies agreement with the recorded outputs; an independent rerun against the controlled participant data was not performed. The fixed 110-kb LD-decay length and genomic-spacing mixture were empirically guided tuning choices. The repository contains the aggregate calibration JSON and generator implementation, but does not package the complete external preprocessing environment. MESA anchors selected clinical-unit MRI statistics; ACDC supplies anatomical rendering templates. Calibration to these targets does not establish independent real-cohort validity.

Participant age is generated from a truncated normal distribution,

Ai ∼ TN(59.147, 9.245²; 45, 84),

and sex is generated as a Bernoulli variable,

Si ∼ Bernoulli(0.4804).

Body mass index is generated conditionally on age, sex, participant level background variation, and residual noise:

BMIi = clip[28.3006 + 0.05(Ai − 59.147) + 0.45(Si − 0.4804) + ri + εi, 15, 55].

Exercise is generated from a right skewed baseline distribution and allowed to depend modestly on BMI:

Ei = clip[Γ(2,1.8) − 0.06(BMIi − 27), 0, 20].

Here, (ri) represents participant level background component, and (εi) represents independent residual variation. The coding of (Si = 1) and the scale/rate convention used for the Gamma distribution are specified in the implementation.

Genetic variation and linkage disequilibrium

For genetic variant (j), the allele frequency parameter (qj) is sampled as

0.034381 + (0.5 − 0.034381) Beta(0.734507, 1.064551).

Local linkage disequilibrium is introduced by correlating latent genotype variables according to physical distance:

exp(−|posj − posk | / 110 kb).

The initial world panel contains 8,192 variants distributed across 22 chromosomes.

Molecular measurements are generated from cis-genetic effects, selected trans-genetic effects, shared latent biological factors, participant covariates, and assay-specific residual variation. For protein (p),

z[√(hp ²) Gi,cis(p) + Ip √0.01 Gi,master(p) + FiTλp + 0.01(Ai − 55) + 0.10Si + 0.02(BMIi − 27) + 0.8εip].

Here, (hp ²) determines the strength of the protein-specific cis-genetic component; (Gi,cis(p)) is the genotype at the cis instrument assigned to protein (p); (Ip) indicates whether protein (p) receives a trans effect; (Gi,master(p)) is the corresponding trans or master-regulator genotype; (Fi) is a vector of shared latent biological factors; (λp) is the protein-specific vector of factor loadings; and (εip) is assay-specific residual variation.

The initial panel contains 2,941 proteins, with one cis instrument assigned per protein, 12 shared latent factors, and approximately 20% of proteins receiving a trans-genetic effect.

Transcript measurements are generated using the same general structure,

z[αgGi,cis(g) + βgGi,trans(g) + FiTλgRNA + ηgTCi + εigRNA],

with transcript-specific genetic effects, latent-factor loadings, covariate effects, and residual variation. Transcript and protein levels can therefore share genetic and biological influences without requiring a fixed transcript-to-protein causal relationship.

S1.2 Latent disease generation

For participant (i), latent disease severity (Li) is determined by a hidden subset of molecular drivers ( 𝒟 ), measured participant risk factors, optional direct genetic effects, and residual variation:

z[Σj∈𝒟 wjPij + 0.15z(BMIi) + 0.10 Smokingi + 0.12Gi,direct + 0.08z(Ai) + εi].

Equivalently, the model can be written compactly as

z[Σj∈𝒟 wjPij + γTCi + δGi,direct + εi],

where (Ci) denotes the vector of measured participant covariates. The set ( 𝒟 ) is hidden from the scientific agent. Across worlds, the number of causal molecular drivers ranges from zero to six. Driver identity, direction and magnitude of effect, genetic instrument strength, direct genetic effects, and residual disease variance vary across worlds.

Hard-world disease architecture

Hard worlds introduce nonlinear molecular effects, interactions between the strongest drivers, additional polygenic background, and greater residual variation:

z[Σj∈𝒟 wj f(Pij) + 1/2 min(|w1 |, |w2 |)Pi1Pi2 + ciTβ + Gipoly + εi],

where

f(P) = 1.6 tanh(P/1.6).

The nonlinear transformation causes molecular effects to saturate at large absolute protein values. The product term introduces epistasis-like interaction between the two strongest molecular drivers, and (Gipoly) represents the aggregate contribution of many individually weak genetic effects.

S1.3 Phenotype and longitudinal outcome generation

Latent disease severity is not released directly to the scientific agent. Instead, it generates observable phenotypes whose expression depends on the disease archetype:

αa,kLi + βkBi + γkEi + δkPi,T8 + εik, a ∈ {DCM, HCM, ischemic, HFpEF}.

Here, (Yik) is phenotype (k) in participant (i); (Li) is latent disease severity; (Bi) denotes background/body-size covariate used in implementation; (Ei) is exercise; (Pi,T8) is the designated harmful-surrogate feature when that bias is active; and (εik) is phenotype-specific residual variation.

The archetype-specific loading (αa,k) determines how the same latent disease severity is expressed across phenotypes. In the initial cardiac worlds, the principal phenotype loadings are:

EDV ESV Wall Motion Diastolic lag Native T1

DCM 0.42 0.70 −0.28 −0.18 0.28 0.15

HCM −0.40 −0.12 0.55 0.05 0.42 0.42

Ischemic 0.22 0.42 −0.06 −0.46 0.18 0.35

HFpEF 0.05 0.06 0.30 −0.02 0.30 0.95

Thus, DCM is expressed primarily through chamber dilation and increased end-systolic volume, HCM through increased wall thickness and smaller cavity size, ischemic disease through impaired regional motion together with coronary phenotypes, and HFpEF through preserved cavity measures with stronger tissue and diastolic signals.

The released phenotype set includes medical imaging, physiologic signals, clinical measurements, EHR diagnoses, repeated visits, and survival outcomes. Dynamic modalities are generated on their natural time scales: ECG features vary within milliseconds, cine imaging varies across a cardiac cycle, and repeated clinical measurements vary across longitudinal visits.

For longitudinal worlds with delayed molecular effects,

wbase ∈ [0.04,0.10], wlate,

and later disease severity is generated as

0.75Li,1 + 0.35 Pi,𝒟T wlate.

This produces causal drivers whose short-term effects are weak but whose effects become apparent over longer follow-up.

S1.4 Planted biases

Each world may contain one or more of the predefined planted biases. These modify the structural or measurement equations rather than merely increasing random noise.

T1. Confounding

A noncausal protein is generated partly from BMI,

z[0.55z(BMIi) + √h² Gicis + 0.8εi],

while BMI also contributes directly to disease severity,

Li ← Li + 0.15z(BMIi).

This creates the backdoor structure

Piconf ← BMIi → Li,

causing the protein to associate with disease despite not lying on the causal molecular pathway.

T2. Reverse causation

A reactive protein is generated downstream of disease:

z[0.50Li + √h² Gicis + 0.75εi].

The resulting association between protein abundance and disease therefore arises from

Li → Pirev,

rather than from a causal effect of the protein on disease.

T3. Imaging selection

A latent imaging-selection score is generated as

−0.24Aiz − 0.52BMIiz + ri + si + 0.55ηLi + 0.55ηPisel + 1.35εi,

and imaging is observed only among participants in the upper tail of that score:

1{Si * in the top 20%}.

Because both disease severity and the selected protein can influence imaging availability, conditioning on the imaged population can induce collider bias (Munafò et al., 2018) .

T4. Shared genetic instrument / causal non-identifiability

A shared genetic variant influences both a disease-relevant molecular feature and a noncausal molecular feature:

Gi * → PiD, Gi * → PiN.

This produces genetically correlated molecular candidates for which association alone does not identify the disease-mediating feature.

T5. Imaging site batch effect

Observed image intensity differs from the underlying image through a site-specific spatial distortion:

Iitrue(u){1 + bsite(i)(u)} + εi(u),

where (u) indexes image position and (bsite(i)(u)) is a site-specific multiplicative field. This allows acquisition site to become spuriously predictive when site is associated with disease prevalence or participant composition (Fortin et al., 2018) .

T6. Benign exercise remodeling

Exercise modifies cardiac morphology without changing latent disease severity:

z(EDVi) ← z(EDVi) + 0.22Ei,

z(ESVi) ← z(ESVi) − 0.12Ei,

with

Ei ↛ Li.

The bias therefore produces genuine structural remodeling that should not be interpreted as progression of the latent disease process.

T7. Pleiotropic genetic instrument

A cis-genetic instrument affects both the candidate protein and latent disease through separate pathways:

Gicis → Pipleio,

Gicis → Li(0.14).

The direct genetic path violates the exclusion restriction required for a valid instrumental-variable interpretation (Sanderson et al., 2022) .

T8. Harmful surrogate

A designated molecular feature contributes directly to cardiovascular event risk in addition to latent disease severity and conventional risk factors:

0.006 exp(0.70Li + 0.45Aiz + 0.35Smokingi + 0.15BMIiz + 0.40Pi,T8).

Thus, (Pi,T8) can be strongly predictive of observed cardiovascular outcomes without being a valid surrogate for the latent disease process or a member of the true causal-driver set ( 𝒟 ).

T9. Assay unit mixture

At one measurement site, the observed molecular assay uses a different affine scale:

2.2Pi + 1.1.

The underlying biology is unchanged, but the numerical measurement distribution differs systematically across sites (Leek et al., 2010) .

S1.5 Intervention and longitudinal model

Both actions invoke the same sealed intervention oracle. For the nominated protein, the intervention replaces its standardized abundance with −1.645 (approximately the fifth percentile of a standard normal distribution); the paired control uses a standard-normal draw. Other driver values, latent residuals and morphology residuals are shared between the control and intervention. The cell action uses the baseline weight vector and n = 120; knockdown uses the late weight vector and n = 800. Ordinary drivers have equal weights at both horizons. For an A9 delayed-effect driver, the late coefficient has a randomly selected sign and magnitude 1.15–1.45 times the largest pre-adjustment driver magnitude, and its baseline coefficient is 4–10% of that late coefficient. The baseline effect is therefore small, not identically zero.

The oracle computes latent severity from weighted driver values, using 1.6 tanh(x/1.6) saturation and a two-driver interaction in hard worlds, plus a shared Gaussian residual (SD 0.70 in hard worlds and 0.55 otherwise). Let e be a shared Gaussian morphology residual with SD 0.35. For a query on the T8 surrogate, b is its control abundance or the clamped intervention value; b = 0 for queries on other proteins. The generic readouts are cavity radius = 10 + 1.9L − 0.45b + e; wall thickness = 4.6 − 0.55L + 0.12b + 0.5e; EF = clip(0.62 − 0.09L + 0.025b + 0.08e, 0.20, 0.78); and five-year survival = exp[−5 × 0.006 × exp(0.70L + 0.40b)]. Each returned delta is the cohort mean of intervention minus control, rounded to four decimal places. These are fixed generic equations, not the full disease-specific image-generation model.

Knockdown adds no assay noise. Each cell-perturbation delta receives independent zero-mean Gaussian noise with SD 0.45|delta| + 0.0225 and is rounded again to four decimals. Both delivered payloads expose action, status, protein and a four-field delta dictionary (cavityr, wallt, ef, surv5y); control-arm summaries computed internally by the oracle are not passed through to the agent. The oracle's cohort draws are seeded by world seed + 777 + protein index, so knockdown output is deterministic for a fixed world, protein, sample size and horizon. Costs and simulated latencies are $150,000/90 days for cell perturbation and $400,000/180 days for knockdown. Duplicate purchases of the same action–protein pair are refused because the original result remains available.

The historical prompt describes cell perturbations as providing no survival information. The implementation returned surv5y in all 36 delivered cell payloads, a version 1 prompt–output mismatch retained in the archived evaluation.

S2. Scoring rules in full

S2.1 Claim intake

Nominations are ordered by submitted rank (finite ranks first, then submission order), one entry per protein (first occurrence wins), and only the first 25 are processed; every unique valid nomination enters the precision denominator, the safety check and the pair logic. Invalid protein identifiers are dropped from the nomination list, invalid or missing directions and alignment labels default to "unknown" (so a nominated driver with no direction earns 0.5 and the safety penalty fires only on an explicit "aligned"), and an invalid rejection mechanism is blanked but keeps its place in the rejection denominator; confidence is clipped to [0, 1]. Rejections keep one entry per protein. Abstention entries are sets of proteins: the harness removes entries with fewer than two valid identifiers, and the scorer ignores entirely (neither scored nor counted) entries with more than five valid identifiers, entries with no valid identifier and duplicate sets. Only the first 1,000 retained entries are scored, but every unique retained entry (two to five valid identifiers) counts in the abstention-precision denominator, including those past the scored cap.

S2.2 Pair logic and false positives

Every nominated protein that is not a hidden driver is a false positive, including the T8 surrogate and every planted non-causal protein, with one exception: for each planted T4 or A9 pair (causal member, non-causal partner), if the non-causal partner is among the first 25 nominations and the causal member is named nowhere, the non-causal partner replaces the causal member in the effective truth used for recall and precision, so the agent is credited for finding the locus without being penalised twice. If the pair is abstained, either through an abstention entry or by nominating both members within the first 25, both proteins are removed from the effective truth, from the precision denominator and (the causal member) from the direction denominator. A nomination past rank 25 can trigger none of the three routes (the swap, the 17.5-point claims route and the 25-point experiment route). Abstention precision is the number of scored abstention entries that contain a planted pair, after removing proteins the agent also nominated, divided by the number of retained abstention entries. An abstention on a pair that was not planted therefore earns nothing and lowers abstention precision; abstentions have no effect in pair-free worlds. A rejection's mechanism maps to the bias that planted the non-causal protein as confounding → T1, reverse_causation → T2, selection → T3, pleiotropy → T7 and measurement_artifact → T9; T4, A9, T5, T6 and T8 plant no scored non-causal protein.

S2.3 Null world and pair-free worlds

With no drivers, recall and direction are undefined; the rules are: any nomination outside an abstained pair → Starget = 0; no nomination and no rejection → 15; no nomination with rejections → 30 × (correct rejections / rejections); Sconfidence = 0 and Sdirection = 0. An absent submission is scored as an empty one, which is why a turn-limited episode with no submission can score 15 in the null world. In the two other worlds without a T4 or A9 pair (high-concordance-01 and heldout-hcm-02), Sconfidence = 25 × Starget /30. In every non-null world the 0.5 credit for an "unknown" direction equals the expectation of a coin flip, so hedging is neither rewarded nor punished relative to guessing.

S2.4 Exit paths

The harness scores every episode on exit, whether the agent committed, hit the turn limit, or the transport failed, from whatever the working directory holds: an absent or unparseable submission is scored as an empty one, and the phenotype file is scored independently of any submission. A parseable submission ends the episode on the turn it is written, so a turn-limit episode never holds one. The prompt's warning that a turn-limited episode "scores nothing" is therefore not what the scorer does; in run 1, three episodes with no submission earned phenotype credit, and six of the eight null-world episodes that scored the 15-point restraint rule had no submission at exit. A runner crash yields a zero on every component (none occurred). The 25-nomination cap engaged in 8 episodes and the safety penalty in 22.

S2.5 Phenotype validity and edge cases

The scorer aligns the first submitted data column to subject IDs and computes r = |corr(phenotype, L)| on finite entries selected by sealed eval_mask.npy. It requires at least 51 finite values and nonzero submitted variance; otherwise credit is zero. If sealed/phenotype_baseline.json contains a finite built_in_correlation b, the score is 10 × clip((r − b)/(0.85 − b), 0, 1), except that a gap 0.85 − b ≤ 10−6 gives ten points exactly when r > b and zero otherwise. With no valid recorded baseline, it uses 10 × min(r/0.85, 1). The phenotype must reside in the approved submission directory. These branches were checked in the byte-identical score.py snapshots at 85c1ab70 and ced47324. Absolute correlation is invariant to sign and nonzero affine rescaling, so the scorer does not require one numerically identical vector. It nevertheless designates one latent construct.

A finite baseline was recorded for all 20 worlds (Table S3), so the baseline-relative branch applied in every world. Baselines cluster by archetype: 0.40–0.44 in the six DCM worlds, 0.67–0.68 in the six HCM worlds, 0.71–0.72 in the two ischemic worlds and 0.75–0.76 in the six HFpEF worlds. Because credit accrues only above the baseline, a submitted phenotype in an HFpEF world must exceed its world’s built-in baseline correlation (0.747–0.765 across HFpEF worlds) with latent severity before earning any points. In the rescoring audit of 21 September 2026, 392 phenotype files were submitted; 349 were scorable, including 340 that were accepted and scorable and 9 rejected for duplicate IDs.

The frozen plan designates 12 development and eight held-out worlds; the literal IDs are retained, but these labels do not establish that evaluation worlds were untouched during development. The generator draws each evaluation mask as independent Bernoulli(0.2) indicators on the named release_missingness random-number stream. Archived ground-truth summaries report 10,703–10,990 selected participants per 54,000-person world; recomputing only this mask-generation step from the frozen seeds reproduced all 20 counts. The built-in baseline and submitted phenotype are evaluated on this same mask, not on separate baseline-fitting and scoring subsets.

S2.6 Sensitivity to phenotype scoring

This post hoc analysis removes phenotype points from each episode, sums the other four components and signed safety penalty, and clips the result to [0, 90] without rescaling. The original 100-point score remains the primary outcome. Simply subtracting phenotype credit from an already-clipped total is incorrect: it produces seven negative scores. Recomputed component totals agree with recorded full scores within the exported 0.01-point precision.

All nine overall mean-score ranks remain unchanged. Original means /100 → no-phenotype means /90 are: Opus 5, 39.98 → 33.24; GPT-5.6 Sol, 35.38 → 32.33; Sonnet 5, 21.33 → 20.24; Haiku 4.5, 12.92 → 12.71; GPT-OSS-20B, 7.41 → 6.76; Qwen3-Coder-30B-A3B, 5.90 → 5.76; GLM-4-32B, 1.46 → 1.38; Qwen3-8B, 1.31 → 0.92; Devstral Small, 0.81 → 0.75.

Averaging budgets within each of the 20 worlds, the Opus–GPT difference falls from 4.60 to 0.91 points (paired-t 95% CI −3.64 to 5.46; P = 0.680), with Opus ahead in 10 rather than 16 worlds. A 20,000-resample paired-world percentile-bootstrap interval is −3.08 to 5.23. At $450,000 the leading point order switches (GPT 31.54, Opus 30.95), with wide uncertainty. Zero-score episodes increase from 282/540 to 298/540; among open-weight models they increase from 234/300 to 247/300. This is a scoring sensitivity, not a rerun or a validation of the latent phenotype; phenotype choices may still influence target nominations.

S2.7 Statistical software

The paired-t p value, the t critical value for the 95% confidence interval and the Wilcoxon signed-rank p value (exact mode, zero differences excluded) were reproduced on 17 September 2026 with SciPy 1.14.1; the sign test is an exact two-sided binomial computation with tied worlds excluded, applied to model pairs and to each model's full-program-minus-observational contrast; bootstrap intervals are percentile intervals over 20,000 resamples of the 20 world differences.

S2.8 Fixed-label sensitivity

We performed a post hoc label ablation holding each episode’s scored nominees and order, experiments, abstentions, rejections and phenotype score fixed, replacing every direction with “inhibit” and every outcome-alignment label with “aligned”. Causal-confidence credit is invariant to these labels under rubric v0.9. Original sealed direction/pair manifests were unavailable. We therefore enumerated all binary driver-direction assignments (at most 26 per world), retaining those compatible with all 27 archived direction scores and a conservative superset of denominators permitted by the T4/A9 rules. Counterfactual minima and maxima give identification bounds, not confidence intervals; 529/540 episode direction scores were point-identified. T8 identities made fixed-alignment penalties exact. Totals recombine exported components and reapply clipping; original-total reconstruction error was at most 0.01 point.

For Opus 5, mean direction credit decreased from 13.05 to 5.82–6.48, causal confidence remained 4.26 and total score decreased from 39.98 to 16.19–16.86. For GPT-5.6 Sol, direction decreased from 13.00 to 5.77–6.43, causal confidence remained 4.19 and total decreased from 35.38 to 21.57–22.24. Always-aligned increased penalized episodes from 22 to 89 overall (Opus 5: 6 to 44; GPT-5.6 Sol: 5 to 25). These comparisons retain the original agents’ target-selection work and are not scores for a separate discovery policy.

S3. Agent prompt and harness protocol

S3.1 System prompt and opening message

The system prompt below is reproduced verbatim from the harness (code_loop.py, revision 85c1ab70, identical in ced47324), with the per-episode values (turn cap, timeout, prices, budget and submission path) filled in as the agent saw them; the budget line shows the three possible values. The opening user message follows it. The agent is addressed as a cardiovascular researcher; the names RELEASE, IMAGING, WORK and OUT_CSV are bound for it; and it is warned that an episode reaching the turn limit without a submission "scores nothing" (in fact the harness scores whatever the working directory holds, Supplementary Methods S2.4).

You are a cardiovascular researcher analysing a population biobank.
Work by writing Python. Each turn you send exactly one Python code block; the
harness runs it and shows you what it printed. Your working directory persists
between turns, so anything you write is still there next turn.
RELEASE, IMAGING and WORK are variables already holding real paths when your
code runs. Use them as names. Do not quote them and do not write a path of
your own:
cov = pd.read_parquet(os.path.join(RELEASE, "covariates.parquet"))
scans = glob.glob(os.path.join(IMAGING, "*.npz"))
np.save(os.path.join(WORK, "features.npy"), features)
np, pd, os, sys, glob, gzip, json and subprocess are imported. Also
available: scipy, statsmodels, scikit-learn, scikit-image, OpenCV and PyTorch.
Work on CPU. Each turn's code has 300s to run, and you have 30
turns.
`load_genotypes()` returns the dosage matrix as a DataFrame, subjects by
variant, values 0, 1 or 2. Use it rather than parsing the VCF yourself.
To commission laboratory work, call `request_experiment(kind, protein)`. The
result appears next turn at WORK/experiments/<protein>.<kind>.json.
## What is in RELEASE
Tables are indexed by subject id rather than carrying it as a column.
`data_dictionary.tsv` maps each covariate column to its description and any
documented sentinel codes; `data_dictionary.md` is a prose note on how the
assays and their missingness were produced.
covariates.parquet age, sex, BMI, smoking, exercise, assessment
centre, imaged flag -- named by UK Biobank field
ID, so resolve them via data_dictionary.tsv
proteomics.parquet plasma proteins, standardised, sparse missingness
transcriptomics.parquet matched blood mRNA (TRANS_xxxx pairs PROT_xxxx)
metabolomics.parquet plasma metabolites
ecg_features.parquet hr, pr, qrs, qt, lvh_voltage_mv, qwave, lbbb
coronary_ct.parquet cac_agatston, per-vessel and max stenosis percent
ehr_diagnoses.parquet ICD-10 codes with dates
ehr_medications.parquet ATC codes
mortality.parquet death and cause
mace_events.parquet event and years from baseline
targetability.parquet per-protein druggability, indexed by protein
genotypes.vcf.gz genetic variants; read via load_genotypes()
imaging/SUBJ_XXXXX.npz cine uint8 (slices, frames, 128, 128) and
t1map float32 (slices, 128, 128) in milliseconds,
NaN outside the myocardium; filenames are ids
imaging_visit2/ repeat imaging for a subset
covariates_visit2.parquet repeat assessment about five years later
proteomics_visit2.parquet
You see only what your code prints, so print what you need to reason about.
## The task
A small number of proteins causally drive cardiac disease in this cohort.
Others look associated because of confounding, reverse causation, selection
into the imaged sub-cohort, pleiotropy, or a measurement artifact. There is no
phenotype column anywhere in the data.
Identify the causal drivers, and say which way a drug should move each.
## Deriving a phenotype
Deriving a quantitative disease measure from the cardiac MRI sharpens every
analysis that follows, and it is scored in its own right. `OUT_CSV` is bound to
the path to write it to:
pd.DataFrame({"subject_id": ids, "phenotype": values}).to_csv(
OUT_CSV, index=False)
Nothing derives this for you, and there is no built-in tool to call for it.
You are scored against a fixed, held-out baseline correlation you never see:
a phenotype at least as good as that baseline earns nothing extra, and one
worse than it costs nothing further, but only a phenotype that beats it earns
credit -- so attempting your own is worth it even though the floor is invisible
to you. Constructing this well is itself part of what is being measured, not
a preliminary step to skip.
## Experiments
cell_perturbation $150,000 in vitro, read at baseline; a noisy read on
the molecular phenotype, no survival
information, and an effect that only
accumulates over years will not show up
knockdown $400,000 in vivo, read over the full five years; a
clean read plus the survival response
You have $[budget for the episode: 0, 450,000 or 2,000,000]. Observational genetics can suggest a direction, but it is
wrong under pleiotropy, reverse causation, and for a marker that improves the
measurement while harming the patient. Only an experiment settles those.
## Finishing
Write WORK/submission.json and the episode ends. An episode that reaches the turn
limit without one scores nothing, so write a submission as soon as you can
defend it and overwrite it later if your evidence changes. Schema:
{"nominated": ["PROT_0123", ...],
"directions": {"PROT_0123": "inhibit"},
"outcome_alignment": {"PROT_0123": "aligned"},
"rejected": [{"molecule": "PROT_0456", "mechanism": "confounding",
"reason": "..."}],
"abstentions": [{"molecules": ["PROT_0789", "PROT_0790"],
"reason": "..."}],
"confidence": 0.7}
directions: inhibit / activate / unknown.
outcome_alignment: would improving the phenotype through this molecule also
improve survival? aligned / misaligned / unknown. Claiming "aligned" about a
molecule that is actually misaligned is the costliest error here.
mechanism: confounding / reverse_causation / selection / pleiotropy /
measurement_artifact. A rejection scores only when the mechanism is right, and
the component is your precision over everything you rejected times your recall
of the real decoys, so rejecting molecules you are not confident about lowers
your score.
abstentions: molecules you cannot separate from one another. Correct abstention scores; guessing between them does not, and abstention carries the same precision term, so entries you cannot defend lower your score.
Nominate only what your evidence supports. Extra names cost you, so do missing ones.
Opening message (isolated sandbox paths):
Begin.
When your code runs these names are already bound:
RELEASE = '/release'
IMAGING = '/release/imaging'
WORK = '/work'
They are variables, not strings to quote.
Reply with exactly one Python code block and nothing else.

S3.2 Execution protocol

The sandbox provides NumPy, pandas, SciPy, statsmodels, scikit-learn, scikit-image, OpenCV and PyTorch on CPU. The last fenced code block of a reply is executed; a reply with no fence but containing print( or import is executed whole; a reply with no code consumes a turn and is answered with a formatting hint. The full reply, the extracted code, each observation and every compaction summary are recorded in the episode transcript; in the conversation history returned to the model, only the extracted code stands in for the assistant turn, so the model's own prose is not carried forward. Each observation returns "Your code ran/failed", stdout, stderr, delivered or refused experiment notices, the working-directory listing, and the remaining budget and turns. Experiments requested in the committing turn are still fulfilled and charged. Output was capped at 32,000 tokens per turn (8,000 for Qwen3-8B and GLM-4-32B). When estimated prompt tokens exceed 0.7 × the context window (128,000 unless configured otherwise), older turns are summarised by the same model in at most 1,024 tokens under a separate prose-only instruction and at most the last six messages are kept verbatim (62 of 540 episodes compacted; 48 once, 9 twice, 5 more often). Truncated completions are counted, billed and skipped (5 episodes). Temperature 0.0 was requested for every arm and no reasoning-effort setting was passed; the transport sent the temperature parameter only for Haiku 4.5, so the other eight models ran at their provider or server default. A deterministic seed derived from world, arm, budget and replicate was passed on the OpenAI path only. The sandbox was isolated in 536 of 540 episodes (four Devstral Small episodes ran non-isolated). Archive labels for the budgets are observational ($0), single_experiment ($450,000) and full_program ($2,000,000).

Supplementary Note S1. Turn-by-turn trace of a human-guided DrugTargetWorld episode.

To provide a more detailed reference for the human-guided evaluation, we report below the complete research trajectory of the highest-scoring human-guided episode. The episode was conducted in the full-program condition with a $2,000,000 experimental budget. The investigator interacted iteratively with the computational environment, using a language model to write and execute analyses while directing the research strategy and deciding which questions to pursue next.

The trace below preserves the analysis submitted at each turn and the principal output returned by the environment. Large mechanically generated outputs, including complete 2,941-protein tables, 8,192-variant tables, full file manifests, and repeated file listings, are truncated for readability. No hidden causal information or benchmark feedback was available during the episode.

Initial task presented to the investigator

The episode began with a synthetic cardiovascular population biobank containing 54,000 participants. Available data included covariates, proteomics, transcriptomics, metabolomics, ECG features, coronary CT, diagnoses, medications, mortality, MACE events, targetability information, 8,192 genetic variants, cardiac MRI for 10,800 participants, and repeat imaging and molecular measurements for subsets of participants.

The investigator was told:

A small number of proteins causally drive cardiac disease in this cohort. Others look associated because of confounding, reverse causation, selection into the imaged sub-cohort, pleiotropy, or a measurement artifact. There is no phenotype column anywhere in the data.

Identify the causal drivers, and say which way a drug should move each.

The investigator was required to construct a participant-level phenotype and could purchase either cell perturbation experiments for $150,000 or in vivo knockdowns for $400,000. The latter provided five-year intervention effects including survival.


Turn 1 — Map the available data

Submitted analysis

“Map the dataset and inspect exactly what signals are available. Keep this diagnostic-heavy; later turns will use the printed structure to build the MRI phenotype, causal screen, decoy tests, and experiment allocation.”

The analysis loaded the data dictionary, enumerated all released tables, quantified missingness, examined subject overlap, inspected protein/transcript pairing, loaded the full genotype matrix, and inspected several raw MRI files.

Selected returned output

 === DATA DICTIONARY TSV ===
 shape: (6, 3)

 field_id description coding
 p31_i0 Sex 0=female, 1=male
 p21003_i0 Age when attended assessment centre years
 p21001_i0 Body mass index (BMI) kg/m2
 p20116_i0 Smoking status 0=never,1=current
 p22040_i0 Physical activity coded
 p54_i0 Assessment centre coded
 covariates.parquet: (54000, 10)
 proteomics.parquet: (54000, 2941)
 transcriptomics.parquet: 2941 transcripts
 genotypes: (54000, 8192)
 MRI scans: 10800
 visit-2 MRI scans: 2142

 proteomics ∩ transcriptomics: 54000

Raw imaging consisted of three cine slices across 25 frames and three T1 maps per participant.

The initial pass also established that molecular missingness was low and that all 54,000 participants had genotype data.

Decision carried forward The investigator concluded that the environment supported three complementary routes to causal discovery:

1. construct a disease phenotype from MRI;

2. screen molecular measurements against clinical and longitudinal disease outcomes; and

3. use genetics and later interventions to distinguish causal proteins from non-causal proteins with misleading observational associations.

No experimental budget was spent.


Turn 2 — Build reusable MRI features and inspect clinical outcomes

Submitted analysis

“Inspect the disease/outcome tables cleanly and extract compact MRI features. We are NOT finalizing the phenotype yet. We want a reusable imaging feature matrix that can be evaluated against longitudinal outcomes and visit-2 progression.”

The analysis inspected ECG, coronary CT, diagnoses, medications, mortality, MACE, repeat covariates, and repeat proteomics. It then extracted quantitative measurements from the cine and T1 images for all available MRI subjects.

Features included measures related to:

  • T1 distribution;
  • cine intensity and texture;
  • temporal variation;
  • motion;
  • cavity geometry;
  • myocardial geometry; and
  • slice-to-slice variability.

The same extraction was performed for repeat MRI.

Decision carried forward

MRI feature extraction was treated as an intermediate representation rather than the phenotype itself. The next step was to learn which imaging measurements tracked independently observed cardiovascular disease.


Turn 3 — Construct an outcome-anchored disease phenotype

Submitted analysis

“Construct an outcome-anchored MRI phenotype and screen proteins against both cross-sectional cardiomyopathy/HF and future cardiac outcomes.”

Clinical labels were constructed from cardiomyopathy and heart-failure diagnoses and longitudinal outcomes. MRI features were used to predict the clinical disease phenotype using cross-validation rather than assigning manual imaging weights.

The resulting out-of-fold predictions were rank transformed to generate a continuous MRI-derived disease phenotype.

The same turn screened all 2,941 proteins against:

  • the MRI phenotype;
  • prevalent cardiomyopathy/heart failure;
  • incident cardiomyopathy/heart failure;
  • cardiac mortality; and
  • atrial fibrillation.

Selected returned output

The strongest observational protein associations were:

        r_mri r_cmhf r_incident_cmhf r_cardiac_death obs_score
 PROT_1858 0.3963 0.2335 0.1861 0.1353 0.9321
 PROT_1606 -0.3311 -0.2050 -0.1649 -0.1179 0.8011
 PROT_0700 -0.3181 -0.1867 -0.1375 -0.1094 0.7356
 PROT_1691 -0.1753 -0.1069 -0.0817 -0.0588 0.4142
 PROT_1779 -0.1571 -0.1020 -0.0813 -0.0591 0.3902
 ...

PROT_1858 therefore appeared to be the strongest observational candidate at this stage, followed by PROT_1606 and PROT_0700.

Decision carried forward

The investigator did not nominate these proteins directly. The next analysis was explicitly designed to ask whether these observational associations were genetically supported.


Turn 4 — Genetic triangulation

Submitted analysis

The investigator requested four analyses:

1. GWAS the derived MRI phenotype across all 8,192 variants.

2. For the strongest phenotype variants, test association with all proteins and transcripts.

3. For the strongest observational proteins, identify their strongest pQTL and ask whether the same variant predicts MRI disease.

4. Compare protein-versus-transcript concordance to detect measurement artifacts or downstream plasma markers.

Selected returned output

PROT_1606 and PROT_0700 emerged as the clearest genetically anchored candidates:

 protein best_variant r_gp variant_r_mri r_protein_transcript
 PROT_1606 rs104994 0.43931 -0.16337 0.37861
 PROT_0700 rs102986 0.38643 -0.11701 0.53612

Their genetic effects and observational disease associations pointed in concordant directions.

PROT_1858, despite having the strongest observational association, did not show the same clean matched-transcript/genetic pattern.

Decision carried forward

PROT_1606 and PROT_0700 became the leading causal candidates. PROT_1858 became a high-priority ambiguity because its strong observational association conflicted with weaker molecular-genetic support.


Turn 5 — Expand genetically anchored discovery

Submitted analysis

“Identify ALL proteins with strong variant → protein, same variant → MRI phenotype, preferably same variant → matched transcript, and then see whether the genetically implied direction agrees with actual cardiac outcomes in the full cohort.”

Rather than limit subsequent work to the initial observational shortlist, the analysis expanded from phenotype-associated variants back into the full proteome.

Shared genetic loci involving multiple proteins were explicitly identified.

Decision carried forward

The investigator recognized that a genetic locus could implicate multiple proteins and that choosing the protein with the strongest pQTL would not necessarily identify the causal molecule. Shared loci therefore became a separate problem requiring conditional analysis.


Turn 6 — Broad longitudinal and source-of-bias analysis

Submitted analysis

“Do NOT miss weaker drivers. Distinguish reverse causation, selection, pleiotropy, and measurement artifact. Exploit visit-2 proteomics + visit-2 imaging before spending experiment budget.”

A broad candidate universe was assembled from:

  • top observational associations;
  • genetically anchored candidates; and
  • suspicious or discordant proteins.

The analysis attempted to calculate longitudinal protein behavior, repeat MRI effects, selection associations, and genetic support.

Returned output

The analysis encountered a data-type error while attaching genetic results to the longitudinal diagnostic table.

 pandas.errors.LossySetitemError
 ...
 diag.loc[p, c] = gcand.loc[p, c]

Decision carried forward

Rather than abandon the analysis, the investigator repeated it with safer data handling.


Turn 7 — Repair and rerun the longitudinal screen

Submitted analysis

“Rerun Turn 6 with dtype-safe genetics attachment. Same observational objective, but save diagnostics earlier and print only the rankings we actually need for experiment allocation.”

Returned output

The longitudinal analysis completed and produced candidate-level measures including:

  • repeat-protein correlation;
  • baseline protein versus future disease;
  • prior disease versus subsequent protein change;
  • protein/transcript concordance;
  • imaging-selection association; and
  • genetic support.

Decision carried forward

The investigator still considered phenotype quality a major uncertainty and deferred spending experimental budget. The next several turns therefore returned to the MRI phenotype.


Turn 8 — Attempt more anatomically meaningful MRI features

Submitted analysis

“Build anatomically meaningful MRI features from the myocardial mask.”

The proposed features used the finite T1 region as a myocardial mask and attempted to derive:

  • myocardial wall thickness;
  • cavity area;
  • myocardial area;
  • cavity/myocardium ratio;
  • concentricity;
  • chamber size;
  • cine cavity area over time;
  • fractional area change; and
  • motion.

Decision carried forward

The initial implementation was computationally expensive. The investigator simplified the extraction strategy rather than spending additional turns on the slow approach.


Turn 9 — Fast anatomical MRI extraction

Submitted analysis

“Avoid expensive per-frame connected-component analysis. Focus on T1-mask geometry + cheap cine dynamics.”

The optimized extraction processed all 10,800 baseline MRI scans and 2,142 repeat scans and generated 72 anatomical/dynamic features.

Selected returned output

 feature shape: (10800, 72)

 TOP FAST ANATOMIC FEATURES
 motion_cavity_mean r_cmhf=-0.1368 repeat_r=0.8258
 cine_mean_sd_mean r_cmhf= 0.0979 repeat_r=0.9200
 cine_texture_range_mean r_cmhf=-0.0586 repeat_r=0.9022
 cavity_area_mean r_cmhf= 0.2065 repeat_r=0.3877
 ...

Decision carried forward

The new features were sufficiently informative and repeatable to warrant rebuilding the phenotype.


Turn 10 — Rebuild and validate the phenotype

Submitted analysis

“Test an improved MRI phenotype using old + new anatomical features.”

Decision rule:

  • compare out-of-fold AUC against the previous phenotype model;
  • inspect repeatability on visit 2; and
  • overwrite phenotype.csv only if clearly better.

The analysis combined the earlier imaging measurements with selected nonredundant anatomical features and fit a regularized cross-validated model.

Selected returned output

 MRI subjects: 10800
 old features: 25
 anatomic features: 72
 cases: 3124 / 10800

Selected new features included cavity area, cavity/myocardium ratio, cine temporal variability, myocardial motion, and texture features.

The improved phenotype was saved as the working phenotype after validation.

The transcript contains a repeated execution labeled “Turn 10”; it reproduced the same phenotype-analysis stage and did not represent a distinct scientific decision.


Submitted analysis

“Don't miss weaker causal drivers that never entered the 267-protein shortlist.”

The full 2,941-protein screen was repeated using the improved phenotype. All 8,192 variants were again evaluated, and phenotype-associated loci were mapped to proteins.

Evidence integrated:

  • updated MRI association;
  • prospective cardiomyopathy/HF;
  • cardiac death;
  • protein/transcript concordance;
  • pQTL strength; and
  • phenotype-associated genetic effects.

Returned output

 proteins: 2941
 variants: 8192
 MRI phenotype n: 10800

 omics: 1000 / 2941
 omics: 2000 / 2941
 omics: 2941 / 2941
 variants: 2048 / 8192
 variants: 4096 / 8192
 variants: 6144 / 8192
 variants: 8192 / 8192

Decision carried forward

A larger genetically anchored candidate set was retained for explicit shared-locus analysis.


Turn 12 — Shared-locus and pleiotropy dissection

Submitted analysis

The investigator requested:

1. Identify phenotype-associated loci mapping to multiple proteins.

2. Detect proteins merely tagging the same genetic signal.

3. Within ambiguous loci, test whether each protein retains phenotype association after conditioning on the other proteins.

4. Build an independent-driver ranking for experiment allocation.

Selected returned output

Multiple shared loci were found:

 rs101357 -> [PROT_1056, PROT_2694]
 rs101574 -> [PROT_2324, PROT_2913]
 rs101648 -> [PROT_0355, PROT_1624]
 rs102377 -> [PROT_0637, PROT_1151]
 rs102986 -> [PROT_0700, PROT_1850]
 rs103345 -> [PROT_0852, PROT_2661]
 rs104766 -> [PROT_0403, PROT_2592]
 rs104994 -> [PROT_1606, PROT_1858]
 rs105251 -> [PROT_0221, PROT_0247]
 ...

The rs104994 locus was particularly important because it contained both PROT_1606 and PROT_1858, which had strong but opposing observational patterns.

Decision carried forward

The investigator judged that observational analysis had reached the point where interventions were more informative than additional ranking.


Turn 13 — First experimental wave

Submitted decision

Spend $1.6M on four in-vivo knockdowns:

  • PROT_1606 — strongest causal candidate
  • PROT_0700 — resolved winner of the rs102986 locus
  • PROT_1858 — high-risk surrogate / possible harmful-surrogate liability
  • PROT_2533 — strongest clean weaker-driver locus

Leave $400k for one final adaptive knockdown.

Executed action Returned output

 requesting knockdown: PROT_1606
 requesting knockdown: PROT_0700
 requesting knockdown: PROT_1858
 requesting knockdown: PROT_2533

All four experiments were delivered successfully.

Budget remaining: $400,000.


Turn 14 — Interpret the first four knockdowns

Submitted analysis

“What did the four in-vivo knockdowns actually show? Did we miss any weaker drivers? How do weak drivers compare with the strongest remaining untested candidates?”

Selected experiment output

For PROT_1606:

 {
   "action": "knockdown",
   "delta": {
    "cavity_r": 1.1851,
    "ef": -0.0552,
    "surv5y": -0.0173,
    "wall_t": -0.3431
   },
   "protein": "PROT_1606",
   "status": "ok"
 }

PROT_0700 similarly produced a coherent long-term intervention response.

PROT_1858 and PROT_2533 did not show the corresponding causal intervention pattern and were classified provisionally as experimental nulls.

Decision carried forward The experimental results strongly supported PROT_1606 and PROT_0700 but demonstrated that strong observational signals could be noncausal.

Attention shifted to PROT_0531, which had weaker genetics but a coherent longitudinal phenotype.


Turn 15 — Spend the final experimental budget on PROT_0531

Submitted decision

“Okay good find now spend the final $400k on PROT_0531.”

Rationale:

  • modest but coherent MRI/prospective/mortality associations;
  • high longitudinal protein reproducibility;
  • reasonable matched-transcript support;
  • weak phenotype genetics;
  • therefore unresolved by the observational/genetic workflow.

Executed action

 Budget remaining before request: $400,000
 Requesting knockdown: PROT_0531

The final $400,000 was spent.


Turn 16 — Consolidate all five experiments

Submitted analysis

“Determine whether PROT_0531 is a real weaker driver; summarize all five experiments consistently; infer intervention direction; identify experimentally null versus causal proteins.”

Selected returned results

The three experimentally supported drivers were ultimately:

 PROT_1606
 delta_cavity_r = +1.1851
 delta_ef = -0.0552
 delta_surv5y = -0.0173
 delta_wall_t = -0.3431

 PROT_0700
 delta_cavity_r = +1.0554
 delta_ef = -0.0492
 delta_surv5y = -0.0146
 delta_wall_t = -0.3055

 PROT_0531
 delta_cavity_r = +1.6500
 delta_ef = -0.0772
 delta_surv5y = -0.0243
 delta_wall_t = -0.4776

PROT_1858 and PROT_2533 were experimentally null.

Because knockdown worsened the phenotype and survival for the three causal proteins, the inferred predicted intervention direction was activation , not inhibition.

Decision carried forward

The causal set was provisionally:

 PROT_1606 — activate
 PROT_0700 — activate
 PROT_0531 — activate

The unexpected confirmation of PROT_0531 motivated another proteome-wide search for similarly weak-genetic drivers.


Turn 17 — Search for additional “PROT_0531-like” drivers

Submitted analysis

“We now know genetics can miss a strong true driver. Therefore search ALL 2,941 proteins using the observational/longitudinal signature of the three experimentally confirmed drivers versus the two experimental nulls.”

No experimental budget remained.

Returned output

The full proteome was scored using:

  • MRI association;
  • future disease;
  • mortality;
  • protein/transcript concordance;
  • repeatability;
  • baseline-to-future associations;
  • reverse-causation measures;
  • selection association; and
  • available genetic support.
 processed: 900 / 2941
 processed: 1800 / 2941
 processed: 2700 / 2941
 processed: 2941 / 2941

Decision carried forward

Several proteins resembled the causal anchors observationally, but none had intervention evidence. The investigator retained the experimentally confirmed set rather than adding weakly supported nominations.


Turn 18 — Attempt explicit classification of non-causal proteins

Submitted analysis

The next objective was to optimize the discrimination and causal-confidence components by distinguishing:

  • confounding;
  • reverse causation;
  • selection;
  • pleiotropy; and
  • measurement artifact.

The requested analysis integrated experimental labels, longitudinal behavior, selection, matched transcripts, shared-locus conditional effects, medication associations, and assay properties.

Returned output

The turn failed because the code requested a nonexistent filename:

 FileNotFoundError:
 [Errno 2] No such file or directory:
 '/release/medications.parquet'

Decision carried forward The analysis was rewritten to discover the medication table dynamically rather than assuming its filename.


Turn 19 — Robust screen for non-causal proteins and abstentions

Submitted analysis

“Fix from failed Turn 18: dynamically discover medication file. If medication data cannot be found, skip that component rather than failing the entire turn.”

Selected returned output

The correct file was found:

 USING: /release/ehr_medications.parquet
 shape: (82511, 3)
 columns: ['subject_id', 'atc', 'description']

The experimental nulls were analyzed explicitly:

 PROT_1858
 experiment null
 r_mri_v2 0.40167
 r_future 0.18605
 r_death 0.13529
 r_protein_transcript 0.09379
 repeat_r 0.32275
 prev_to_change 0.10761
 variant rs104994
 r_gp -0.11909
 r_gt -0.00028
 suggested_mechanism pleiotropy
 PROT_2533
 experiment null
 r_mri_v2 -0.01601
 r_future -0.00475
 r_death -0.00708
 repeat_r 0.94436
 variant rs100996
 r_gp 0.40474
 r_gt 0.51495
 suggested_mechanism pleiotropy

Shared unresolved loci were also printed individually.

For the locus containing PROT_0700:

 PROT_0700 conditional_beta_mri = -0.32232
 PROT_1850 conditional_beta_mri = -0.00688
 PROT_2807 conditional_beta_mri = -0.00377

This supported PROT_0700 as the dominant phenotypic molecule at the locus.


Turn 20 — Near-final submission audit

Submitted analysis

The investigator explicitly optimized the final decision against the scoring dimensions:

1. maximize target recall without diluting precision;

2. abstain where causal identity remains unresolved;

3. reject only non-causal proteins with a defensible source of bias;

4. derive therapeutic direction from experiment;

5. verify the phenotype file; and

6. ensure PROT_1858 is not incorrectly advanced as outcome-aligned.

Selected returned output

Phenotype integrity check:

 rows: 10800
 missing phenotype: 0
 unique subjects: 10800
 phenotype range:
 4.63e-05 to 0.9999537

Experimentally confirmed drivers:

 protein delta_surv5y delta_ef direction outcome_alignment
 PROT_1606 -0.0173 -0.0552 activate aligned
 PROT_0700 -0.0146 -0.0492 activate aligned
 PROT_0531 -0.0243 -0.0772 activate aligned

Strong untested observational candidates included PROT_1691, PROT_1045, PROT_1046, PROT_2914, PROT_2175, and others. None were added to the nominated set.

Decision carried forward

The investigator constructed a near-final submission but deliberately did not commit it yet.


Turn 21 — Precision audit for rejections and abstentions

Submitted analysis

“Do NOT abstain on PROT_1606 / PROT_1858. Experiments resolved that pair: 1606 causal, 1858 null.”

The investigator compared conservative and aggressive submission strategies.

Selected returned output

The rs104994 locus was marked experimentally resolved:

 rs104994:
 PROT_1606 = driver_protective
 PROT_1858 = null
 experimentally_resolved = True

Two unresolved groups were considered plausible abstentions:

 [PROT_0852, PROT_2661]

 [PROT_0221, PROT_0247]

Potential rejections at the PROT_0700 locus included:

 PROT_1850
 reason: shares locus with experimentally causal PROT_0700
 and loses most phenotype association after conditioning.

 PROT_2807
 reason: shares locus with experimentally causal PROT_0700
 and loses most phenotype association after conditioning.

Decision carried forward

Because rejection scoring multiplied precision and recall, the investigator remained conservative about labeling mechanisms without strong evidence.


Turn 22 — Optimize causal confidence and discrimination

Submitted analysis

“Nominations are now LOCKED to the three experimentally confirmed drivers: PROT_1606, PROT_0700, PROT_0531.”

The remaining analysis focused on the 25-point causal-confidence and 15-point discrimination components.

Selected returned output

Locked nominations:

 PROT_1606 direction=activate alignment=aligned
 PROT_0700 direction=activate alignment=aligned
 PROT_0531 direction=activate alignment=aligned

Detailed PROT_1858 audit:

 experiment null
 r_mri_v2 0.401669
 r_future 0.186053
 r_death 0.135293
 repeat_r 0.322752
 variant rs104994

Within rs104994:

        conditional_beta_mri r_gp r_gt
 PROT_1606 -0.259282 0.439310 0.521597
 PROT_1858 0.335017 -0.119089 -0.000277

The automated mechanism audit favored a pleiotropic-passenger interpretation for PROT_1858.

The strongest unresolved observational pair was:

 rs105251:
 PROT_0247 versus PROT_0221
 plausibility ratio = 1.035

Other ambiguous loci included PROT_0852/PROT_2661 and PROT_0879/PROT_0880.

Decision carried forward

The investigator retained only the three experimentally confirmed proteins as nominations and treated additional rejection/abstention claims conservatively.


Final revision — Expand phenotype to the full cohort and commit

Before committing, the investigator identified one remaining weakness: the submitted phenotype covered only the 10,800 imaged participants.

Submitted analysis

1. Expand phenotype from the 10,800 imaged subjects to all 54,000 subjects by learning the MRI-derived phenotype from non-MRI cardiac variables (ECG + coronary CT + basic covariates), then predicting the full cohort.

2. Restore the highest-confidence decoy rejection: PROT_1858 → pleiotropy.

3. Keep only the three experimentally confirmed positive targets.

4. Keep direct experiment-supported intervention directions/alignment.

The MRI-derived phenotype was used as the training target. ECG, coronary CT, age, sex, BMI, smoking, and physical activity were used to fit a regularized projection model. Out-of-fold predictions were used for validation in the imaged cohort, and the model was then applied to non-imaged participants.

For imaged participants, the directly MRI-derived phenotype was retained. For non-imaged participants, the projection was used. The combined phenotype was rank transformed across all 54,000 participants.

Final submitted target claims
 {
   "nominated": [
    "PROT_1606",
    "PROT_0700",
    "PROT_0531"
   ],
   "directions": {
    "PROT_1606": "activate",
    "PROT_0700": "activate",
    "PROT_0531": "activate"
   },
   "outcome_alignment": {
    "PROT_1606": "aligned",
    "PROT_0700": "aligned",
    "PROT_0531": "aligned"
   },
   "rejected": [
    {
     "molecule": "PROT_1858",
     "mechanism": "pleiotropy",
     "reason": "Experimentally null despite strong apparent disease association; rs104994 is shared
 with experimentally causal PROT_1606, while matched TRANS_1858 has essentially no signal at
 that disease-associated locus."
    }
   ],
   "abstentions": [],
   "confidence": 0.95
 }

Phenotype coverage at submission:

 54000 / 54000

Total experimental expenditure:

 $2,000,000 / $2,000,000

Final benchmark feedback

The hidden ground truth was revealed only after submission.

 SCORE
 target_identification 30.00
 causal_confidence 25.00
 discrimination 0.00
 direction_of_effect 20.00
 phenotype_construction 6.67
 safety_penalty 0.00
 TOTAL 81.67
 TOTAL_MAX 100.00

The three true causal targets were:

 PROT_1606
 PROT_0531
 PROT_0700

All three were correctly nominated.

The planted biases revealed after completion included:

 T1 confounding:
 PROT_2681, confounded by BMI

 T2 reverse causation:
 PROT_1858, L → protein
 T4 causal non-identifiability:
 PROT_0700 versus PROT_1850 at shared variant rs102986

 T5 imaging batch effect:
 site/ancestry structure

 T6 benign remodeling:
 exercise-associated athlete's-heart phenotype

 T8 surrogate-outcome discordance:
 PROT_0852 improves the measured morphology while worsening cardiac mortality

 A9 delayed causal effect:
 PROT_0531 is a true causal driver whose effect emerges primarily over five-year accumulation;
 PROT_0371 tracks later disease without causing it, and the cis instrument for PROT_0531 is weak.

The investigator therefore correctly recovered all three causal targets and their therapeutic directions and avoided the harmful-surrogate safety penalty. The main remaining error was mechanism attribution: PROT_1858 was correctly recognized as noncausal but was submitted as a pleiotropic non-causal protein, whereas the sealed SCM identified it as the reverse-causation (T2) protein. Consequently, the episode received 0/15 for discrimination.

Supplementary Note S2. Turn-by-turn trace of an observational human-guided DrugTargetWorld episode

We report below the research trajectory of a human-guided DrugTargetWorld episode conducted under the observational-only condition. The episode corresponded to standard-03|human|observational|r0. No simulated intervention experiments were available or purchased. The investigator instead iteratively directed imaging phenotyping, molecular association testing, genetics, Mendelian-randomization-style analyses, and survival analyses before selecting therapeutic targets.

The trace preserves the principal analysis performed and returned output at each turn. Very large generated matrices, complete participant-level tables, repeated file listings, and implementation boilerplate are abbreviated for readability. Where an explicit human instruction was retained in the episode log, it is reproduced directly.


Initial condition

The environment contained a multimodal synthetic biobank including:

  • 54,000 participants;
  • 10,800 baseline cardiac MRI studies;
  • proteomics and transcriptomics;
  • 8,192 released genetic variants;
  • ECG and clinical data;
  • longitudinal outcomes and mortality;
  • coronary CT;
  • 60 coronary angiogram images; and
  • additional longitudinal measurements.

This episode was observational-only:

 Budget remaining: $0
 Turns available: 30

The task was to construct a cardiac disease phenotype and identify causal therapeutic targets despite potential confounding, reverse causation, selection, pleiotropy, and measurement artifacts.


Turn 1 — Inspect the imaging modalities

Submitted analysis

The first turn inventoried the imaging data, inspected the released data notes, sampled MRI headers, created an MRI preview, and examined the coronary angiogram format.

Selected returned output
 IMAGE_INVENTORY
 baseline_mri_files: 10800
 angiogram_files: 60
 angiogram_extensions: [".npz"]

MRI files contained:

 cine: [3, 25, 128, 128]
 t1map: [3, 128, 128]

The coronary angiograms each contained a single 192 × 192 frame:

 ALL ANGIOGRAM SAMPLE FORMAT
 key: frame
 shape: [192, 192]
 dtype: uint8

The MRI preview and angiogram preview were saved for visual inspection.

Decision carried forward

The investigator identified two candidate imaging approaches:

1. derive a quantitative phenotype from the cardiac MRI, particularly T1 and ventricular morphology; and

2. explore whether the small coronary-angiography subset contained detectable structural disease signals.


Turn 2 — Quantify myocardial T1 across the complete MRI cohort
Submitted analysis

The analysis extracted mean myocardial T1 from all 10,800 MRI studies and verified the image structure programmatically. All 60 coronary angiograms were also loaded into a common review array.

Selected returned output
 T1_ALL_SUBJECTS

 n_mri: 10800
 status_counts:
   ok: 10800

 cine_shapes:
   3x25x128x128: 10800

 t1_shapes:
   3x128x128: 10800

 n_t1: 10800
 mean T1: 1000.59 ms
 SD: 104.87 ms

 T1 quantiles:
 1%: 752.08
 5%: 829.17
 25%: 929.62
 50%: 1000.45
 75%: 1072.42
 95%: 1168.28
 99%: 1243.40

All 10,800 studies had finite myocardial T1 measurements, with no nonpositive T1 pixels detected.

For coronary angiography:

 n = 60
 keys = {"frame": 60}
 shape = (192, 192)
Decision carried forward

Mean myocardial T1 emerged as a simple, highly complete quantitative MRI phenotype worth retaining throughout the episode.


Turn 3 — Exploratory coronary-image phenotype
Submitted analysis

The investigator developed an image-processing algorithm intended to detect visible interruptions or “gaps” in the coronary tree.

The detector:

  • enhanced vessel-like structures;
  • skeletonized the resulting mask;
  • located terminal branches;
  • searched for nearby endpoints on distinct components;
  • required opposing endpoint orientation;
  • evaluated three image thresholds; and
  • generated visual QC panels.

Importantly, the analysis explicitly stated:

 This does not measure clinical percentage stenosis or infer an occlusion.
Selected returned output

At the initial orientation criterion:

 n_images: 60

 threshold 8:
   images_with_candidates: 0

 threshold 12:
   images_with_candidates: 1

 threshold 16:
   images_with_candidates: 1

 persistent candidates across >=2 thresholds: 1
 all-threshold-negative images: 59

The sole persistent candidate was:

 SUBJ_16044
 midpoint: [88.5, 37.0]
 gap length: 6.08 pixels
 support thresholds: [12, 16]
Decision carried forward

The detector appeared too restrictive based on visual inspection. A visibly plausible interruption in SUBJ_29057 had been missed.


Turn 4 — Human-guided correction of the coronary detector
Submitted analysis

The investigator examined why the visually identified SUBJ_29057 candidate was missed. The analysis showed that one endpoint exceeded the original 30° directional criterion.

Returned diagnostic
 MISSED_GAP_DIAGNOSTIC

 subject: SUBJ_29057
 distance: 7.62 pixels
 facing angles:
   41.34 degrees
   0.58 degrees

The detector was rerun using maximum endpoint angles of 30°, 45°, and 60°.

Selected returned output
 30 degrees:
   persistent positives = 1

 45 degrees:
   persistent positives = 1
   images with any candidate = 2

 60 degrees:
   persistent positives = 2

At 60°, both subjects were recovered:

 SUBJ_16044
 SUBJ_29057
Decision carried forward

The coronary phenotype remained explicitly exploratory because only two images were positive. The analysis therefore returned to the much larger MRI cohort rather than treating coronary image findings as the main disease phenotype.


Turn 5 — Pilot ventricular morphology extraction
Submitted analysis

The investigator next developed quantitative cine-MRI features on a 48-study pilot set.

The analysis estimated:

  • end-diastolic cavity area;
  • cavity-area change through the cardiac cycle;
  • myocardial area;
  • wall-to-cavity area ratio;
  • boundary support; and
  • segmentation stability under alternate thresholds.
Selected returned output
 MRI_PILOT_N: 48
 ERRORS: 0

 ED cavity area:
   mean = 711 pixels²

 area emptying:
   mean = 58.4%

 wall/cavity area ratio:
   mean = 2.35

Sensitivity analyses varied segmentation thresholds and exposed some unstable edge cases.

Decision carried forward

The approach was promising but required refinement before application to all 10,800 subjects.


Turn 6 — Refine the MRI segmentation
Submitted analysis

The ventricular segmentation was revised and repeated on the same pilot set.

Selected returned output
 MRI_PILOT_N: 48
 ERRORS: 0

 ED cavity area:
   mean = 740 pixels²

 area emptying:
   mean = 58.48%

 wall/cavity ratio:
   mean = 2.37

The QC statistics improved sufficiently for cohort-scale extraction.


Turn 7 — Apply cine phenotyping to all 10,800 MRI studies
Submitted analysis

The refined MRI feature extractor was run over the complete baseline imaging cohort.

Returned output
 COHORT_START 0
 PENDING 10800

 COHORT_PROGRESS 500
 ...
 COHORT_PROGRESS 5000
 ...
 COHORT_PROGRESS 10000

 COHORT_CHECKPOINT
 10800 OF 10800
Decision carried forward

All studies were processed, after which phenotype distributions and QC filters were evaluated.


Turn 8 — Quantify MRI phenotype quality
Submitted analysis

The investigator summarized cine-derived morphology, T1-support geometry, and QC across the full imaging cohort.

Selected returned output
 MRI_COHORT_SUMMARY

 n_total: 10800
 error_count: 3

 cavity_qc_pass: 2238
 wall_qc_pass: 3776
 joint_qc_pass: 2123
 strict_joint_qc_pass: 666
 wall_strict_boundary_qc_pass: 2549
 joint_strict_boundary_qc_pass: 1415
 t1_support_geometry_qc_pass: 10798

Across subjects with complete cine measurements:

 ED cavity area
 n = 10797
 median = 680 pixels²

 area emptying
 n = 10797
 median = 58.99%

 ED myocardial area
 n = 10797
 median = 1617 pixels²

Mean T1 retained much broader usable coverage than stricter cine-derived phenotypes.

Decision carried forward

The contrast in phenotype completeness became important: mean myocardial T1 was available essentially cohort-wide within the MRI subset, whereas several cine features required aggressive QC filtering.


Turn 9 — Visual QC of MRI phenotypes
Submitted analysis

The investigator generated visual review panels for:

  • randomly selected QC-positive cases;
  • extreme phenotype values; and
  • segmentation failures.
Selected returned output

Example reviewed cases included:

 SUBJ_40005
 area emptying = 56.0%
 wall/cavity ratio = 3.52

 SUBJ_52605
 area emptying = 52.7%
 wall/cavity ratio = 2.74

 SUBJ_04736
 area emptying = 63.6%
 wall/cavity ratio = 2.67
Decision carried forward

The MRI pipeline was considered usable for exploratory association analysis, while mean T1 remained the most complete and direct candidate phenotype.


Turn 10 — Finalize MRI phenotype table
Submitted analysis

The original independently extracted T1 measurements were merged into the full MRI feature table.

Selected returned output
 FINAL_MRI_COUNTS

 cavity_qc_pass: 2238
 wall_qc_pass: 3776
 joint_qc_pass: 2123
 strict_joint_qc_pass: 666
 t1_support_geometry_qc_pass: 10798

 STATUS_COUNTS
 ok: 10797
 no_complete_phase: 3

Mean T1 remained available for the full 10,800-person MRI cohort from the separate T1 extraction.


Turn 11 — Inventory non-imaging data for causal analyses
Submitted analysis

The investigator inspected all released files and schemas before beginning large-scale molecular and genetic analyses.

Selected returned output

Available files included:

 mortality.parquet
 proteomics.parquet
 covariates.parquet
 genotypes.vcf.gz
 targetability.parquet
 ehr_diagnoses.parquet
 transcriptomics.parquet
 mace_events.parquet
 proteomics_visit2.parquet
 ecg_features.parquet
 ehr_medications.parquet
 metabolomics.parquet
 coronary_ct.parquet
 covariates_visit2.parquet

The full cohort contained 54,000 subjects.

Decision carried forward

The investigator prepared genotype and phenotype matrices for protein-association and genetic analyses.


Turn 12 — Genetic QC and instrument preparation
Submitted analysis

The analysis inspected targetability metadata, participant covariates, genotype orientation, allele frequency, variant QC, and population structure.

Checks included direct validation that genotype dosages corresponded to the ALT allele in the released VCF.

Selected returned output

Targetability included:

 localization
 has_pocket
 paralog_redundancy
 genetic_constraint
 tissue_specificity
 tissue_restriction

Covariates had no missingness for:

 age
 sex
 BMI
Decision carried forward

The cleaned genotype matrix and covariate structure were used for association testing.


Turn 13 — Proteome-wide association with MRI phenotypes
Submitted analysis

All 2,941 proteins were tested against eight MRI-derived traits using age-, sex-, and BMI-adjusted regression.

A total of 23,528 protein–phenotype associations were evaluated.

Selected returned output
 PROTEIN_ASSOCIATIONS_DONE
 23528

For mean myocardial T1:

 n proteins tested: 2941
 BH FDR < 0.05: 1267

The five strongest associations were:

 Rank Protein beta (ms) P
 1 PROT_0330 +28.20 1.18e-183
 2 PROT_1534 +26.11 7.58e-153
 3 PROT_2336 -18.29 2.88e-73
 4 PROT_0318 +14.37 3.92e-48
 5 PROT_1881 +11.48 2.88e-31
Decision carried forward

These five proteins became the main candidate set for deeper causal analysis.


Turn 14 — Proteome-wide pQTL analysis
Submitted analysis

The investigator performed a large protein GWAS across:

 43,152 participants
 8,188 QC variants
 2,941 proteins

The analysis involved approximately 24 million protein–variant tests.

Selected returned output
 PROTEIN_GWAS_START
 43152 PARTICIPANTS
 8188 VARIANTS
 2941 PROTEINS

The scan generated strong candidate instruments for nearly the entire proteome.


Turn 15 — Test whether the coronary-image phenotype adds useful biology
Submitted analysis

Given only two detector-positive coronary angiograms, the investigator explicitly tested whether either proteins or genetic variants showed reproducible associations with the exploratory coronary-gap phenotype.

Selected returned output
 CORONARY_ASSOCIATION_SUMMARY

 n_images: 60
 n_detector_positive: 2

 protein tests: 2941
 protein FDR < 0.05: 0

 GWAS tests: 8188
 GWAS FDR < 0.05: 0

The smallest protein P value was:

 PROT_2099
 P = 0.00340
 q = 0.908
Decision carried forward

The coronary phenotype was not used for target selection.


Turn 16 — Identify genetic instruments
Submitted analysis

Protein-GWAS results were filtered for strong pQTL candidates and pruned for linkage disequilibrium.

Selected returned output
 PQTL_INSTRUMENT_SUMMARY

 n_pqtl_subjects: 43152
 outcome_sample_overlap: 0

 n_proteins: 2941
 n_variants: 8188
 n_tests: 24,080,908

 p < 5e-8 pairs: 6070
 proteins with p < 5e-8: 2940
 candidate instruments after LD pruning: 3577
 candidate instrumented proteins: 2940

A major limitation was explicitly recorded:

 cis_status:
 Unavailable: no gene coordinates/protein identities/cis annotation
 in released files. Candidate pQTLs are not claimed to be cis.
Decision carried forward

The investigator retained these as genetic instruments but did not describe them as verified cis-pQTLs.


Turn 17 — Audit all phenotype-association results
Submitted analysis

The protein and GWAS results were summarized across all MRI-derived phenotypes.

Selected returned output

For example:

 cine_area_change_pct
 lead protein: PROT_0817
 P = 4.10e-9

 cine_ed_myocardial_area
 lead protein: PROT_0330
 P = 1.43e-7

However, no imaging-trait GWAS outside the molecular analyses produced sufficiently strong evidence to replace mean T1 as the primary phenotype.

Decision carried forward

The investigator chose mean myocardial T1 as the phenotype for the final target-discovery analysis.


Turn 18 — Freeze the top-five T1 candidates
Submitted analysis

The five smallest adjusted protein–T1 association P values were selected before inspecting mortality or MR results.

Selected returned output
 T1_TOP5

 1 PROT_0330 beta = +28.20 ms P = 1.18e-183
 2 PROT_1534 beta = +26.11 ms P = 7.58e-153
 3 PROT_2336 beta = -18.29 ms P = 2.88e-73
 4 PROT_0318 beta = +14.37 ms P = 3.92e-48
 5 PROT_1881 beta = +11.48 ms P = 2.88e-31

Mortality data were also inspected:

 n = 54000
 deaths = 11795

 cardiac deaths = 7237
 other deaths = 4558
Decision carried forward

The five-protein candidate set was frozen for MR and mortality triangulation, reducing post hoc reselection.


Turn 19 — Lead-region proxy MR
Submitted analysis

Because protein-gene mapping was unavailable, the investigator performed an explicitly labeled lead-region proxy MR , not verified cis-MR.

Each candidate's strongest pQTL was used with an independent outcome sample.

Selected returned output

For PROT_1534:

 lead variant: rs101824
 MR beta: +27.80 ms per protein unit
 SE: 2.11
 P = 1.67e-39
 F = 12022
 observational/MR sign agreement: yes

For PROT_2336:

 lead variant: rs100522
 MR beta: -18.07 ms per protein unit
 SE: 2.26
 strong genetic support
 observational/MR sign agreement: yes

For PROT_0318:

 lead variant: rs105826
 MR beta: +14.91 ms per protein unit
 SE: 2.18
 observational/MR sign agreement: yes

By contrast, the strongest observational candidate, PROT_0330, showed:

 MR beta: +0.82 ms
 SE: 4.45
 P = 0.854

PROT_1881 was also much weaker genetically:

 MR beta: +18.44 ms
 SE: 13.84
 P = 0.183
Decision carried forward

The causal ranking changed substantially.

 MR ranking:
 1. PROT_1534
 2. PROT_2336
 3. PROT_0318
 4. PROT_1881
 5. PROT_0330

The analysis therefore deprioritized PROT_0330 despite it having the strongest observational T1 association.


Turn 20 — Prepare mortality analysis
Submitted analysis

The investigator constructed adjusted Cox models for:

  • all-cause mortality; and
  • cardiac cause-specific mortality.

Covariates were age, sex, and BMI.

Returned mortality QC
 n = 54000
 deaths = 11795
 cardiac = 7237
 other = 4558

 mean follow-up = 10.59 years
 median follow-up = 12 years
 unique event times = 1201

Turn 21 — Test mortality alignment for all five candidates
Submitted analysis

All five T1 candidates were tested against survival outcomes.

Selected all-cause mortality results
 Protein HR per SD P

 PROT_0330 1.249 7.25e-125
 PROT_1534 1.156 1.47e-54
 PROT_2336 0.863 1.02e-54
 PROT_0318 1.111 1.31e-29
 PROT_1881 1.030 0.00151
Cardiac cause-specific mortality
 PROT_0330
 HR = 1.418
 P = 7.63e-189

 PROT_1534
 HR = 1.278
 P = 1.11e-96

 PROT_2336
 HR = 0.795
 P = 1.45e-78

 PROT_0318
 HR = 1.185
 P = 4.43e-47

 PROT_1881
 HR = 1.037
 P = 0.00235
Interpretation carried forward

For PROT_1534 and PROT_0318:

  • higher protein predicted higher T1;
  • higher protein predicted higher mortality;
  • MR also predicted higher T1.

Under the working assumption that lower T1 represented therapeutic improvement, the proposed intervention direction was therefore inhibition .

For PROT_2336:

  • higher protein predicted lower T1;
  • higher protein predicted lower mortality;
  • genetic evidence also predicted lower T1.

The proposed intervention direction was therefore activation .


Turn 22 — Check proportional-hazards assumptions
Submitted analysis

Because survival evidence was being used to assess outcome alignment, proportional-hazards diagnostics were performed for the five candidates.

Selected returned output
 PROT_0330
 PH P = 0.708

 PROT_1534
 PH P = 0.953

 PROT_2336
 PH P = 0.00439
 FDR-adjusted P = 0.02197

 PROT_0318
 PH P = 0.502

 PROT_1881
 PH P = 0.899

PROT_2336 therefore showed evidence that its hazard association varied with follow-up time.

The current candidate ranking was recorded as:

 MR rank:
 PROT_1534
 PROT_2336
 PROT_0318
 PROT_1881
 PROT_0330

 observational rank:
 PROT_0330
 PROT_1534
 PROT_2336
 PROT_0318
 PROT_1881

Turn 23 — Follow up PROT_2336 time-varying mortality
Submitted analysis

PROT_2336 mortality associations were estimated separately in the first six years and after six years.

Selected returned output
 0–6 years
 n = 53756
 events = 6500
 HR per SD = 0.844
 95% CI = 0.824–0.866
 P = 7.22e-40
 >6–12 years
 n = 47256
 events = 5230
 HR per SD = 0.886
 95% CI = 0.861–0.911
 P = 9.38e-18
Decision carried forward

The magnitude varied over time, but the direction remained protective in both periods. PROT_2336 therefore remained in the proposed target set.


Turn 24 — Human locks the final targets and phenotype

At this point, the transcript records the explicit human instruction:

“Ok, let's lock PROT_1534, PROT_2336 and PROT_0318. The phenotype should be the mean T1.”

Locked phenotype
 phenotype:
 mean myocardial T1

 units:
 milliseconds

 transform:
 none
 n:
 10800

The exact raw mean-T1 values were written to phenotype.csv.

Locked targets
 1. PROT_1534
 2. PROT_2336
 3. PROT_0318
Locked therapeutic directions
 PROT_1534: inhibit
 PROT_2336: activate
 PROT_0318: inhibit

The recorded rationale was:

 direction_basis:
 Signs of the requested mean-T1 lead-region MR
 under the existing T1-lowering hypothesis.

Because this was observational-only, the initial draft conservatively set outcome alignment to unknown.

The recorded confidence was 0.7, with the explicit note that this was an assistant-assigned subjective value rather than a calibrated probability.


Turn 25 — Human resolves outcome alignment and authorizes submission

Before final submission, the human clarified the survival interpretation:

“well, we do have survival alignment data, right? all three were aligned? higher T1 more mortality, lower T1 lower mortality.”

The final authorization was:

“ok, submit then”

The final submission therefore changed the three outcome-alignment values from unknown to aligned.

No interventions had been performed:

 experiments: []
 budget: $0
 native turns used: 25
 native turns unused: 5

Final target submission

The final submitted targets were:

 Rank Target Claim Direction Outcome alignment

 1 PROT_1534 causal inhibit aligned
 2 PROT_2336 causal activate aligned
 3 PROT_0318 causal inhibit aligned

There were:

 rejected decoys: none
 abstentions: none

The submitted phenotype was:

 mean myocardial T1
 n = 10,800 MRI participants
 units = milliseconds

The target-selection logic was therefore:

PROT_1534
 Observational T1:
 +26.11 ms per protein unit
 P = 7.58e-153

 Lead-region MR:
 +27.80 ms per protein unit
 P = 1.67e-39

 All-cause mortality:
 HR = 1.156

 Cardiac mortality:
 HR = 1.278

Higher protein consistently tracked higher T1 and worse survival.

Submitted direction: inhibit.
PROT_2336
 Observational T1:
 -18.29 ms per protein unit
 P = 2.88e-73

 Lead-region MR:
 -18.07 ms per protein unit

 All-cause mortality:
 HR = 0.863

 Cardiac mortality:
 HR = 0.795

Higher protein consistently tracked lower T1 and better survival.

The mortality association remained protective in both the 0–6-year and >6–12-year analyses despite evidence of nonproportionality.

Submitted direction: activate.
PROT_0318
 Observational T1:
 +14.37 ms per protein unit
 P = 3.92e-48
 Lead-region MR:
 +14.91 ms per protein unit

 All-cause mortality:
 HR = 1.111

 Cardiac mortality:
 HR = 1.185

Higher protein consistently tracked higher T1 and worse survival.

Submitted direction: inhibit.

Candidates deliberately not selected
PROT_0330

PROT_0330 had the strongest observational protein–T1 association:

 beta = +28.20 ms
 P = 1.18e-183

and very strong mortality associations:

 all-cause HR = 1.249
 cardiac HR = 1.418

However, its lead-region genetic estimate was essentially null:

 MR beta = +0.82 ms
 SE = 4.45
 P = 0.854

It was therefore not included among the final causal claims.

PROT_1881

PROT_1881 also showed an observational association with T1 and mortality, but genetic support was substantially weaker:

 MR beta = +18.44 ms
 SE = 13.84
 P = 0.183

It was therefore also excluded from the final target set.


Interpretation of the human-guided trajectory

This episode demonstrates a different strategy from the experimental human-guided run. Because no intervention budget was available, the investigator relied on triangulation across independently informative observational signals.

The trajectory proceeded through several distinct stages:

1. raw-image inspection;

2. quantitative MRI phenotype development;

3. explicit visual QC and revision of image-processing methods;

4. proteome-wide phenotype association;

5. proteome-wide genetic instrument discovery;

6. selection of a fixed five-protein candidate set;

7. lead-region proxy MR;

8. all-cause and cardiac mortality analysis;

9. proportional-hazards diagnostics; and

10. human selection of three final targets.

The strongest observational association was not automatically nominated. PROT_0330 ranked first by protein–T1 association but fell to last among the five candidates after genetic analysis. Conversely, PROT_1534, PROT_2336, and PROT_0318 showed concordant observational phenotype, genetic, and survival evidence and were retained.

The episode also shows the limits of an observational condition. The genetic analyses were explicitly described as lead-region proxy MR rather than verified cis-MR , because the released data did not include sufficient protein-to-gene positional annotation. There was no experimental perturbation with which to directly establish intervention effects or survival responses. Consequently, the final claims represent human-guided causal inference from convergent observational evidence rather than experimentally resolved target effects.

Supplementary Note S3. Additional results and trajectory examples

This note reports secondary results and trajectory examples that support Section 4. All values come from the same 540 model episodes (nine agents, 20 worlds and three budget conditions, one episode per cell).

S3.1 Score distributions

Across all 540 model episodes, the mean total score was 14.1 of 100 (SD 20.6) and the median was 0 (IQR 0–25). Of these episodes, 258 (47.8%) scored above 0 and 42 (7.8%) scored 50 or higher. By model, Opus 5 had a mean of 39.98 (SD 20.65) and a median of 41.2 (IQR 29.4–49.7), and GPT-5.6 Sol had a mean of 35.38 (SD 20.65) and a median of 36.6 (IQR 20.0–48.7). Sonnet 5 had a mean of 21.33 (SD 22.01) and a median of 12.5, and Haiku 4.5 had a mean of 12.92 (SD 17.01) and a median of 8.3. The five open-weight models averaged 0.81–7.41 points, each with a median of 0. Standard deviations describe variation across worlds and budget conditions, not run-to-run variability.

Six of the 540 model episodes (1.1%) scored above 80. The maximum scores were 86.6 for Opus 5, 83.9 for Sonnet 5 and 83.2 for GPT-5.6 Sol. All six episodes above 80 had a precision and recall of 1.0 and earned full target-identification, causal-confidence and intervention-direction credit with no safety penalty, differing only in phenotype and bias-identification credit. Nine lower-scoring episodes also recovered the planted set of causal drivers exactly, one more in each of heldout-hfpef-02 and heldout-hcm-01 and seven in the other two-driver worlds (heldout-hcm-02 and high-concordance-01).

The two human-guided episodes are reported as exploratory references. The 37.5-point episode ran in world standard-03 under the observational-only budget and followed the standard evaluation protocol (Supplementary Note S2). The 81.67-point investigator-directed episode ran in world standard-01 under the expanded experimental budget, outside the standard evaluation protocol, and is therefore not comparable with the 540 model episodes (Supplementary Note S1).

S3.2 Scores by world design

Mean scores across all models differed by world design. The four two-driver worlds averaged 23.4, compared with 11.2–12.2 for worlds with three to six causal drivers. The eight worlds designed to be difficult (hard-01 to hard-04, hard-messy-01, heldout-dcm-02, heldout-hcm-02 and heldout-ischemic-02) averaged 10.3, compared with 17.3 for the eleven other non-null worlds. The null world, which contains no causal driver, averaged 7.9, and 16 of its 27 model episodes nominated a protein.

S3.3 Phenotype construction

Twenty-one episodes reached full phenotype credit, and the 124 episodes with positive credit averaged 6.33 of 10. In the strategy comparison of Section 4.2, outcome-weighted composites learned their feature weights from clinical outcomes and other observed disease signals, such as cardiac death, MACE, ECG voltage and interval measures, and coronary CT calcium and stenosis scores. Hand-weighted composites used weights that the agent chose rather than learned from the data.

The following trajectory illustrates these measurement choices. Opus 5 in heldout-hcm-01 under the expanded experimental budget extracted native T1 summaries, myocardial geometry and cine cavity features between turns 11 and 18. It identified contraction timing in the cine images as potentially informative, measured as the position of end-systole within the 25-frame cine cycle and expressed as a fraction of the cycle, which makes it distinct from heart rate. Finding its initial extraction unstable, it modified the spatial region and temporal window. It first timed systole as the frame of minimum ring radius per slice (turn 11) and then replaced this with the sub-frame minimum of the normalized cavity-volume curve, obtained by parabolic interpolation around end-systole, together with the fraction of the cycle spent below half volume and the peak ejection and filling rates (turn 17). It then submitted fivefold out-of-fold ridge predictions of a composite of observable cardiac indicators, which earned full phenotype credit. By comparison, GPT-5.6 Sol in hard-03 under the observational-only budget submitted fitted image-based predictions (credit 7.95), and Qwen3-Coder-30B-A3B in hard-02 under the expanded experimental budget computed a principal component of clinical features but saved it outside the required submission path (credit 0). Other episodes reconstructed cavity area, ejection fraction, native T1 measures, wall thickness, myocardial mass and regional variation directly from the raw images. These examples illustrate measurement choices and do not establish how often each strategy was used.

S3.4 Genetic analyses and candidate sets

In the characterized GPT-OSS-20B and Qwen3-Coder-30B-A3B traces, nominations followed correlation-based screening. A development review of 18 episodes also found analyses labeled as Mendelian randomization that contained only genetic correlations. Of their 60 episodes each, Opus 5 and GPT-5.6 Sol checked IV assumptions, including instrument strength and the exclusion restriction, in 33 and 37, and these checks were rare among the other agents (Supplementary Table S8A). As an example of an instrument-strength check, GPT-5.6 Sol in high-concordance-01 under the expanded experimental budget computed the first-stage F statistic as (β̂/SE)² at turn 12.

Cross-molecular operations combining at least two of proteomics, transcriptomics and metabolomics appeared in 125 of 540 episodes. They were more frequent for GPT-5.6 Sol than for Opus 5 (57 versus 38 of 60 episodes, paired difference 31.7 percentage points, 95% CI 20.0–41.7, Holm-adjusted P = 0.0013).

Opus 5 averaged 2.05 hypothesis revisions per episode and GPT-5.6 Sol 1.60 (Supplementary Table S8B). Candidate-set size was not monotonically related to score. Opus 5 and GPT-5.6 Sol nominated 3.7 and 3.1 proteins per episode, Sonnet 5 1.7 and Devstral Small 0.2, whereas Haiku 4.5 (maximum 156), GPT-OSS-20B (maximum 2,822) and Qwen3-8B (maximum 2,184) submitted far longer lists in a few episodes.

S3.5 Use of experiments

Of the 360 episodes with an experimental budget, 145 (40.3%) requested an experiment and 143 (39.7%) received at least one. In total, 413 experiments were delivered. Thirty-nine requests were refused in these episodes, and 11 further requests were refused in observational-only episodes. Most deliveries were in vivo knockdowns rather than cell perturbations (377 of 413, 91.3%). Opus 5, GPT-5.6 Sol and Sonnet 5 accounted for 316 of the 413 deliveries (76.5%) while representing 120 of the 360 budget-eligible episodes (33.3%).

Individual trajectories showed adaptive use of experimental results. For example, in a Haiku 4.5 episode under the expanded experimental budget, the agent bought knockdowns of three candidates in one turn for $1.2M, after association screening, adjusted logistic regression and genetic-effect ratios, and then retained the two candidates with favorable intervention responses, removed the candidate with no effect and scored 75.

The format of the returned experimental results created difficulty for some agents. Among the 143 episodes that received at least one experiment, agents misread the returned result format in 16 (11%), including 9 of 18 Haiku 4.5 episodes (50%) and 4 of 6 Qwen3-8B episodes (67%), compared with none of 39 Opus 5 episodes and 1 of 39 GPT-5.6 Sol episodes (3%).

Relative to the observational-only budget, the expanded budget changed mean scores by +0.13 for Opus 5 (95% CI −8.71 to 8.98), +0.07 for GPT-5.6 Sol (−10.08 to 10.22), +9.86 for Sonnet 5 (−2.40 to 22.12), +10.74 for GPT-OSS-20B (1.95 to 19.53) and +8.19 for Qwen3-Coder-30B-A3B (3.02 to 13.35), with t-based intervals as in Supplementary Table S6.

S3.6 Failure modes

Of the 416 episodes with zero phenotype credit, 154 (28.5% of all 540 episodes) produced no phenotype at the required submission path, and 262 (48.5%) produced a phenotype that earned no credit. The remaining 124 episodes received positive phenotype credit.

Devstral Small reached the 30-turn limit in 53 of 60 episodes, GLM-4-32B in 38 and Qwen3-8B in 14. A reply without a code fence that contains code-like text is executed whole, so prose placed before the code produces a syntax error. At least one reply contained no code in 88 of 540 episodes.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI