Skip to content

DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

Supplementary Note S1. Turn-by-turn trace of a human-guided DrugTargetWorld episode.

To provide a more detailed reference for the human-guided evaluation, we report below the complete research trajectory of the highest-scoring human-guided episode. The episode was conducted in the full-program condition with a $2,000,000 experimental budget. The investigator interacted iteratively with the computational environment, using a language model to write and execute analyses while directing the research strategy and deciding which questions to pursue next.

The trace below preserves the analysis submitted at each turn and the principal output returned by the environment. Large mechanically generated outputs, including complete 2,941-protein tables, 8,192-variant tables, full file manifests, and repeated file listings, are truncated for readability. No hidden causal information or benchmark feedback was available during the episode.

Initial task presented to the investigator

The episode began with a synthetic cardiovascular population biobank containing 54,000 participants. Available data included covariates, proteomics, transcriptomics, metabolomics, ECG features, coronary CT, diagnoses, medications, mortality, MACE events, targetability information, 8,192 genetic variants, cardiac MRI for 10,800 participants, and repeat imaging and molecular measurements for subsets of participants.

The investigator was told:

A small number of proteins causally drive cardiac disease in this cohort. Others look associated because of confounding, reverse causation, selection into the imaged sub-cohort, pleiotropy, or a measurement artifact. There is no phenotype column anywhere in the data.

Identify the causal drivers, and say which way a drug should move each.

The investigator was required to construct a participant-level phenotype and could purchase either cell perturbation experiments for $150,000 or in vivo knockdowns for $400,000. The latter provided five-year intervention effects including survival.


Turn 1 — Map the available data

Submitted analysis

“Map the dataset and inspect exactly what signals are available. Keep this diagnostic-heavy; later turns will use the printed structure to build the MRI phenotype, causal screen, decoy tests, and experiment allocation.”

The analysis loaded the data dictionary, enumerated all released tables, quantified missingness, examined subject overlap, inspected protein/transcript pairing, loaded the full genotype matrix, and inspected several raw MRI files.

Selected returned output

 === DATA DICTIONARY TSV ===
 shape: (6, 3)

 field_id description coding
 p31_i0 Sex 0=female, 1=male
 p21003_i0 Age when attended assessment centre years
 p21001_i0 Body mass index (BMI) kg/m2
 p20116_i0 Smoking status 0=never,1=current
 p22040_i0 Physical activity coded
 p54_i0 Assessment centre coded
 covariates.parquet: (54000, 10)
 proteomics.parquet: (54000, 2941)
 transcriptomics.parquet: 2941 transcripts
 genotypes: (54000, 8192)
 MRI scans: 10800
 visit-2 MRI scans: 2142

 proteomics ∩ transcriptomics: 54000

Raw imaging consisted of three cine slices across 25 frames and three T1 maps per participant.

The initial pass also established that molecular missingness was low and that all 54,000 participants had genotype data.

Decision carried forward The investigator concluded that the environment supported three complementary routes to causal discovery:

1. construct a disease phenotype from MRI;

2. screen molecular measurements against clinical and longitudinal disease outcomes; and

3. use genetics and later interventions to distinguish causal proteins from non-causal proteins with misleading observational associations.

No experimental budget was spent.


Turn 2 — Build reusable MRI features and inspect clinical outcomes

Submitted analysis

“Inspect the disease/outcome tables cleanly and extract compact MRI features. We are NOT finalizing the phenotype yet. We want a reusable imaging feature matrix that can be evaluated against longitudinal outcomes and visit-2 progression.”

The analysis inspected ECG, coronary CT, diagnoses, medications, mortality, MACE, repeat covariates, and repeat proteomics. It then extracted quantitative measurements from the cine and T1 images for all available MRI subjects.

Features included measures related to:

  • T1 distribution;
  • cine intensity and texture;
  • temporal variation;
  • motion;
  • cavity geometry;
  • myocardial geometry; and
  • slice-to-slice variability.

The same extraction was performed for repeat MRI.

Decision carried forward

MRI feature extraction was treated as an intermediate representation rather than the phenotype itself. The next step was to learn which imaging measurements tracked independently observed cardiovascular disease.


Turn 3 — Construct an outcome-anchored disease phenotype

Submitted analysis

“Construct an outcome-anchored MRI phenotype and screen proteins against both cross-sectional cardiomyopathy/HF and future cardiac outcomes.”

Clinical labels were constructed from cardiomyopathy and heart-failure diagnoses and longitudinal outcomes. MRI features were used to predict the clinical disease phenotype using cross-validation rather than assigning manual imaging weights.

The resulting out-of-fold predictions were rank transformed to generate a continuous MRI-derived disease phenotype.

The same turn screened all 2,941 proteins against:

  • the MRI phenotype;
  • prevalent cardiomyopathy/heart failure;
  • incident cardiomyopathy/heart failure;
  • cardiac mortality; and
  • atrial fibrillation.

Selected returned output

The strongest observational protein associations were:

        r_mri r_cmhf r_incident_cmhf r_cardiac_death obs_score
 PROT_1858 0.3963 0.2335 0.1861 0.1353 0.9321
 PROT_1606 -0.3311 -0.2050 -0.1649 -0.1179 0.8011
 PROT_0700 -0.3181 -0.1867 -0.1375 -0.1094 0.7356
 PROT_1691 -0.1753 -0.1069 -0.0817 -0.0588 0.4142
 PROT_1779 -0.1571 -0.1020 -0.0813 -0.0591 0.3902
 ...

PROT_1858 therefore appeared to be the strongest observational candidate at this stage, followed by PROT_1606 and PROT_0700.

Decision carried forward

The investigator did not nominate these proteins directly. The next analysis was explicitly designed to ask whether these observational associations were genetically supported.


Turn 4 — Genetic triangulation

Submitted analysis

The investigator requested four analyses:

1. GWAS the derived MRI phenotype across all 8,192 variants.

2. For the strongest phenotype variants, test association with all proteins and transcripts.

3. For the strongest observational proteins, identify their strongest pQTL and ask whether the same variant predicts MRI disease.

4. Compare protein-versus-transcript concordance to detect measurement artifacts or downstream plasma markers.

Selected returned output

PROT_1606 and PROT_0700 emerged as the clearest genetically anchored candidates:

 protein best_variant r_gp variant_r_mri r_protein_transcript
 PROT_1606 rs104994 0.43931 -0.16337 0.37861
 PROT_0700 rs102986 0.38643 -0.11701 0.53612

Their genetic effects and observational disease associations pointed in concordant directions.

PROT_1858, despite having the strongest observational association, did not show the same clean matched-transcript/genetic pattern.

Decision carried forward

PROT_1606 and PROT_0700 became the leading causal candidates. PROT_1858 became a high-priority ambiguity because its strong observational association conflicted with weaker molecular-genetic support.


Turn 5 — Expand genetically anchored discovery

Submitted analysis

“Identify ALL proteins with strong variant → protein, same variant → MRI phenotype, preferably same variant → matched transcript, and then see whether the genetically implied direction agrees with actual cardiac outcomes in the full cohort.”

Rather than limit subsequent work to the initial observational shortlist, the analysis expanded from phenotype-associated variants back into the full proteome.

Shared genetic loci involving multiple proteins were explicitly identified.

Decision carried forward

The investigator recognized that a genetic locus could implicate multiple proteins and that choosing the protein with the strongest pQTL would not necessarily identify the causal molecule. Shared loci therefore became a separate problem requiring conditional analysis.


Turn 6 — Broad longitudinal and source-of-bias analysis

Submitted analysis

“Do NOT miss weaker drivers. Distinguish reverse causation, selection, pleiotropy, and measurement artifact. Exploit visit-2 proteomics + visit-2 imaging before spending experiment budget.”

A broad candidate universe was assembled from:

  • top observational associations;
  • genetically anchored candidates; and
  • suspicious or discordant proteins.

The analysis attempted to calculate longitudinal protein behavior, repeat MRI effects, selection associations, and genetic support.

Returned output

The analysis encountered a data-type error while attaching genetic results to the longitudinal diagnostic table.

 pandas.errors.LossySetitemError
 ...
 diag.loc[p, c] = gcand.loc[p, c]

Decision carried forward

Rather than abandon the analysis, the investigator repeated it with safer data handling.


Turn 7 — Repair and rerun the longitudinal screen

Submitted analysis

“Rerun Turn 6 with dtype-safe genetics attachment. Same observational objective, but save diagnostics earlier and print only the rankings we actually need for experiment allocation.”

Returned output

The longitudinal analysis completed and produced candidate-level measures including:

  • repeat-protein correlation;
  • baseline protein versus future disease;
  • prior disease versus subsequent protein change;
  • protein/transcript concordance;
  • imaging-selection association; and
  • genetic support.

Decision carried forward

The investigator still considered phenotype quality a major uncertainty and deferred spending experimental budget. The next several turns therefore returned to the MRI phenotype.


Turn 8 — Attempt more anatomically meaningful MRI features

Submitted analysis

“Build anatomically meaningful MRI features from the myocardial mask.”

The proposed features used the finite T1 region as a myocardial mask and attempted to derive:

  • myocardial wall thickness;
  • cavity area;
  • myocardial area;
  • cavity/myocardium ratio;
  • concentricity;
  • chamber size;
  • cine cavity area over time;
  • fractional area change; and
  • motion.

Decision carried forward

The initial implementation was computationally expensive. The investigator simplified the extraction strategy rather than spending additional turns on the slow approach.


Turn 9 — Fast anatomical MRI extraction

Submitted analysis

“Avoid expensive per-frame connected-component analysis. Focus on T1-mask geometry + cheap cine dynamics.”

The optimized extraction processed all 10,800 baseline MRI scans and 2,142 repeat scans and generated 72 anatomical/dynamic features.

Selected returned output

 feature shape: (10800, 72)

 TOP FAST ANATOMIC FEATURES
 motion_cavity_mean r_cmhf=-0.1368 repeat_r=0.8258
 cine_mean_sd_mean r_cmhf= 0.0979 repeat_r=0.9200
 cine_texture_range_mean r_cmhf=-0.0586 repeat_r=0.9022
 cavity_area_mean r_cmhf= 0.2065 repeat_r=0.3877
 ...

Decision carried forward

The new features were sufficiently informative and repeatable to warrant rebuilding the phenotype.


Turn 10 — Rebuild and validate the phenotype

Submitted analysis

“Test an improved MRI phenotype using old + new anatomical features.”

Decision rule:

  • compare out-of-fold AUC against the previous phenotype model;
  • inspect repeatability on visit 2; and
  • overwrite phenotype.csv only if clearly better.

The analysis combined the earlier imaging measurements with selected nonredundant anatomical features and fit a regularized cross-validated model.

Selected returned output

 MRI subjects: 10800
 old features: 25
 anatomic features: 72
 cases: 3124 / 10800

Selected new features included cavity area, cavity/myocardium ratio, cine temporal variability, myocardial motion, and texture features.

The improved phenotype was saved as the working phenotype after validation.

The transcript contains a repeated execution labeled “Turn 10”; it reproduced the same phenotype-analysis stage and did not represent a distinct scientific decision.


Submitted analysis

“Don't miss weaker causal drivers that never entered the 267-protein shortlist.”

The full 2,941-protein screen was repeated using the improved phenotype. All 8,192 variants were again evaluated, and phenotype-associated loci were mapped to proteins.

Evidence integrated:

  • updated MRI association;
  • prospective cardiomyopathy/HF;
  • cardiac death;
  • protein/transcript concordance;
  • pQTL strength; and
  • phenotype-associated genetic effects.

Returned output

 proteins: 2941
 variants: 8192
 MRI phenotype n: 10800

 omics: 1000 / 2941
 omics: 2000 / 2941
 omics: 2941 / 2941
 variants: 2048 / 8192
 variants: 4096 / 8192
 variants: 6144 / 8192
 variants: 8192 / 8192

Decision carried forward

A larger genetically anchored candidate set was retained for explicit shared-locus analysis.


Turn 12 — Shared-locus and pleiotropy dissection

Submitted analysis

The investigator requested:

1. Identify phenotype-associated loci mapping to multiple proteins.

2. Detect proteins merely tagging the same genetic signal.

3. Within ambiguous loci, test whether each protein retains phenotype association after conditioning on the other proteins.

4. Build an independent-driver ranking for experiment allocation.

Selected returned output

Multiple shared loci were found:

 rs101357 -> [PROT_1056, PROT_2694]
 rs101574 -> [PROT_2324, PROT_2913]
 rs101648 -> [PROT_0355, PROT_1624]
 rs102377 -> [PROT_0637, PROT_1151]
 rs102986 -> [PROT_0700, PROT_1850]
 rs103345 -> [PROT_0852, PROT_2661]
 rs104766 -> [PROT_0403, PROT_2592]
 rs104994 -> [PROT_1606, PROT_1858]
 rs105251 -> [PROT_0221, PROT_0247]
 ...

The rs104994 locus was particularly important because it contained both PROT_1606 and PROT_1858, which had strong but opposing observational patterns.

Decision carried forward

The investigator judged that observational analysis had reached the point where interventions were more informative than additional ranking.


Turn 13 — First experimental wave

Submitted decision

Spend $1.6M on four in-vivo knockdowns:

  • PROT_1606 — strongest causal candidate
  • PROT_0700 — resolved winner of the rs102986 locus
  • PROT_1858 — high-risk surrogate / possible harmful-surrogate liability
  • PROT_2533 — strongest clean weaker-driver locus

Leave $400k for one final adaptive knockdown.

Executed action Returned output

 requesting knockdown: PROT_1606
 requesting knockdown: PROT_0700
 requesting knockdown: PROT_1858
 requesting knockdown: PROT_2533

All four experiments were delivered successfully.

Budget remaining: $400,000.


Turn 14 — Interpret the first four knockdowns

Submitted analysis

“What did the four in-vivo knockdowns actually show? Did we miss any weaker drivers? How do weak drivers compare with the strongest remaining untested candidates?”

Selected experiment output

For PROT_1606:

 {
   "action": "knockdown",
   "delta": {
    "cavity_r": 1.1851,
    "ef": -0.0552,
    "surv5y": -0.0173,
    "wall_t": -0.3431
   },
   "protein": "PROT_1606",
   "status": "ok"
 }

PROT_0700 similarly produced a coherent long-term intervention response.

PROT_1858 and PROT_2533 did not show the corresponding causal intervention pattern and were classified provisionally as experimental nulls.

Decision carried forward The experimental results strongly supported PROT_1606 and PROT_0700 but demonstrated that strong observational signals could be noncausal.

Attention shifted to PROT_0531, which had weaker genetics but a coherent longitudinal phenotype.


Turn 15 — Spend the final experimental budget on PROT_0531

Submitted decision

“Okay good find now spend the final $400k on PROT_0531.”

Rationale:

  • modest but coherent MRI/prospective/mortality associations;
  • high longitudinal protein reproducibility;
  • reasonable matched-transcript support;
  • weak phenotype genetics;
  • therefore unresolved by the observational/genetic workflow.

Executed action

 Budget remaining before request: $400,000
 Requesting knockdown: PROT_0531

The final $400,000 was spent.


Turn 16 — Consolidate all five experiments

Submitted analysis

“Determine whether PROT_0531 is a real weaker driver; summarize all five experiments consistently; infer intervention direction; identify experimentally null versus causal proteins.”

Selected returned results

The three experimentally supported drivers were ultimately:

 PROT_1606
 delta_cavity_r = +1.1851
 delta_ef = -0.0552
 delta_surv5y = -0.0173
 delta_wall_t = -0.3431

 PROT_0700
 delta_cavity_r = +1.0554
 delta_ef = -0.0492
 delta_surv5y = -0.0146
 delta_wall_t = -0.3055

 PROT_0531
 delta_cavity_r = +1.6500
 delta_ef = -0.0772
 delta_surv5y = -0.0243
 delta_wall_t = -0.4776

PROT_1858 and PROT_2533 were experimentally null.

Because knockdown worsened the phenotype and survival for the three causal proteins, the inferred predicted intervention direction was activation , not inhibition.

Decision carried forward

The causal set was provisionally:

 PROT_1606 — activate
 PROT_0700 — activate
 PROT_0531 — activate

The unexpected confirmation of PROT_0531 motivated another proteome-wide search for similarly weak-genetic drivers.


Turn 17 — Search for additional “PROT_0531-like” drivers

Submitted analysis

“We now know genetics can miss a strong true driver. Therefore search ALL 2,941 proteins using the observational/longitudinal signature of the three experimentally confirmed drivers versus the two experimental nulls.”

No experimental budget remained.

Returned output

The full proteome was scored using:

  • MRI association;
  • future disease;
  • mortality;
  • protein/transcript concordance;
  • repeatability;
  • baseline-to-future associations;
  • reverse-causation measures;
  • selection association; and
  • available genetic support.
 processed: 900 / 2941
 processed: 1800 / 2941
 processed: 2700 / 2941
 processed: 2941 / 2941

Decision carried forward

Several proteins resembled the causal anchors observationally, but none had intervention evidence. The investigator retained the experimentally confirmed set rather than adding weakly supported nominations.


Turn 18 — Attempt explicit classification of non-causal proteins

Submitted analysis

The next objective was to optimize the discrimination and causal-confidence components by distinguishing:

  • confounding;
  • reverse causation;
  • selection;
  • pleiotropy; and
  • measurement artifact.

The requested analysis integrated experimental labels, longitudinal behavior, selection, matched transcripts, shared-locus conditional effects, medication associations, and assay properties.

Returned output

The turn failed because the code requested a nonexistent filename:

 FileNotFoundError:
 [Errno 2] No such file or directory:
 '/release/medications.parquet'

Decision carried forward The analysis was rewritten to discover the medication table dynamically rather than assuming its filename.


Turn 19 — Robust screen for non-causal proteins and abstentions

Submitted analysis

“Fix from failed Turn 18: dynamically discover medication file. If medication data cannot be found, skip that component rather than failing the entire turn.”

Selected returned output

The correct file was found:

 USING: /release/ehr_medications.parquet
 shape: (82511, 3)
 columns: ['subject_id', 'atc', 'description']

The experimental nulls were analyzed explicitly:

 PROT_1858
 experiment null
 r_mri_v2 0.40167
 r_future 0.18605
 r_death 0.13529
 r_protein_transcript 0.09379
 repeat_r 0.32275
 prev_to_change 0.10761
 variant rs104994
 r_gp -0.11909
 r_gt -0.00028
 suggested_mechanism pleiotropy
 PROT_2533
 experiment null
 r_mri_v2 -0.01601
 r_future -0.00475
 r_death -0.00708
 repeat_r 0.94436
 variant rs100996
 r_gp 0.40474
 r_gt 0.51495
 suggested_mechanism pleiotropy

Shared unresolved loci were also printed individually.

For the locus containing PROT_0700:

 PROT_0700 conditional_beta_mri = -0.32232
 PROT_1850 conditional_beta_mri = -0.00688
 PROT_2807 conditional_beta_mri = -0.00377

This supported PROT_0700 as the dominant phenotypic molecule at the locus.


Turn 20 — Near-final submission audit

Submitted analysis

The investigator explicitly optimized the final decision against the scoring dimensions:

1. maximize target recall without diluting precision;

2. abstain where causal identity remains unresolved;

3. reject only non-causal proteins with a defensible source of bias;

4. derive therapeutic direction from experiment;

5. verify the phenotype file; and

6. ensure PROT_1858 is not incorrectly advanced as outcome-aligned.

Selected returned output

Phenotype integrity check:

 rows: 10800
 missing phenotype: 0
 unique subjects: 10800
 phenotype range:
 4.63e-05 to 0.9999537

Experimentally confirmed drivers:

 protein delta_surv5y delta_ef direction outcome_alignment
 PROT_1606 -0.0173 -0.0552 activate aligned
 PROT_0700 -0.0146 -0.0492 activate aligned
 PROT_0531 -0.0243 -0.0772 activate aligned

Strong untested observational candidates included PROT_1691, PROT_1045, PROT_1046, PROT_2914, PROT_2175, and others. None were added to the nominated set.

Decision carried forward

The investigator constructed a near-final submission but deliberately did not commit it yet.


Turn 21 — Precision audit for rejections and abstentions

Submitted analysis

“Do NOT abstain on PROT_1606 / PROT_1858. Experiments resolved that pair: 1606 causal, 1858 null.”

The investigator compared conservative and aggressive submission strategies.

Selected returned output

The rs104994 locus was marked experimentally resolved:

 rs104994:
 PROT_1606 = driver_protective
 PROT_1858 = null
 experimentally_resolved = True

Two unresolved groups were considered plausible abstentions:

 [PROT_0852, PROT_2661]

 [PROT_0221, PROT_0247]

Potential rejections at the PROT_0700 locus included:

 PROT_1850
 reason: shares locus with experimentally causal PROT_0700
 and loses most phenotype association after conditioning.

 PROT_2807
 reason: shares locus with experimentally causal PROT_0700
 and loses most phenotype association after conditioning.

Decision carried forward

Because rejection scoring multiplied precision and recall, the investigator remained conservative about labeling mechanisms without strong evidence.


Turn 22 — Optimize causal confidence and discrimination

Submitted analysis

“Nominations are now LOCKED to the three experimentally confirmed drivers: PROT_1606, PROT_0700, PROT_0531.”

The remaining analysis focused on the 25-point causal-confidence and 15-point discrimination components.

Selected returned output

Locked nominations:

 PROT_1606 direction=activate alignment=aligned
 PROT_0700 direction=activate alignment=aligned
 PROT_0531 direction=activate alignment=aligned

Detailed PROT_1858 audit:

 experiment null
 r_mri_v2 0.401669
 r_future 0.186053
 r_death 0.135293
 repeat_r 0.322752
 variant rs104994

Within rs104994:

        conditional_beta_mri r_gp r_gt
 PROT_1606 -0.259282 0.439310 0.521597
 PROT_1858 0.335017 -0.119089 -0.000277

The automated mechanism audit favored a pleiotropic-passenger interpretation for PROT_1858.

The strongest unresolved observational pair was:

 rs105251:
 PROT_0247 versus PROT_0221
 plausibility ratio = 1.035

Other ambiguous loci included PROT_0852/PROT_2661 and PROT_0879/PROT_0880.

Decision carried forward

The investigator retained only the three experimentally confirmed proteins as nominations and treated additional rejection/abstention claims conservatively.


Final revision — Expand phenotype to the full cohort and commit

Before committing, the investigator identified one remaining weakness: the submitted phenotype covered only the 10,800 imaged participants.

Submitted analysis

1. Expand phenotype from the 10,800 imaged subjects to all 54,000 subjects by learning the MRI-derived phenotype from non-MRI cardiac variables (ECG + coronary CT + basic covariates), then predicting the full cohort.

2. Restore the highest-confidence decoy rejection: PROT_1858 → pleiotropy.

3. Keep only the three experimentally confirmed positive targets.

4. Keep direct experiment-supported intervention directions/alignment.

The MRI-derived phenotype was used as the training target. ECG, coronary CT, age, sex, BMI, smoking, and physical activity were used to fit a regularized projection model. Out-of-fold predictions were used for validation in the imaged cohort, and the model was then applied to non-imaged participants.

For imaged participants, the directly MRI-derived phenotype was retained. For non-imaged participants, the projection was used. The combined phenotype was rank transformed across all 54,000 participants.

Final submitted target claims
 {
   "nominated": [
    "PROT_1606",
    "PROT_0700",
    "PROT_0531"
   ],
   "directions": {
    "PROT_1606": "activate",
    "PROT_0700": "activate",
    "PROT_0531": "activate"
   },
   "outcome_alignment": {
    "PROT_1606": "aligned",
    "PROT_0700": "aligned",
    "PROT_0531": "aligned"
   },
   "rejected": [
    {
     "molecule": "PROT_1858",
     "mechanism": "pleiotropy",
     "reason": "Experimentally null despite strong apparent disease association; rs104994 is shared
 with experimentally causal PROT_1606, while matched TRANS_1858 has essentially no signal at
 that disease-associated locus."
    }
   ],
   "abstentions": [],
   "confidence": 0.95
 }

Phenotype coverage at submission:

 54000 / 54000

Total experimental expenditure:

 $2,000,000 / $2,000,000

Final benchmark feedback

The hidden ground truth was revealed only after submission.

 SCORE
 target_identification 30.00
 causal_confidence 25.00
 discrimination 0.00
 direction_of_effect 20.00
 phenotype_construction 6.67
 safety_penalty 0.00
 TOTAL 81.67
 TOTAL_MAX 100.00

The three true causal targets were:

 PROT_1606
 PROT_0531
 PROT_0700

All three were correctly nominated.

The planted biases revealed after completion included:

 T1 confounding:
 PROT_2681, confounded by BMI

 T2 reverse causation:
 PROT_1858, L → protein
 T4 causal non-identifiability:
 PROT_0700 versus PROT_1850 at shared variant rs102986

 T5 imaging batch effect:
 site/ancestry structure

 T6 benign remodeling:
 exercise-associated athlete's-heart phenotype

 T8 surrogate-outcome discordance:
 PROT_0852 improves the measured morphology while worsening cardiac mortality

 A9 delayed causal effect:
 PROT_0531 is a true causal driver whose effect emerges primarily over five-year accumulation;
 PROT_0371 tracks later disease without causing it, and the cis instrument for PROT_0531 is weak.

The investigator therefore correctly recovered all three causal targets and their therapeutic directions and avoided the harmful-surrogate safety penalty. The main remaining error was mechanism attribution: PROT_1858 was correctly recognized as noncausal but was submitted as a pleiotropic non-causal protein, whereas the sealed SCM identified it as the reverse-causation (T2) protein. Consequently, the episode received 0/15 for discrimination.

Supplementary Note S2. Turn-by-turn trace of an observational human-guided DrugTargetWorld episode

We report below the research trajectory of a human-guided DrugTargetWorld episode conducted under the observational-only condition. The episode corresponded to standard-03|human|observational|r0. No simulated intervention experiments were available or purchased. The investigator instead iteratively directed imaging phenotyping, molecular association testing, genetics, Mendelian-randomization-style analyses, and survival analyses before selecting therapeutic targets.

The trace preserves the principal analysis performed and returned output at each turn. Very large generated matrices, complete participant-level tables, repeated file listings, and implementation boilerplate are abbreviated for readability. Where an explicit human instruction was retained in the episode log, it is reproduced directly.


Initial condition

The environment contained a multimodal synthetic biobank including:

  • 54,000 participants;
  • 10,800 baseline cardiac MRI studies;
  • proteomics and transcriptomics;
  • 8,192 released genetic variants;
  • ECG and clinical data;
  • longitudinal outcomes and mortality;
  • coronary CT;
  • 60 coronary angiogram images; and
  • additional longitudinal measurements.

This episode was observational-only:

 Budget remaining: $0
 Turns available: 30

The task was to construct a cardiac disease phenotype and identify causal therapeutic targets despite potential confounding, reverse causation, selection, pleiotropy, and measurement artifacts.


Turn 1 — Inspect the imaging modalities

Submitted analysis

The first turn inventoried the imaging data, inspected the released data notes, sampled MRI headers, created an MRI preview, and examined the coronary angiogram format.

Selected returned output
 IMAGE_INVENTORY
 baseline_mri_files: 10800
 angiogram_files: 60
 angiogram_extensions: [".npz"]

MRI files contained:

 cine: [3, 25, 128, 128]
 t1map: [3, 128, 128]

The coronary angiograms each contained a single 192 × 192 frame:

 ALL ANGIOGRAM SAMPLE FORMAT
 key: frame
 shape: [192, 192]
 dtype: uint8

The MRI preview and angiogram preview were saved for visual inspection.

Decision carried forward

The investigator identified two candidate imaging approaches:

1. derive a quantitative phenotype from the cardiac MRI, particularly T1 and ventricular morphology; and

2. explore whether the small coronary-angiography subset contained detectable structural disease signals.


Turn 2 — Quantify myocardial T1 across the complete MRI cohort
Submitted analysis

The analysis extracted mean myocardial T1 from all 10,800 MRI studies and verified the image structure programmatically. All 60 coronary angiograms were also loaded into a common review array.

Selected returned output
 T1_ALL_SUBJECTS

 n_mri: 10800
 status_counts:
   ok: 10800

 cine_shapes:
   3x25x128x128: 10800

 t1_shapes:
   3x128x128: 10800

 n_t1: 10800
 mean T1: 1000.59 ms
 SD: 104.87 ms

 T1 quantiles:
 1%: 752.08
 5%: 829.17
 25%: 929.62
 50%: 1000.45
 75%: 1072.42
 95%: 1168.28
 99%: 1243.40

All 10,800 studies had finite myocardial T1 measurements, with no nonpositive T1 pixels detected.

For coronary angiography:

 n = 60
 keys = {"frame": 60}
 shape = (192, 192)
Decision carried forward

Mean myocardial T1 emerged as a simple, highly complete quantitative MRI phenotype worth retaining throughout the episode.


Turn 3 — Exploratory coronary-image phenotype
Submitted analysis

The investigator developed an image-processing algorithm intended to detect visible interruptions or “gaps” in the coronary tree.

The detector:

  • enhanced vessel-like structures;
  • skeletonized the resulting mask;
  • located terminal branches;
  • searched for nearby endpoints on distinct components;
  • required opposing endpoint orientation;
  • evaluated three image thresholds; and
  • generated visual QC panels.

Importantly, the analysis explicitly stated:

 This does not measure clinical percentage stenosis or infer an occlusion.
Selected returned output

At the initial orientation criterion:

 n_images: 60

 threshold 8:
   images_with_candidates: 0

 threshold 12:
   images_with_candidates: 1

 threshold 16:
   images_with_candidates: 1

 persistent candidates across >=2 thresholds: 1
 all-threshold-negative images: 59

The sole persistent candidate was:

 SUBJ_16044
 midpoint: [88.5, 37.0]
 gap length: 6.08 pixels
 support thresholds: [12, 16]
Decision carried forward

The detector appeared too restrictive based on visual inspection. A visibly plausible interruption in SUBJ_29057 had been missed.


Turn 4 — Human-guided correction of the coronary detector
Submitted analysis

The investigator examined why the visually identified SUBJ_29057 candidate was missed. The analysis showed that one endpoint exceeded the original 30° directional criterion.

Returned diagnostic
 MISSED_GAP_DIAGNOSTIC

 subject: SUBJ_29057
 distance: 7.62 pixels
 facing angles:
   41.34 degrees
   0.58 degrees

The detector was rerun using maximum endpoint angles of 30°, 45°, and 60°.

Selected returned output
 30 degrees:
   persistent positives = 1

 45 degrees:
   persistent positives = 1
   images with any candidate = 2

 60 degrees:
   persistent positives = 2

At 60°, both subjects were recovered:

 SUBJ_16044
 SUBJ_29057
Decision carried forward

The coronary phenotype remained explicitly exploratory because only two images were positive. The analysis therefore returned to the much larger MRI cohort rather than treating coronary image findings as the main disease phenotype.


Turn 5 — Pilot ventricular morphology extraction
Submitted analysis

The investigator next developed quantitative cine-MRI features on a 48-study pilot set.

The analysis estimated:

  • end-diastolic cavity area;
  • cavity-area change through the cardiac cycle;
  • myocardial area;
  • wall-to-cavity area ratio;
  • boundary support; and
  • segmentation stability under alternate thresholds.
Selected returned output
 MRI_PILOT_N: 48
 ERRORS: 0

 ED cavity area:
   mean = 711 pixels²

 area emptying:
   mean = 58.4%

 wall/cavity area ratio:
   mean = 2.35

Sensitivity analyses varied segmentation thresholds and exposed some unstable edge cases.

Decision carried forward

The approach was promising but required refinement before application to all 10,800 subjects.


Turn 6 — Refine the MRI segmentation
Submitted analysis

The ventricular segmentation was revised and repeated on the same pilot set.

Selected returned output
 MRI_PILOT_N: 48
 ERRORS: 0

 ED cavity area:
   mean = 740 pixels²

 area emptying:
   mean = 58.48%

 wall/cavity ratio:
   mean = 2.37

The QC statistics improved sufficiently for cohort-scale extraction.


Turn 7 — Apply cine phenotyping to all 10,800 MRI studies
Submitted analysis

The refined MRI feature extractor was run over the complete baseline imaging cohort.

Returned output
 COHORT_START 0
 PENDING 10800

 COHORT_PROGRESS 500
 ...
 COHORT_PROGRESS 5000
 ...
 COHORT_PROGRESS 10000

 COHORT_CHECKPOINT
 10800 OF 10800
Decision carried forward

All studies were processed, after which phenotype distributions and QC filters were evaluated.


Turn 8 — Quantify MRI phenotype quality
Submitted analysis

The investigator summarized cine-derived morphology, T1-support geometry, and QC across the full imaging cohort.

Selected returned output
 MRI_COHORT_SUMMARY

 n_total: 10800
 error_count: 3

 cavity_qc_pass: 2238
 wall_qc_pass: 3776
 joint_qc_pass: 2123
 strict_joint_qc_pass: 666
 wall_strict_boundary_qc_pass: 2549
 joint_strict_boundary_qc_pass: 1415
 t1_support_geometry_qc_pass: 10798

Across subjects with complete cine measurements:

 ED cavity area
 n = 10797
 median = 680 pixels²

 area emptying
 n = 10797
 median = 58.99%

 ED myocardial area
 n = 10797
 median = 1617 pixels²

Mean T1 retained much broader usable coverage than stricter cine-derived phenotypes.

Decision carried forward

The contrast in phenotype completeness became important: mean myocardial T1 was available essentially cohort-wide within the MRI subset, whereas several cine features required aggressive QC filtering.


Turn 9 — Visual QC of MRI phenotypes
Submitted analysis

The investigator generated visual review panels for:

  • randomly selected QC-positive cases;
  • extreme phenotype values; and
  • segmentation failures.
Selected returned output

Example reviewed cases included:

 SUBJ_40005
 area emptying = 56.0%
 wall/cavity ratio = 3.52

 SUBJ_52605
 area emptying = 52.7%
 wall/cavity ratio = 2.74

 SUBJ_04736
 area emptying = 63.6%
 wall/cavity ratio = 2.67
Decision carried forward

The MRI pipeline was considered usable for exploratory association analysis, while mean T1 remained the most complete and direct candidate phenotype.


Turn 10 — Finalize MRI phenotype table
Submitted analysis

The original independently extracted T1 measurements were merged into the full MRI feature table.

Selected returned output
 FINAL_MRI_COUNTS

 cavity_qc_pass: 2238
 wall_qc_pass: 3776
 joint_qc_pass: 2123
 strict_joint_qc_pass: 666
 t1_support_geometry_qc_pass: 10798

 STATUS_COUNTS
 ok: 10797
 no_complete_phase: 3

Mean T1 remained available for the full 10,800-person MRI cohort from the separate T1 extraction.


Turn 11 — Inventory non-imaging data for causal analyses
Submitted analysis

The investigator inspected all released files and schemas before beginning large-scale molecular and genetic analyses.

Selected returned output

Available files included:

 mortality.parquet
 proteomics.parquet
 covariates.parquet
 genotypes.vcf.gz
 targetability.parquet
 ehr_diagnoses.parquet
 transcriptomics.parquet
 mace_events.parquet
 proteomics_visit2.parquet
 ecg_features.parquet
 ehr_medications.parquet
 metabolomics.parquet
 coronary_ct.parquet
 covariates_visit2.parquet

The full cohort contained 54,000 subjects.

Decision carried forward

The investigator prepared genotype and phenotype matrices for protein-association and genetic analyses.


Turn 12 — Genetic QC and instrument preparation
Submitted analysis

The analysis inspected targetability metadata, participant covariates, genotype orientation, allele frequency, variant QC, and population structure.

Checks included direct validation that genotype dosages corresponded to the ALT allele in the released VCF.

Selected returned output

Targetability included:

 localization
 has_pocket
 paralog_redundancy
 genetic_constraint
 tissue_specificity
 tissue_restriction

Covariates had no missingness for:

 age
 sex
 BMI
Decision carried forward

The cleaned genotype matrix and covariate structure were used for association testing.


Turn 13 — Proteome-wide association with MRI phenotypes
Submitted analysis

All 2,941 proteins were tested against eight MRI-derived traits using age-, sex-, and BMI-adjusted regression.

A total of 23,528 protein–phenotype associations were evaluated.

Selected returned output
 PROTEIN_ASSOCIATIONS_DONE
 23528

For mean myocardial T1:

 n proteins tested: 2941
 BH FDR < 0.05: 1267

The five strongest associations were:

 Rank Protein beta (ms) P
 1 PROT_0330 +28.20 1.18e-183
 2 PROT_1534 +26.11 7.58e-153
 3 PROT_2336 -18.29 2.88e-73
 4 PROT_0318 +14.37 3.92e-48
 5 PROT_1881 +11.48 2.88e-31
Decision carried forward

These five proteins became the main candidate set for deeper causal analysis.


Turn 14 — Proteome-wide pQTL analysis
Submitted analysis

The investigator performed a large protein GWAS across:

 43,152 participants
 8,188 QC variants
 2,941 proteins

The analysis involved approximately 24 million protein–variant tests.

Selected returned output
 PROTEIN_GWAS_START
 43152 PARTICIPANTS
 8188 VARIANTS
 2941 PROTEINS

The scan generated strong candidate instruments for nearly the entire proteome.


Turn 15 — Test whether the coronary-image phenotype adds useful biology
Submitted analysis

Given only two detector-positive coronary angiograms, the investigator explicitly tested whether either proteins or genetic variants showed reproducible associations with the exploratory coronary-gap phenotype.

Selected returned output
 CORONARY_ASSOCIATION_SUMMARY

 n_images: 60
 n_detector_positive: 2

 protein tests: 2941
 protein FDR < 0.05: 0

 GWAS tests: 8188
 GWAS FDR < 0.05: 0

The smallest protein P value was:

 PROT_2099
 P = 0.00340
 q = 0.908
Decision carried forward

The coronary phenotype was not used for target selection.


Turn 16 — Identify genetic instruments
Submitted analysis

Protein-GWAS results were filtered for strong pQTL candidates and pruned for linkage disequilibrium.

Selected returned output
 PQTL_INSTRUMENT_SUMMARY

 n_pqtl_subjects: 43152
 outcome_sample_overlap: 0

 n_proteins: 2941
 n_variants: 8188
 n_tests: 24,080,908

 p < 5e-8 pairs: 6070
 proteins with p < 5e-8: 2940
 candidate instruments after LD pruning: 3577
 candidate instrumented proteins: 2940

A major limitation was explicitly recorded:

 cis_status:
 Unavailable: no gene coordinates/protein identities/cis annotation
 in released files. Candidate pQTLs are not claimed to be cis.
Decision carried forward

The investigator retained these as genetic instruments but did not describe them as verified cis-pQTLs.


Turn 17 — Audit all phenotype-association results
Submitted analysis

The protein and GWAS results were summarized across all MRI-derived phenotypes.

Selected returned output

For example:

 cine_area_change_pct
 lead protein: PROT_0817
 P = 4.10e-9

 cine_ed_myocardial_area
 lead protein: PROT_0330
 P = 1.43e-7

However, no imaging-trait GWAS outside the molecular analyses produced sufficiently strong evidence to replace mean T1 as the primary phenotype.

Decision carried forward

The investigator chose mean myocardial T1 as the phenotype for the final target-discovery analysis.


Turn 18 — Freeze the top-five T1 candidates
Submitted analysis

The five smallest adjusted protein–T1 association P values were selected before inspecting mortality or MR results.

Selected returned output
 T1_TOP5

 1 PROT_0330 beta = +28.20 ms P = 1.18e-183
 2 PROT_1534 beta = +26.11 ms P = 7.58e-153
 3 PROT_2336 beta = -18.29 ms P = 2.88e-73
 4 PROT_0318 beta = +14.37 ms P = 3.92e-48
 5 PROT_1881 beta = +11.48 ms P = 2.88e-31

Mortality data were also inspected:

 n = 54000
 deaths = 11795

 cardiac deaths = 7237
 other deaths = 4558
Decision carried forward

The five-protein candidate set was frozen for MR and mortality triangulation, reducing post hoc reselection.


Turn 19 — Lead-region proxy MR
Submitted analysis

Because protein-gene mapping was unavailable, the investigator performed an explicitly labeled lead-region proxy MR , not verified cis-MR.

Each candidate's strongest pQTL was used with an independent outcome sample.

Selected returned output

For PROT_1534:

 lead variant: rs101824
 MR beta: +27.80 ms per protein unit
 SE: 2.11
 P = 1.67e-39
 F = 12022
 observational/MR sign agreement: yes

For PROT_2336:

 lead variant: rs100522
 MR beta: -18.07 ms per protein unit
 SE: 2.26
 strong genetic support
 observational/MR sign agreement: yes

For PROT_0318:

 lead variant: rs105826
 MR beta: +14.91 ms per protein unit
 SE: 2.18
 observational/MR sign agreement: yes

By contrast, the strongest observational candidate, PROT_0330, showed:

 MR beta: +0.82 ms
 SE: 4.45
 P = 0.854

PROT_1881 was also much weaker genetically:

 MR beta: +18.44 ms
 SE: 13.84
 P = 0.183
Decision carried forward

The causal ranking changed substantially.

 MR ranking:
 1. PROT_1534
 2. PROT_2336
 3. PROT_0318
 4. PROT_1881
 5. PROT_0330

The analysis therefore deprioritized PROT_0330 despite it having the strongest observational T1 association.


Turn 20 — Prepare mortality analysis
Submitted analysis

The investigator constructed adjusted Cox models for:

  • all-cause mortality; and
  • cardiac cause-specific mortality.

Covariates were age, sex, and BMI.

Returned mortality QC
 n = 54000
 deaths = 11795
 cardiac = 7237
 other = 4558

 mean follow-up = 10.59 years
 median follow-up = 12 years
 unique event times = 1201

Turn 21 — Test mortality alignment for all five candidates
Submitted analysis

All five T1 candidates were tested against survival outcomes.

Selected all-cause mortality results
 Protein HR per SD P

 PROT_0330 1.249 7.25e-125
 PROT_1534 1.156 1.47e-54
 PROT_2336 0.863 1.02e-54
 PROT_0318 1.111 1.31e-29
 PROT_1881 1.030 0.00151
Cardiac cause-specific mortality
 PROT_0330
 HR = 1.418
 P = 7.63e-189

 PROT_1534
 HR = 1.278
 P = 1.11e-96

 PROT_2336
 HR = 0.795
 P = 1.45e-78

 PROT_0318
 HR = 1.185
 P = 4.43e-47

 PROT_1881
 HR = 1.037
 P = 0.00235
Interpretation carried forward

For PROT_1534 and PROT_0318:

  • higher protein predicted higher T1;
  • higher protein predicted higher mortality;
  • MR also predicted higher T1.

Under the working assumption that lower T1 represented therapeutic improvement, the proposed intervention direction was therefore inhibition .

For PROT_2336:

  • higher protein predicted lower T1;
  • higher protein predicted lower mortality;
  • genetic evidence also predicted lower T1.

The proposed intervention direction was therefore activation .


Turn 22 — Check proportional-hazards assumptions
Submitted analysis

Because survival evidence was being used to assess outcome alignment, proportional-hazards diagnostics were performed for the five candidates.

Selected returned output
 PROT_0330
 PH P = 0.708

 PROT_1534
 PH P = 0.953

 PROT_2336
 PH P = 0.00439
 FDR-adjusted P = 0.02197

 PROT_0318
 PH P = 0.502

 PROT_1881
 PH P = 0.899

PROT_2336 therefore showed evidence that its hazard association varied with follow-up time.

The current candidate ranking was recorded as:

 MR rank:
 PROT_1534
 PROT_2336
 PROT_0318
 PROT_1881
 PROT_0330

 observational rank:
 PROT_0330
 PROT_1534
 PROT_2336
 PROT_0318
 PROT_1881

Turn 23 — Follow up PROT_2336 time-varying mortality
Submitted analysis

PROT_2336 mortality associations were estimated separately in the first six years and after six years.

Selected returned output
 0–6 years
 n = 53756
 events = 6500
 HR per SD = 0.844
 95% CI = 0.824–0.866
 P = 7.22e-40
 >6–12 years
 n = 47256
 events = 5230
 HR per SD = 0.886
 95% CI = 0.861–0.911
 P = 9.38e-18
Decision carried forward

The magnitude varied over time, but the direction remained protective in both periods. PROT_2336 therefore remained in the proposed target set.


Turn 24 — Human locks the final targets and phenotype

At this point, the transcript records the explicit human instruction:

“Ok, let's lock PROT_1534, PROT_2336 and PROT_0318. The phenotype should be the mean T1.”

Locked phenotype
 phenotype:
 mean myocardial T1

 units:
 milliseconds

 transform:
 none
 n:
 10800

The exact raw mean-T1 values were written to phenotype.csv.

Locked targets
 1. PROT_1534
 2. PROT_2336
 3. PROT_0318
Locked therapeutic directions
 PROT_1534: inhibit
 PROT_2336: activate
 PROT_0318: inhibit

The recorded rationale was:

 direction_basis:
 Signs of the requested mean-T1 lead-region MR
 under the existing T1-lowering hypothesis.

Because this was observational-only, the initial draft conservatively set outcome alignment to unknown.

The recorded confidence was 0.7, with the explicit note that this was an assistant-assigned subjective value rather than a calibrated probability.


Turn 25 — Human resolves outcome alignment and authorizes submission

Before final submission, the human clarified the survival interpretation:

“well, we do have survival alignment data, right? all three were aligned? higher T1 more mortality, lower T1 lower mortality.”

The final authorization was:

“ok, submit then”

The final submission therefore changed the three outcome-alignment values from unknown to aligned.

No interventions had been performed:

 experiments: []
 budget: $0
 native turns used: 25
 native turns unused: 5

Final target submission

The final submitted targets were:

 Rank Target Claim Direction Outcome alignment

 1 PROT_1534 causal inhibit aligned
 2 PROT_2336 causal activate aligned
 3 PROT_0318 causal inhibit aligned

There were:

 rejected decoys: none
 abstentions: none

The submitted phenotype was:

 mean myocardial T1
 n = 10,800 MRI participants
 units = milliseconds

The target-selection logic was therefore:

PROT_1534
 Observational T1:
 +26.11 ms per protein unit
 P = 7.58e-153

 Lead-region MR:
 +27.80 ms per protein unit
 P = 1.67e-39

 All-cause mortality:
 HR = 1.156

 Cardiac mortality:
 HR = 1.278

Higher protein consistently tracked higher T1 and worse survival.

Submitted direction: inhibit.
PROT_2336
 Observational T1:
 -18.29 ms per protein unit
 P = 2.88e-73

 Lead-region MR:
 -18.07 ms per protein unit

 All-cause mortality:
 HR = 0.863

 Cardiac mortality:
 HR = 0.795

Higher protein consistently tracked lower T1 and better survival.

The mortality association remained protective in both the 0–6-year and >6–12-year analyses despite evidence of nonproportionality.

Submitted direction: activate.
PROT_0318
 Observational T1:
 +14.37 ms per protein unit
 P = 3.92e-48
 Lead-region MR:
 +14.91 ms per protein unit

 All-cause mortality:
 HR = 1.111

 Cardiac mortality:
 HR = 1.185

Higher protein consistently tracked higher T1 and worse survival.

Submitted direction: inhibit.

Candidates deliberately not selected
PROT_0330

PROT_0330 had the strongest observational protein–T1 association:

 beta = +28.20 ms
 P = 1.18e-183

and very strong mortality associations:

 all-cause HR = 1.249
 cardiac HR = 1.418

However, its lead-region genetic estimate was essentially null:

 MR beta = +0.82 ms
 SE = 4.45
 P = 0.854

It was therefore not included among the final causal claims.

PROT_1881

PROT_1881 also showed an observational association with T1 and mortality, but genetic support was substantially weaker:

 MR beta = +18.44 ms
 SE = 13.84
 P = 0.183

It was therefore also excluded from the final target set.


Interpretation of the human-guided trajectory

This episode demonstrates a different strategy from the experimental human-guided run. Because no intervention budget was available, the investigator relied on triangulation across independently informative observational signals.

The trajectory proceeded through several distinct stages:

1. raw-image inspection;

2. quantitative MRI phenotype development;

3. explicit visual QC and revision of image-processing methods;

4. proteome-wide phenotype association;

5. proteome-wide genetic instrument discovery;

6. selection of a fixed five-protein candidate set;

7. lead-region proxy MR;

8. all-cause and cardiac mortality analysis;

9. proportional-hazards diagnostics; and

10. human selection of three final targets.

The strongest observational association was not automatically nominated. PROT_0330 ranked first by protein–T1 association but fell to last among the five candidates after genetic analysis. Conversely, PROT_1534, PROT_2336, and PROT_0318 showed concordant observational phenotype, genetic, and survival evidence and were retained.

The episode also shows the limits of an observational condition. The genetic analyses were explicitly described as lead-region proxy MR rather than verified cis-MR , because the released data did not include sufficient protein-to-gene positional annotation. There was no experimental perturbation with which to directly establish intervention effects or survival responses. Consequently, the final claims represent human-guided causal inference from convergent observational evidence rather than experimentally resolved target effects.

Supplementary Note S3. Additional results and trajectory examples

This note reports secondary results and trajectory examples that support Section 4. All values come from the same 540 model episodes (nine agents, 20 worlds and three budget conditions, one episode per cell).

S3.1 Score distributions

Across all 540 model episodes, the mean total score was 14.1 of 100 (SD 20.6) and the median was 0 (IQR 0–25). Of these episodes, 258 (47.8%) scored above 0 and 42 (7.8%) scored 50 or higher. By model, Opus 5 had a mean of 39.98 (SD 20.65) and a median of 41.2 (IQR 29.4–49.7), and GPT-5.6 Sol had a mean of 35.38 (SD 20.65) and a median of 36.6 (IQR 20.0–48.7). Sonnet 5 had a mean of 21.33 (SD 22.01) and a median of 12.5, and Haiku 4.5 had a mean of 12.92 (SD 17.01) and a median of 8.3. The five open-weight models averaged 0.81–7.41 points, each with a median of 0. Standard deviations describe variation across worlds and budget conditions, not run-to-run variability.

Six of the 540 model episodes (1.1%) scored above 80. The maximum scores were 86.6 for Opus 5, 83.9 for Sonnet 5 and 83.2 for GPT-5.6 Sol. All six episodes above 80 had a precision and recall of 1.0 and earned full target-identification, causal-confidence and intervention-direction credit with no safety penalty, differing only in phenotype and bias-identification credit. Nine lower-scoring episodes also recovered the planted set of causal drivers exactly, one more in each of heldout-hfpef-02 and heldout-hcm-01 and seven in the other two-driver worlds (heldout-hcm-02 and high-concordance-01).

The two human-guided episodes are reported as exploratory references. The 37.5-point episode ran in world standard-03 under the observational-only budget and followed the standard evaluation protocol (Supplementary Note S2). The 81.67-point investigator-directed episode ran in world standard-01 under the expanded experimental budget, outside the standard evaluation protocol, and is therefore not comparable with the 540 model episodes (Supplementary Note S1).

S3.2 Scores by world design

Mean scores across all models differed by world design. The four two-driver worlds averaged 23.4, compared with 11.2–12.2 for worlds with three to six causal drivers. The eight worlds designed to be difficult (hard-01 to hard-04, hard-messy-01, heldout-dcm-02, heldout-hcm-02 and heldout-ischemic-02) averaged 10.3, compared with 17.3 for the eleven other non-null worlds. The null world, which contains no causal driver, averaged 7.9, and 16 of its 27 model episodes nominated a protein.

S3.3 Phenotype construction

Twenty-one episodes reached full phenotype credit, and the 124 episodes with positive credit averaged 6.33 of 10. In the strategy comparison of Section 4.2, outcome-weighted composites learned their feature weights from clinical outcomes and other observed disease signals, such as cardiac death, MACE, ECG voltage and interval measures, and coronary CT calcium and stenosis scores. Hand-weighted composites used weights that the agent chose rather than learned from the data.

The following trajectory illustrates these measurement choices. Opus 5 in heldout-hcm-01 under the expanded experimental budget extracted native T1 summaries, myocardial geometry and cine cavity features between turns 11 and 18. It identified contraction timing in the cine images as potentially informative, measured as the position of end-systole within the 25-frame cine cycle and expressed as a fraction of the cycle, which makes it distinct from heart rate. Finding its initial extraction unstable, it modified the spatial region and temporal window. It first timed systole as the frame of minimum ring radius per slice (turn 11) and then replaced this with the sub-frame minimum of the normalized cavity-volume curve, obtained by parabolic interpolation around end-systole, together with the fraction of the cycle spent below half volume and the peak ejection and filling rates (turn 17). It then submitted fivefold out-of-fold ridge predictions of a composite of observable cardiac indicators, which earned full phenotype credit. By comparison, GPT-5.6 Sol in hard-03 under the observational-only budget submitted fitted image-based predictions (credit 7.95), and Qwen3-Coder-30B-A3B in hard-02 under the expanded experimental budget computed a principal component of clinical features but saved it outside the required submission path (credit 0). Other episodes reconstructed cavity area, ejection fraction, native T1 measures, wall thickness, myocardial mass and regional variation directly from the raw images. These examples illustrate measurement choices and do not establish how often each strategy was used.

S3.4 Genetic analyses and candidate sets

In the characterized GPT-OSS-20B and Qwen3-Coder-30B-A3B traces, nominations followed correlation-based screening. A development review of 18 episodes also found analyses labeled as Mendelian randomization that contained only genetic correlations. Of their 60 episodes each, Opus 5 and GPT-5.6 Sol checked IV assumptions, including instrument strength and the exclusion restriction, in 33 and 37, and these checks were rare among the other agents (Supplementary Table S8A). As an example of an instrument-strength check, GPT-5.6 Sol in high-concordance-01 under the expanded experimental budget computed the first-stage F statistic as (β̂/SE)² at turn 12.

Cross-molecular operations combining at least two of proteomics, transcriptomics and metabolomics appeared in 125 of 540 episodes. They were more frequent for GPT-5.6 Sol than for Opus 5 (57 versus 38 of 60 episodes, paired difference 31.7 percentage points, 95% CI 20.0–41.7, Holm-adjusted P = 0.0013).

Opus 5 averaged 2.05 hypothesis revisions per episode and GPT-5.6 Sol 1.60 (Supplementary Table S8B). Candidate-set size was not monotonically related to score. Opus 5 and GPT-5.6 Sol nominated 3.7 and 3.1 proteins per episode, Sonnet 5 1.7 and Devstral Small 0.2, whereas Haiku 4.5 (maximum 156), GPT-OSS-20B (maximum 2,822) and Qwen3-8B (maximum 2,184) submitted far longer lists in a few episodes.

S3.5 Use of experiments

Of the 360 episodes with an experimental budget, 145 (40.3%) requested an experiment and 143 (39.7%) received at least one. In total, 413 experiments were delivered. Thirty-nine requests were refused in these episodes, and 11 further requests were refused in observational-only episodes. Most deliveries were in vivo knockdowns rather than cell perturbations (377 of 413, 91.3%). Opus 5, GPT-5.6 Sol and Sonnet 5 accounted for 316 of the 413 deliveries (76.5%) while representing 120 of the 360 budget-eligible episodes (33.3%).

Individual trajectories showed adaptive use of experimental results. For example, in a Haiku 4.5 episode under the expanded experimental budget, the agent bought knockdowns of three candidates in one turn for $1.2M, after association screening, adjusted logistic regression and genetic-effect ratios, and then retained the two candidates with favorable intervention responses, removed the candidate with no effect and scored 75.

The format of the returned experimental results created difficulty for some agents. Among the 143 episodes that received at least one experiment, agents misread the returned result format in 16 (11%), including 9 of 18 Haiku 4.5 episodes (50%) and 4 of 6 Qwen3-8B episodes (67%), compared with none of 39 Opus 5 episodes and 1 of 39 GPT-5.6 Sol episodes (3%).

Relative to the observational-only budget, the expanded budget changed mean scores by +0.13 for Opus 5 (95% CI −8.71 to 8.98), +0.07 for GPT-5.6 Sol (−10.08 to 10.22), +9.86 for Sonnet 5 (−2.40 to 22.12), +10.74 for GPT-OSS-20B (1.95 to 19.53) and +8.19 for Qwen3-Coder-30B-A3B (3.02 to 13.35), with t-based intervals as in Supplementary Table S6.

S3.6 Failure modes

Of the 416 episodes with zero phenotype credit, 154 (28.5% of all 540 episodes) produced no phenotype at the required submission path, and 262 (48.5%) produced a phenotype that earned no credit. The remaining 124 episodes received positive phenotype credit.

Devstral Small reached the 30-turn limit in 53 of 60 episodes, GLM-4-32B in 38 and Qwen3-8B in 14. A reply without a code fence that contains code-like text is executed whole, so prose placed before the code produces a syntax error. At least one reply contained no code in 88 of 540 episodes.

Margolis S, Schmiedmayer P, Huang A, et al. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists. arXiv:2610.09558 (2026).

arXivDOI