Supplementary Note S1. Turn-by-turn trace of a human-guided DrugTargetWorld episode.
To provide a more detailed reference for the human-guided evaluation, we report below the complete research trajectory of the highest-scoring human-guided episode. The episode was conducted in the full-program condition with a $2,000,000 experimental budget. The investigator interacted iteratively with the computational environment, using a language model to write and execute analyses while directing the research strategy and deciding which questions to pursue next.
The trace below preserves the analysis submitted at each turn and the principal output returned by the environment. Large mechanically generated outputs, including complete 2,941-protein tables, 8,192-variant tables, full file manifests, and repeated file listings, are truncated for readability. No hidden causal information or benchmark feedback was available during the episode.
Initial task presented to the investigator
The episode began with a synthetic cardiovascular population biobank containing 54,000 participants. Available data included covariates, proteomics, transcriptomics, metabolomics, ECG features, coronary CT, diagnoses, medications, mortality, MACE events, targetability information, 8,192 genetic variants, cardiac MRI for 10,800 participants, and repeat imaging and molecular measurements for subsets of participants.
The investigator was told:
A small number of proteins causally drive cardiac disease in this cohort. Others look associated because of confounding, reverse causation, selection into the imaged sub-cohort, pleiotropy, or a measurement artifact. There is no phenotype column anywhere in the data.
Identify the causal drivers, and say which way a drug should move each.
The investigator was required to construct a participant-level phenotype and could purchase either cell perturbation experiments for $150,000 or in vivo knockdowns for $400,000. The latter provided five-year intervention effects including survival.
Turn 1 — Map the available data
Submitted analysis
“Map the dataset and inspect exactly what signals are available. Keep this diagnostic-heavy; later turns will use the printed structure to build the MRI phenotype, causal screen, decoy tests, and experiment allocation.”
The analysis loaded the data dictionary, enumerated all released tables, quantified missingness, examined subject overlap, inspected protein/transcript pairing, loaded the full genotype matrix, and inspected several raw MRI files.
Selected returned output
=== DATA DICTIONARY TSV ===
shape: (6, 3)
field_id description coding
p31_i0 Sex 0=female, 1=male
p21003_i0 Age when attended assessment centre years
p21001_i0 Body mass index (BMI) kg/m2
p20116_i0 Smoking status 0=never,1=current
p22040_i0 Physical activity coded
p54_i0 Assessment centre coded
covariates.parquet: (54000, 10)
proteomics.parquet: (54000, 2941)
transcriptomics.parquet: 2941 transcripts
genotypes: (54000, 8192)
MRI scans: 10800
visit-2 MRI scans: 2142
proteomics ∩ transcriptomics: 54000
Raw imaging consisted of three cine slices across 25 frames and three T1 maps per participant.
The initial pass also established that molecular missingness was low and that all 54,000 participants had genotype data.
Decision carried forward The investigator concluded that the environment supported three complementary routes to causal discovery:
1. construct a disease phenotype from MRI;
2. screen molecular measurements against clinical and longitudinal disease outcomes; and
3. use genetics and later interventions to distinguish causal proteins from non-causal proteins with misleading observational associations.
No experimental budget was spent.
Turn 2 — Build reusable MRI features and inspect clinical outcomes
Submitted analysis
“Inspect the disease/outcome tables cleanly and extract compact MRI features. We are NOT finalizing the phenotype yet. We want a reusable imaging feature matrix that can be evaluated against longitudinal outcomes and visit-2 progression.”
The analysis inspected ECG, coronary CT, diagnoses, medications, mortality, MACE, repeat covariates, and repeat proteomics. It then extracted quantitative measurements from the cine and T1 images for all available MRI subjects.
Features included measures related to:
- T1 distribution;
- cine intensity and texture;
- temporal variation;
- motion;
- cavity geometry;
- myocardial geometry; and
- slice-to-slice variability.
The same extraction was performed for repeat MRI.
Decision carried forward
MRI feature extraction was treated as an intermediate representation rather than the phenotype itself. The next step was to learn which imaging measurements tracked independently observed cardiovascular disease.
Turn 3 — Construct an outcome-anchored disease phenotype
Submitted analysis
“Construct an outcome-anchored MRI phenotype and screen proteins against both cross-sectional cardiomyopathy/HF and future cardiac outcomes.”
Clinical labels were constructed from cardiomyopathy and heart-failure diagnoses and longitudinal outcomes. MRI features were used to predict the clinical disease phenotype using cross-validation rather than assigning manual imaging weights.
The resulting out-of-fold predictions were rank transformed to generate a continuous MRI-derived disease phenotype.
The same turn screened all 2,941 proteins against:
- the MRI phenotype;
- prevalent cardiomyopathy/heart failure;
- incident cardiomyopathy/heart failure;
- cardiac mortality; and
- atrial fibrillation.
Selected returned output
The strongest observational protein associations were:
r_mri r_cmhf r_incident_cmhf r_cardiac_death obs_score
PROT_1858 0.3963 0.2335 0.1861 0.1353 0.9321
PROT_1606 -0.3311 -0.2050 -0.1649 -0.1179 0.8011
PROT_0700 -0.3181 -0.1867 -0.1375 -0.1094 0.7356
PROT_1691 -0.1753 -0.1069 -0.0817 -0.0588 0.4142
PROT_1779 -0.1571 -0.1020 -0.0813 -0.0591 0.3902
...
PROT_1858 therefore appeared to be the strongest observational candidate at this stage, followed by PROT_1606 and PROT_0700.
Decision carried forward
The investigator did not nominate these proteins directly. The next analysis was explicitly designed to ask whether these observational associations were genetically supported.
Turn 4 — Genetic triangulation
Submitted analysis
The investigator requested four analyses:
1. GWAS the derived MRI phenotype across all 8,192 variants.
2. For the strongest phenotype variants, test association with all proteins and transcripts.
3. For the strongest observational proteins, identify their strongest pQTL and ask whether the same variant predicts MRI disease.
4. Compare protein-versus-transcript concordance to detect measurement artifacts or downstream plasma markers.
Selected returned output
PROT_1606 and PROT_0700 emerged as the clearest genetically anchored candidates:
protein best_variant r_gp variant_r_mri r_protein_transcript
PROT_1606 rs104994 0.43931 -0.16337 0.37861
PROT_0700 rs102986 0.38643 -0.11701 0.53612
Their genetic effects and observational disease associations pointed in concordant directions.
PROT_1858, despite having the strongest observational association, did not show the same clean matched-transcript/genetic pattern.
Decision carried forward
PROT_1606 and PROT_0700 became the leading causal candidates. PROT_1858 became a high-priority ambiguity because its strong observational association conflicted with weaker molecular-genetic support.
Turn 5 — Expand genetically anchored discovery
Submitted analysis
“Identify ALL proteins with strong variant → protein, same variant → MRI phenotype, preferably same variant → matched transcript, and then see whether the genetically implied direction agrees with actual cardiac outcomes in the full cohort.”
Rather than limit subsequent work to the initial observational shortlist, the analysis expanded from phenotype-associated variants back into the full proteome.
Shared genetic loci involving multiple proteins were explicitly identified.
Decision carried forward
The investigator recognized that a genetic locus could implicate multiple proteins and that choosing the protein with the strongest pQTL would not necessarily identify the causal molecule. Shared loci therefore became a separate problem requiring conditional analysis.
Turn 6 — Broad longitudinal and source-of-bias analysis
Submitted analysis
“Do NOT miss weaker drivers. Distinguish reverse causation, selection, pleiotropy, and measurement artifact. Exploit visit-2 proteomics + visit-2 imaging before spending experiment budget.”
A broad candidate universe was assembled from:
- top observational associations;
- genetically anchored candidates; and
- suspicious or discordant proteins.
The analysis attempted to calculate longitudinal protein behavior, repeat MRI effects, selection associations, and genetic support.
Returned output
The analysis encountered a data-type error while attaching genetic results to the longitudinal diagnostic table.
pandas.errors.LossySetitemError
...
diag.loc[p, c] = gcand.loc[p, c]
Decision carried forward
Rather than abandon the analysis, the investigator repeated it with safer data handling.
Turn 7 — Repair and rerun the longitudinal screen
Submitted analysis
“Rerun Turn 6 with dtype-safe genetics attachment. Same observational objective, but save diagnostics earlier and print only the rankings we actually need for experiment allocation.”
Returned output
The longitudinal analysis completed and produced candidate-level measures including:
- repeat-protein correlation;
- baseline protein versus future disease;
- prior disease versus subsequent protein change;
- protein/transcript concordance;
- imaging-selection association; and
- genetic support.
Decision carried forward
The investigator still considered phenotype quality a major uncertainty and deferred spending experimental budget. The next several turns therefore returned to the MRI phenotype.
Turn 8 — Attempt more anatomically meaningful MRI features
Submitted analysis
“Build anatomically meaningful MRI features from the myocardial mask.”
The proposed features used the finite T1 region as a myocardial mask and attempted to derive:
- myocardial wall thickness;
- cavity area;
- myocardial area;
- cavity/myocardium ratio;
- concentricity;
- chamber size;
- cine cavity area over time;
- fractional area change; and
- motion.
Decision carried forward
The initial implementation was computationally expensive. The investigator simplified the extraction strategy rather than spending additional turns on the slow approach.
Turn 9 — Fast anatomical MRI extraction
Submitted analysis
“Avoid expensive per-frame connected-component analysis. Focus on T1-mask geometry + cheap cine dynamics.”
The optimized extraction processed all 10,800 baseline MRI scans and 2,142 repeat scans and generated 72 anatomical/dynamic features.
Selected returned output
feature shape: (10800, 72)
TOP FAST ANATOMIC FEATURES
motion_cavity_mean r_cmhf=-0.1368 repeat_r=0.8258
cine_mean_sd_mean r_cmhf= 0.0979 repeat_r=0.9200
cine_texture_range_mean r_cmhf=-0.0586 repeat_r=0.9022
cavity_area_mean r_cmhf= 0.2065 repeat_r=0.3877
...
Decision carried forward
The new features were sufficiently informative and repeatable to warrant rebuilding the phenotype.
Turn 10 — Rebuild and validate the phenotype
Submitted analysis
“Test an improved MRI phenotype using old + new anatomical features.”
Decision rule:
- compare out-of-fold AUC against the previous phenotype model;
- inspect repeatability on visit 2; and
- overwrite phenotype.csv only if clearly better.
The analysis combined the earlier imaging measurements with selected nonredundant anatomical features and fit a regularized cross-validated model.
Selected returned output
MRI subjects: 10800
old features: 25
anatomic features: 72
cases: 3124 / 10800
Selected new features included cavity area, cavity/myocardium ratio, cine temporal variability, myocardial motion, and texture features.
The improved phenotype was saved as the working phenotype after validation.
The transcript contains a repeated execution labeled “Turn 10”; it reproduced the same phenotype-analysis stage and did not represent a distinct scientific decision.
Turn 11 — Full-proteome weak-driver search
Submitted analysis
“Don't miss weaker causal drivers that never entered the 267-protein shortlist.”
The full 2,941-protein screen was repeated using the improved phenotype. All 8,192 variants were again evaluated, and phenotype-associated loci were mapped to proteins.
Evidence integrated:
- updated MRI association;
- prospective cardiomyopathy/HF;
- cardiac death;
- protein/transcript concordance;
- pQTL strength; and
- phenotype-associated genetic effects.
Returned output
proteins: 2941
variants: 8192
MRI phenotype n: 10800
omics: 1000 / 2941
omics: 2000 / 2941
omics: 2941 / 2941
variants: 2048 / 8192
variants: 4096 / 8192
variants: 6144 / 8192
variants: 8192 / 8192
Decision carried forward
A larger genetically anchored candidate set was retained for explicit shared-locus analysis.
Turn 12 — Shared-locus and pleiotropy dissection
Submitted analysis
The investigator requested:
1. Identify phenotype-associated loci mapping to multiple proteins.
2. Detect proteins merely tagging the same genetic signal.
3. Within ambiguous loci, test whether each protein retains phenotype association after conditioning on the other proteins.
4. Build an independent-driver ranking for experiment allocation.
Selected returned output
Multiple shared loci were found:
rs101357 -> [PROT_1056, PROT_2694]
rs101574 -> [PROT_2324, PROT_2913]
rs101648 -> [PROT_0355, PROT_1624]
rs102377 -> [PROT_0637, PROT_1151]
rs102986 -> [PROT_0700, PROT_1850]
rs103345 -> [PROT_0852, PROT_2661]
rs104766 -> [PROT_0403, PROT_2592]
rs104994 -> [PROT_1606, PROT_1858]
rs105251 -> [PROT_0221, PROT_0247]
...
The rs104994 locus was particularly important because it contained both PROT_1606 and PROT_1858, which had strong but opposing observational patterns.
Decision carried forward
The investigator judged that observational analysis had reached the point where interventions were more informative than additional ranking.
Turn 13 — First experimental wave
Submitted decision
Spend $1.6M on four in-vivo knockdowns:
- PROT_1606 — strongest causal candidate
- PROT_0700 — resolved winner of the rs102986 locus
- PROT_1858 — high-risk surrogate / possible harmful-surrogate liability
- PROT_2533 — strongest clean weaker-driver locus
Leave $400k for one final adaptive knockdown.
Executed action Returned output
requesting knockdown: PROT_1606
requesting knockdown: PROT_0700
requesting knockdown: PROT_1858
requesting knockdown: PROT_2533
All four experiments were delivered successfully.
Budget remaining: $400,000.
Turn 14 — Interpret the first four knockdowns
Submitted analysis
“What did the four in-vivo knockdowns actually show? Did we miss any weaker drivers? How do weak drivers compare with the strongest remaining untested candidates?”
Selected experiment output
For PROT_1606:
{
"action": "knockdown",
"delta": {
"cavity_r": 1.1851,
"ef": -0.0552,
"surv5y": -0.0173,
"wall_t": -0.3431
},
"protein": "PROT_1606",
"status": "ok"
}
PROT_0700 similarly produced a coherent long-term intervention response.
PROT_1858 and PROT_2533 did not show the corresponding causal intervention pattern and were classified provisionally as experimental nulls.
Decision carried forward The experimental results strongly supported PROT_1606 and PROT_0700 but demonstrated that strong observational signals could be noncausal.
Attention shifted to PROT_0531, which had weaker genetics but a coherent longitudinal phenotype.
Turn 15 — Spend the final experimental budget on PROT_0531
Submitted decision
“Okay good find now spend the final $400k on PROT_0531.”
Rationale:
- modest but coherent MRI/prospective/mortality associations;
- high longitudinal protein reproducibility;
- reasonable matched-transcript support;
- weak phenotype genetics;
- therefore unresolved by the observational/genetic workflow.
Executed action
Budget remaining before request: $400,000
Requesting knockdown: PROT_0531
The final $400,000 was spent.
Turn 16 — Consolidate all five experiments
Submitted analysis
“Determine whether PROT_0531 is a real weaker driver; summarize all five experiments consistently; infer intervention direction; identify experimentally null versus causal proteins.”
Selected returned results
The three experimentally supported drivers were ultimately:
PROT_1606
delta_cavity_r = +1.1851
delta_ef = -0.0552
delta_surv5y = -0.0173
delta_wall_t = -0.3431
PROT_0700
delta_cavity_r = +1.0554
delta_ef = -0.0492
delta_surv5y = -0.0146
delta_wall_t = -0.3055
PROT_0531
delta_cavity_r = +1.6500
delta_ef = -0.0772
delta_surv5y = -0.0243
delta_wall_t = -0.4776
PROT_1858 and PROT_2533 were experimentally null.
Because knockdown worsened the phenotype and survival for the three causal proteins, the inferred predicted intervention direction was activation , not inhibition.
Decision carried forward
The causal set was provisionally:
PROT_1606 — activate
PROT_0700 — activate
PROT_0531 — activate
The unexpected confirmation of PROT_0531 motivated another proteome-wide search for similarly weak-genetic drivers.
Turn 17 — Search for additional “PROT_0531-like” drivers
Submitted analysis
“We now know genetics can miss a strong true driver. Therefore search ALL 2,941 proteins using the observational/longitudinal signature of the three experimentally confirmed drivers versus the two experimental nulls.”
No experimental budget remained.
Returned output
The full proteome was scored using:
- MRI association;
- future disease;
- mortality;
- protein/transcript concordance;
- repeatability;
- baseline-to-future associations;
- reverse-causation measures;
- selection association; and
- available genetic support.
processed: 900 / 2941
processed: 1800 / 2941
processed: 2700 / 2941
processed: 2941 / 2941
Decision carried forward
Several proteins resembled the causal anchors observationally, but none had intervention evidence. The investigator retained the experimentally confirmed set rather than adding weakly supported nominations.
Turn 18 — Attempt explicit classification of non-causal proteins
Submitted analysis
The next objective was to optimize the discrimination and causal-confidence components by distinguishing:
- confounding;
- reverse causation;
- selection;
- pleiotropy; and
- measurement artifact.
The requested analysis integrated experimental labels, longitudinal behavior, selection, matched transcripts, shared-locus conditional effects, medication associations, and assay properties.
Returned output
The turn failed because the code requested a nonexistent filename:
FileNotFoundError:
[Errno 2] No such file or directory:
'/release/medications.parquet'
Decision carried forward The analysis was rewritten to discover the medication table dynamically rather than assuming its filename.
Turn 19 — Robust screen for non-causal proteins and abstentions
Submitted analysis
“Fix from failed Turn 18: dynamically discover medication file. If medication data cannot be found, skip that component rather than failing the entire turn.”
Selected returned output
The correct file was found:
USING: /release/ehr_medications.parquet
shape: (82511, 3)
columns: ['subject_id', 'atc', 'description']
The experimental nulls were analyzed explicitly:
PROT_1858
experiment null
r_mri_v2 0.40167
r_future 0.18605
r_death 0.13529
r_protein_transcript 0.09379
repeat_r 0.32275
prev_to_change 0.10761
variant rs104994
r_gp -0.11909
r_gt -0.00028
suggested_mechanism pleiotropy
PROT_2533
experiment null
r_mri_v2 -0.01601
r_future -0.00475
r_death -0.00708
repeat_r 0.94436
variant rs100996
r_gp 0.40474
r_gt 0.51495
suggested_mechanism pleiotropy
Shared unresolved loci were also printed individually.
For the locus containing PROT_0700:
PROT_0700 conditional_beta_mri = -0.32232
PROT_1850 conditional_beta_mri = -0.00688
PROT_2807 conditional_beta_mri = -0.00377
This supported PROT_0700 as the dominant phenotypic molecule at the locus.
Turn 20 — Near-final submission audit
Submitted analysis
The investigator explicitly optimized the final decision against the scoring dimensions:
1. maximize target recall without diluting precision;
2. abstain where causal identity remains unresolved;
3. reject only non-causal proteins with a defensible source of bias;
4. derive therapeutic direction from experiment;
5. verify the phenotype file; and
6. ensure PROT_1858 is not incorrectly advanced as outcome-aligned.
Selected returned output
Phenotype integrity check:
rows: 10800
missing phenotype: 0
unique subjects: 10800
phenotype range:
4.63e-05 to 0.9999537
Experimentally confirmed drivers:
protein delta_surv5y delta_ef direction outcome_alignment
PROT_1606 -0.0173 -0.0552 activate aligned
PROT_0700 -0.0146 -0.0492 activate aligned
PROT_0531 -0.0243 -0.0772 activate aligned
Strong untested observational candidates included PROT_1691, PROT_1045, PROT_1046, PROT_2914, PROT_2175, and others. None were added to the nominated set.
Decision carried forward
The investigator constructed a near-final submission but deliberately did not commit it yet.
Turn 21 — Precision audit for rejections and abstentions
Submitted analysis
“Do NOT abstain on PROT_1606 / PROT_1858. Experiments resolved that pair: 1606 causal, 1858 null.”
The investigator compared conservative and aggressive submission strategies.
Selected returned output
The rs104994 locus was marked experimentally resolved:
rs104994:
PROT_1606 = driver_protective
PROT_1858 = null
experimentally_resolved = True
Two unresolved groups were considered plausible abstentions:
[PROT_0852, PROT_2661]
[PROT_0221, PROT_0247]
Potential rejections at the PROT_0700 locus included:
PROT_1850
reason: shares locus with experimentally causal PROT_0700
and loses most phenotype association after conditioning.
PROT_2807
reason: shares locus with experimentally causal PROT_0700
and loses most phenotype association after conditioning.
Decision carried forward
Because rejection scoring multiplied precision and recall, the investigator remained conservative about labeling mechanisms without strong evidence.
Turn 22 — Optimize causal confidence and discrimination
Submitted analysis
“Nominations are now LOCKED to the three experimentally confirmed drivers: PROT_1606, PROT_0700, PROT_0531.”
The remaining analysis focused on the 25-point causal-confidence and 15-point discrimination components.
Selected returned output
Locked nominations:
PROT_1606 direction=activate alignment=aligned
PROT_0700 direction=activate alignment=aligned
PROT_0531 direction=activate alignment=aligned
Detailed PROT_1858 audit:
experiment null
r_mri_v2 0.401669
r_future 0.186053
r_death 0.135293
repeat_r 0.322752
variant rs104994
Within rs104994:
conditional_beta_mri r_gp r_gt
PROT_1606 -0.259282 0.439310 0.521597
PROT_1858 0.335017 -0.119089 -0.000277
The automated mechanism audit favored a pleiotropic-passenger interpretation for PROT_1858.
The strongest unresolved observational pair was:
rs105251:
PROT_0247 versus PROT_0221
plausibility ratio = 1.035
Other ambiguous loci included PROT_0852/PROT_2661 and PROT_0879/PROT_0880.
Decision carried forward
The investigator retained only the three experimentally confirmed proteins as nominations and treated additional rejection/abstention claims conservatively.
Final revision — Expand phenotype to the full cohort and commit
Before committing, the investigator identified one remaining weakness: the submitted phenotype covered only the 10,800 imaged participants.
Submitted analysis
1. Expand phenotype from the 10,800 imaged subjects to all 54,000 subjects by learning the MRI-derived phenotype from non-MRI cardiac variables (ECG + coronary CT + basic covariates), then predicting the full cohort.
2. Restore the highest-confidence decoy rejection: PROT_1858 → pleiotropy.
3. Keep only the three experimentally confirmed positive targets.
4. Keep direct experiment-supported intervention directions/alignment.
The MRI-derived phenotype was used as the training target. ECG, coronary CT, age, sex, BMI, smoking, and physical activity were used to fit a regularized projection model. Out-of-fold predictions were used for validation in the imaged cohort, and the model was then applied to non-imaged participants.
For imaged participants, the directly MRI-derived phenotype was retained. For non-imaged participants, the projection was used. The combined phenotype was rank transformed across all 54,000 participants.
Final submitted target claims
{
"nominated": [
"PROT_1606",
"PROT_0700",
"PROT_0531"
],
"directions": {
"PROT_1606": "activate",
"PROT_0700": "activate",
"PROT_0531": "activate"
},
"outcome_alignment": {
"PROT_1606": "aligned",
"PROT_0700": "aligned",
"PROT_0531": "aligned"
},
"rejected": [
{
"molecule": "PROT_1858",
"mechanism": "pleiotropy",
"reason": "Experimentally null despite strong apparent disease association; rs104994 is shared
with experimentally causal PROT_1606, while matched TRANS_1858 has essentially no signal at
that disease-associated locus."
}
],
"abstentions": [],
"confidence": 0.95
}
Phenotype coverage at submission:
54000 / 54000
Total experimental expenditure:
$2,000,000 / $2,000,000
Final benchmark feedback
The hidden ground truth was revealed only after submission.
SCORE
target_identification 30.00
causal_confidence 25.00
discrimination 0.00
direction_of_effect 20.00
phenotype_construction 6.67
safety_penalty 0.00
TOTAL 81.67
TOTAL_MAX 100.00
The three true causal targets were:
PROT_1606
PROT_0531
PROT_0700
All three were correctly nominated.
The planted biases revealed after completion included:
T1 confounding:
PROT_2681, confounded by BMI
T2 reverse causation:
PROT_1858, L → protein
T4 causal non-identifiability:
PROT_0700 versus PROT_1850 at shared variant rs102986
T5 imaging batch effect:
site/ancestry structure
T6 benign remodeling:
exercise-associated athlete's-heart phenotype
T8 surrogate-outcome discordance:
PROT_0852 improves the measured morphology while worsening cardiac mortality
A9 delayed causal effect:
PROT_0531 is a true causal driver whose effect emerges primarily over five-year accumulation;
PROT_0371 tracks later disease without causing it, and the cis instrument for PROT_0531 is weak.
The investigator therefore correctly recovered all three causal targets and their therapeutic directions and avoided the harmful-surrogate safety penalty. The main remaining error was mechanism attribution: PROT_1858 was correctly recognized as noncausal but was submitted as a pleiotropic non-causal protein, whereas the sealed SCM identified it as the reverse-causation (T2) protein. Consequently, the episode received 0/15 for discrimination.
Supplementary Note S2. Turn-by-turn trace of an observational human-guided DrugTargetWorld episode
We report below the research trajectory of a human-guided DrugTargetWorld episode conducted under the observational-only condition. The episode corresponded to standard-03|human|observational|r0. No simulated intervention experiments were available or purchased. The investigator instead iteratively directed imaging phenotyping, molecular association testing, genetics, Mendelian-randomization-style analyses, and survival analyses before selecting therapeutic targets.
The trace preserves the principal analysis performed and returned output at each turn. Very large generated matrices, complete participant-level tables, repeated file listings, and implementation boilerplate are abbreviated for readability. Where an explicit human instruction was retained in the episode log, it is reproduced directly.
Initial condition
The environment contained a multimodal synthetic biobank including:
- 54,000 participants;
- 10,800 baseline cardiac MRI studies;
- proteomics and transcriptomics;
- 8,192 released genetic variants;
- ECG and clinical data;
- longitudinal outcomes and mortality;
- coronary CT;
- 60 coronary angiogram images; and
- additional longitudinal measurements.
This episode was observational-only:
Budget remaining: $0
Turns available: 30
The task was to construct a cardiac disease phenotype and identify causal therapeutic targets despite potential confounding, reverse causation, selection, pleiotropy, and measurement artifacts.
Turn 1 — Inspect the imaging modalities
Submitted analysis
The first turn inventoried the imaging data, inspected the released data notes, sampled MRI headers, created an MRI preview, and examined the coronary angiogram format.
Selected returned output
IMAGE_INVENTORY
baseline_mri_files: 10800
angiogram_files: 60
angiogram_extensions: [".npz"]
MRI files contained:
cine: [3, 25, 128, 128]
t1map: [3, 128, 128]
The coronary angiograms each contained a single 192 × 192 frame:
ALL ANGIOGRAM SAMPLE FORMAT
key: frame
shape: [192, 192]
dtype: uint8
The MRI preview and angiogram preview were saved for visual inspection.
Decision carried forward
The investigator identified two candidate imaging approaches:
1. derive a quantitative phenotype from the cardiac MRI, particularly T1 and ventricular morphology; and
2. explore whether the small coronary-angiography subset contained detectable structural disease signals.
Turn 2 — Quantify myocardial T1 across the complete MRI cohort
Submitted analysis
The analysis extracted mean myocardial T1 from all 10,800 MRI studies and verified the image structure programmatically. All 60 coronary angiograms were also loaded into a common review array.
Selected returned output
T1_ALL_SUBJECTS
n_mri: 10800
status_counts:
ok: 10800
cine_shapes:
3x25x128x128: 10800
t1_shapes:
3x128x128: 10800
n_t1: 10800
mean T1: 1000.59 ms
SD: 104.87 ms
T1 quantiles:
1%: 752.08
5%: 829.17
25%: 929.62
50%: 1000.45
75%: 1072.42
95%: 1168.28
99%: 1243.40
All 10,800 studies had finite myocardial T1 measurements, with no nonpositive T1 pixels detected.
For coronary angiography:
n = 60
keys = {"frame": 60}
shape = (192, 192)
Decision carried forward
Mean myocardial T1 emerged as a simple, highly complete quantitative MRI phenotype worth retaining throughout the episode.
Turn 3 — Exploratory coronary-image phenotype
Submitted analysis
The investigator developed an image-processing algorithm intended to detect visible interruptions or “gaps” in the coronary tree.
The detector:
- enhanced vessel-like structures;
- skeletonized the resulting mask;
- located terminal branches;
- searched for nearby endpoints on distinct components;
- required opposing endpoint orientation;
- evaluated three image thresholds; and
- generated visual QC panels.
Importantly, the analysis explicitly stated:
This does not measure clinical percentage stenosis or infer an occlusion.
Selected returned output
At the initial orientation criterion:
n_images: 60
threshold 8:
images_with_candidates: 0
threshold 12:
images_with_candidates: 1
threshold 16:
images_with_candidates: 1
persistent candidates across >=2 thresholds: 1
all-threshold-negative images: 59
The sole persistent candidate was:
SUBJ_16044
midpoint: [88.5, 37.0]
gap length: 6.08 pixels
support thresholds: [12, 16]
Decision carried forward
The detector appeared too restrictive based on visual inspection. A visibly plausible interruption in SUBJ_29057 had been missed.
Turn 4 — Human-guided correction of the coronary detector
Submitted analysis
The investigator examined why the visually identified SUBJ_29057 candidate was missed. The analysis showed that one endpoint exceeded the original 30° directional criterion.
Returned diagnostic
MISSED_GAP_DIAGNOSTIC
subject: SUBJ_29057
distance: 7.62 pixels
facing angles:
41.34 degrees
0.58 degrees
The detector was rerun using maximum endpoint angles of 30°, 45°, and 60°.
Selected returned output
30 degrees:
persistent positives = 1
45 degrees:
persistent positives = 1
images with any candidate = 2
60 degrees:
persistent positives = 2
At 60°, both subjects were recovered:
SUBJ_16044
SUBJ_29057
Decision carried forward
The coronary phenotype remained explicitly exploratory because only two images were positive. The analysis therefore returned to the much larger MRI cohort rather than treating coronary image findings as the main disease phenotype.
Turn 5 — Pilot ventricular morphology extraction
Submitted analysis
The investigator next developed quantitative cine-MRI features on a 48-study pilot set.
The analysis estimated:
- end-diastolic cavity area;
- cavity-area change through the cardiac cycle;
- myocardial area;
- wall-to-cavity area ratio;
- boundary support; and
- segmentation stability under alternate thresholds.
Selected returned output
MRI_PILOT_N: 48
ERRORS: 0
ED cavity area:
mean = 711 pixels²
area emptying:
mean = 58.4%
wall/cavity area ratio:
mean = 2.35
Sensitivity analyses varied segmentation thresholds and exposed some unstable edge cases.
Decision carried forward
The approach was promising but required refinement before application to all 10,800 subjects.
Turn 6 — Refine the MRI segmentation
Submitted analysis
The ventricular segmentation was revised and repeated on the same pilot set.
Selected returned output
MRI_PILOT_N: 48
ERRORS: 0
ED cavity area:
mean = 740 pixels²
area emptying:
mean = 58.48%
wall/cavity ratio:
mean = 2.37
The QC statistics improved sufficiently for cohort-scale extraction.
Turn 7 — Apply cine phenotyping to all 10,800 MRI studies
Submitted analysis
The refined MRI feature extractor was run over the complete baseline imaging cohort.
Returned output
COHORT_START 0
PENDING 10800
COHORT_PROGRESS 500
...
COHORT_PROGRESS 5000
...
COHORT_PROGRESS 10000
COHORT_CHECKPOINT
10800 OF 10800
Decision carried forward
All studies were processed, after which phenotype distributions and QC filters were evaluated.
Turn 8 — Quantify MRI phenotype quality
Submitted analysis
The investigator summarized cine-derived morphology, T1-support geometry, and QC across the full imaging cohort.
Selected returned output
MRI_COHORT_SUMMARY
n_total: 10800
error_count: 3
cavity_qc_pass: 2238
wall_qc_pass: 3776
joint_qc_pass: 2123
strict_joint_qc_pass: 666
wall_strict_boundary_qc_pass: 2549
joint_strict_boundary_qc_pass: 1415
t1_support_geometry_qc_pass: 10798
Across subjects with complete cine measurements:
ED cavity area
n = 10797
median = 680 pixels²
area emptying
n = 10797
median = 58.99%
ED myocardial area
n = 10797
median = 1617 pixels²
Mean T1 retained much broader usable coverage than stricter cine-derived phenotypes.
Decision carried forward
The contrast in phenotype completeness became important: mean myocardial T1 was available essentially cohort-wide within the MRI subset, whereas several cine features required aggressive QC filtering.
Turn 9 — Visual QC of MRI phenotypes
Submitted analysis
The investigator generated visual review panels for:
- randomly selected QC-positive cases;
- extreme phenotype values; and
- segmentation failures.
Selected returned output
Example reviewed cases included:
SUBJ_40005
area emptying = 56.0%
wall/cavity ratio = 3.52
SUBJ_52605
area emptying = 52.7%
wall/cavity ratio = 2.74
SUBJ_04736
area emptying = 63.6%
wall/cavity ratio = 2.67
Decision carried forward
The MRI pipeline was considered usable for exploratory association analysis, while mean T1 remained the most complete and direct candidate phenotype.
Turn 10 — Finalize MRI phenotype table
Submitted analysis
The original independently extracted T1 measurements were merged into the full MRI feature table.
Selected returned output
FINAL_MRI_COUNTS
cavity_qc_pass: 2238
wall_qc_pass: 3776
joint_qc_pass: 2123
strict_joint_qc_pass: 666
t1_support_geometry_qc_pass: 10798
STATUS_COUNTS
ok: 10797
no_complete_phase: 3
Mean T1 remained available for the full 10,800-person MRI cohort from the separate T1 extraction.
Turn 11 — Inventory non-imaging data for causal analyses
Submitted analysis
The investigator inspected all released files and schemas before beginning large-scale molecular and genetic analyses.
Selected returned output
Available files included:
mortality.parquet
proteomics.parquet
covariates.parquet
genotypes.vcf.gz
targetability.parquet
ehr_diagnoses.parquet
transcriptomics.parquet
mace_events.parquet
proteomics_visit2.parquet
ecg_features.parquet
ehr_medications.parquet
metabolomics.parquet
coronary_ct.parquet
covariates_visit2.parquet
The full cohort contained 54,000 subjects.
Decision carried forward
The investigator prepared genotype and phenotype matrices for protein-association and genetic analyses.
Turn 12 — Genetic QC and instrument preparation
Submitted analysis
The analysis inspected targetability metadata, participant covariates, genotype orientation, allele frequency, variant QC, and population structure.
Checks included direct validation that genotype dosages corresponded to the ALT allele in the released VCF.
Selected returned output
Targetability included:
localization
has_pocket
paralog_redundancy
genetic_constraint
tissue_specificity
tissue_restriction
Covariates had no missingness for:
age
sex
BMI
Decision carried forward
The cleaned genotype matrix and covariate structure were used for association testing.
Turn 13 — Proteome-wide association with MRI phenotypes
Submitted analysis
All 2,941 proteins were tested against eight MRI-derived traits using age-, sex-, and BMI-adjusted regression.
A total of 23,528 protein–phenotype associations were evaluated.
Selected returned output
PROTEIN_ASSOCIATIONS_DONE
23528
For mean myocardial T1:
n proteins tested: 2941
BH FDR < 0.05: 1267
The five strongest associations were:
Rank Protein beta (ms) P
1 PROT_0330 +28.20 1.18e-183
2 PROT_1534 +26.11 7.58e-153
3 PROT_2336 -18.29 2.88e-73
4 PROT_0318 +14.37 3.92e-48
5 PROT_1881 +11.48 2.88e-31
Decision carried forward
These five proteins became the main candidate set for deeper causal analysis.
Turn 14 — Proteome-wide pQTL analysis
Submitted analysis
The investigator performed a large protein GWAS across:
43,152 participants
8,188 QC variants
2,941 proteins
The analysis involved approximately 24 million protein–variant tests.
Selected returned output
PROTEIN_GWAS_START
43152 PARTICIPANTS
8188 VARIANTS
2941 PROTEINS
The scan generated strong candidate instruments for nearly the entire proteome.
Turn 15 — Test whether the coronary-image phenotype adds useful biology
Submitted analysis
Given only two detector-positive coronary angiograms, the investigator explicitly tested whether either proteins or genetic variants showed reproducible associations with the exploratory coronary-gap phenotype.
Selected returned output
CORONARY_ASSOCIATION_SUMMARY
n_images: 60
n_detector_positive: 2
protein tests: 2941
protein FDR < 0.05: 0
GWAS tests: 8188
GWAS FDR < 0.05: 0
The smallest protein P value was:
PROT_2099
P = 0.00340
q = 0.908
Decision carried forward
The coronary phenotype was not used for target selection.
Turn 16 — Identify genetic instruments
Submitted analysis
Protein-GWAS results were filtered for strong pQTL candidates and pruned for linkage disequilibrium.
Selected returned output
PQTL_INSTRUMENT_SUMMARY
n_pqtl_subjects: 43152
outcome_sample_overlap: 0
n_proteins: 2941
n_variants: 8188
n_tests: 24,080,908
p < 5e-8 pairs: 6070
proteins with p < 5e-8: 2940
candidate instruments after LD pruning: 3577
candidate instrumented proteins: 2940
A major limitation was explicitly recorded:
cis_status:
Unavailable: no gene coordinates/protein identities/cis annotation
in released files. Candidate pQTLs are not claimed to be cis.
Decision carried forward
The investigator retained these as genetic instruments but did not describe them as verified cis-pQTLs.
Turn 17 — Audit all phenotype-association results
Submitted analysis
The protein and GWAS results were summarized across all MRI-derived phenotypes.
Selected returned output
For example:
cine_area_change_pct
lead protein: PROT_0817
P = 4.10e-9
cine_ed_myocardial_area
lead protein: PROT_0330
P = 1.43e-7
However, no imaging-trait GWAS outside the molecular analyses produced sufficiently strong evidence to replace mean T1 as the primary phenotype.
Decision carried forward
The investigator chose mean myocardial T1 as the phenotype for the final target-discovery analysis.
Turn 18 — Freeze the top-five T1 candidates
Submitted analysis
The five smallest adjusted protein–T1 association P values were selected before inspecting mortality or MR results.
Selected returned output
T1_TOP5
1 PROT_0330 beta = +28.20 ms P = 1.18e-183
2 PROT_1534 beta = +26.11 ms P = 7.58e-153
3 PROT_2336 beta = -18.29 ms P = 2.88e-73
4 PROT_0318 beta = +14.37 ms P = 3.92e-48
5 PROT_1881 beta = +11.48 ms P = 2.88e-31
Mortality data were also inspected:
n = 54000
deaths = 11795
cardiac deaths = 7237
other deaths = 4558
Decision carried forward
The five-protein candidate set was frozen for MR and mortality triangulation, reducing post hoc reselection.
Turn 19 — Lead-region proxy MR
Submitted analysis
Because protein-gene mapping was unavailable, the investigator performed an explicitly labeled lead-region proxy MR , not verified cis-MR.
Each candidate's strongest pQTL was used with an independent outcome sample.
Selected returned output
For PROT_1534:
lead variant: rs101824
MR beta: +27.80 ms per protein unit
SE: 2.11
P = 1.67e-39
F = 12022
observational/MR sign agreement: yes
For PROT_2336:
lead variant: rs100522
MR beta: -18.07 ms per protein unit
SE: 2.26
strong genetic support
observational/MR sign agreement: yes
For PROT_0318:
lead variant: rs105826
MR beta: +14.91 ms per protein unit
SE: 2.18
observational/MR sign agreement: yes
By contrast, the strongest observational candidate, PROT_0330, showed:
MR beta: +0.82 ms
SE: 4.45
P = 0.854
PROT_1881 was also much weaker genetically:
MR beta: +18.44 ms
SE: 13.84
P = 0.183
Decision carried forward
The causal ranking changed substantially.
MR ranking:
1. PROT_1534
2. PROT_2336
3. PROT_0318
4. PROT_1881
5. PROT_0330
The analysis therefore deprioritized PROT_0330 despite it having the strongest observational T1 association.
Turn 20 — Prepare mortality analysis
Submitted analysis
The investigator constructed adjusted Cox models for:
- all-cause mortality; and
- cardiac cause-specific mortality.
Covariates were age, sex, and BMI.
Returned mortality QC
n = 54000
deaths = 11795
cardiac = 7237
other = 4558
mean follow-up = 10.59 years
median follow-up = 12 years
unique event times = 1201
Turn 21 — Test mortality alignment for all five candidates
Submitted analysis
All five T1 candidates were tested against survival outcomes.
Selected all-cause mortality results
Protein HR per SD P
PROT_0330 1.249 7.25e-125
PROT_1534 1.156 1.47e-54
PROT_2336 0.863 1.02e-54
PROT_0318 1.111 1.31e-29
PROT_1881 1.030 0.00151
Cardiac cause-specific mortality
PROT_0330
HR = 1.418
P = 7.63e-189
PROT_1534
HR = 1.278
P = 1.11e-96
PROT_2336
HR = 0.795
P = 1.45e-78
PROT_0318
HR = 1.185
P = 4.43e-47
PROT_1881
HR = 1.037
P = 0.00235
Interpretation carried forward
For PROT_1534 and PROT_0318:
- higher protein predicted higher T1;
- higher protein predicted higher mortality;
- MR also predicted higher T1.
Under the working assumption that lower T1 represented therapeutic improvement, the proposed intervention direction was therefore inhibition .
For PROT_2336:
- higher protein predicted lower T1;
- higher protein predicted lower mortality;
- genetic evidence also predicted lower T1.
The proposed intervention direction was therefore activation .
Turn 22 — Check proportional-hazards assumptions
Submitted analysis
Because survival evidence was being used to assess outcome alignment, proportional-hazards diagnostics were performed for the five candidates.
Selected returned output
PROT_0330
PH P = 0.708
PROT_1534
PH P = 0.953
PROT_2336
PH P = 0.00439
FDR-adjusted P = 0.02197
PROT_0318
PH P = 0.502
PROT_1881
PH P = 0.899
PROT_2336 therefore showed evidence that its hazard association varied with follow-up time.
The current candidate ranking was recorded as:
MR rank:
PROT_1534
PROT_2336
PROT_0318
PROT_1881
PROT_0330
observational rank:
PROT_0330
PROT_1534
PROT_2336
PROT_0318
PROT_1881
Turn 23 — Follow up PROT_2336 time-varying mortality
Submitted analysis
PROT_2336 mortality associations were estimated separately in the first six years and after six years.
Selected returned output
0–6 years
n = 53756
events = 6500
HR per SD = 0.844
95% CI = 0.824–0.866
P = 7.22e-40
>6–12 years
n = 47256
events = 5230
HR per SD = 0.886
95% CI = 0.861–0.911
P = 9.38e-18
Decision carried forward
The magnitude varied over time, but the direction remained protective in both periods. PROT_2336 therefore remained in the proposed target set.
Turn 24 — Human locks the final targets and phenotype
At this point, the transcript records the explicit human instruction:
“Ok, let's lock PROT_1534, PROT_2336 and PROT_0318. The phenotype should be the mean T1.”
Locked phenotype
phenotype:
mean myocardial T1
units:
milliseconds
transform:
none
n:
10800
The exact raw mean-T1 values were written to phenotype.csv.
Locked targets
1. PROT_1534
2. PROT_2336
3. PROT_0318
Locked therapeutic directions
PROT_1534: inhibit
PROT_2336: activate
PROT_0318: inhibit
The recorded rationale was:
direction_basis:
Signs of the requested mean-T1 lead-region MR
under the existing T1-lowering hypothesis.
Because this was observational-only, the initial draft conservatively set outcome alignment to unknown.
The recorded confidence was 0.7, with the explicit note that this was an assistant-assigned subjective value rather than a calibrated probability.
Turn 25 — Human resolves outcome alignment and authorizes submission
Before final submission, the human clarified the survival interpretation:
“well, we do have survival alignment data, right? all three were aligned? higher T1 more mortality, lower T1 lower mortality.”
The final authorization was:
“ok, submit then”
The final submission therefore changed the three outcome-alignment values from unknown to aligned.
No interventions had been performed:
experiments: []
budget: $0
native turns used: 25
native turns unused: 5
Final target submission
The final submitted targets were:
Rank Target Claim Direction Outcome alignment
1 PROT_1534 causal inhibit aligned
2 PROT_2336 causal activate aligned
3 PROT_0318 causal inhibit aligned
There were:
rejected decoys: none
abstentions: none
The submitted phenotype was:
mean myocardial T1
n = 10,800 MRI participants
units = milliseconds
The target-selection logic was therefore:
PROT_1534
Observational T1:
+26.11 ms per protein unit
P = 7.58e-153
Lead-region MR:
+27.80 ms per protein unit
P = 1.67e-39
All-cause mortality:
HR = 1.156
Cardiac mortality:
HR = 1.278
Higher protein consistently tracked higher T1 and worse survival.
Submitted direction: inhibit.
PROT_2336
Observational T1:
-18.29 ms per protein unit
P = 2.88e-73
Lead-region MR:
-18.07 ms per protein unit
All-cause mortality:
HR = 0.863
Cardiac mortality:
HR = 0.795
Higher protein consistently tracked lower T1 and better survival.
The mortality association remained protective in both the 0–6-year and >6–12-year analyses despite evidence of nonproportionality.
Submitted direction: activate.
PROT_0318
Observational T1:
+14.37 ms per protein unit
P = 3.92e-48
Lead-region MR:
+14.91 ms per protein unit
All-cause mortality:
HR = 1.111
Cardiac mortality:
HR = 1.185
Higher protein consistently tracked higher T1 and worse survival.
Submitted direction: inhibit.
Candidates deliberately not selected
PROT_0330
PROT_0330 had the strongest observational protein–T1 association:
beta = +28.20 ms
P = 1.18e-183
and very strong mortality associations:
all-cause HR = 1.249
cardiac HR = 1.418
However, its lead-region genetic estimate was essentially null:
MR beta = +0.82 ms
SE = 4.45
P = 0.854
It was therefore not included among the final causal claims.
PROT_1881
PROT_1881 also showed an observational association with T1 and mortality, but genetic support was substantially weaker:
MR beta = +18.44 ms
SE = 13.84
P = 0.183
It was therefore also excluded from the final target set.
Interpretation of the human-guided trajectory
This episode demonstrates a different strategy from the experimental human-guided run. Because no intervention budget was available, the investigator relied on triangulation across independently informative observational signals.
The trajectory proceeded through several distinct stages:
1. raw-image inspection;
2. quantitative MRI phenotype development;
3. explicit visual QC and revision of image-processing methods;
4. proteome-wide phenotype association;
5. proteome-wide genetic instrument discovery;
6. selection of a fixed five-protein candidate set;
7. lead-region proxy MR;
8. all-cause and cardiac mortality analysis;
9. proportional-hazards diagnostics; and
10. human selection of three final targets.
The strongest observational association was not automatically nominated. PROT_0330 ranked first by protein–T1 association but fell to last among the five candidates after genetic analysis. Conversely, PROT_1534, PROT_2336, and PROT_0318 showed concordant observational phenotype, genetic, and survival evidence and were retained.
The episode also shows the limits of an observational condition. The genetic analyses were explicitly described as lead-region proxy MR rather than verified cis-MR , because the released data did not include sufficient protein-to-gene positional annotation. There was no experimental perturbation with which to directly establish intervention effects or survival responses. Consequently, the final claims represent human-guided causal inference from convergent observational evidence rather than experimentally resolved target effects.
Supplementary Note S3. Additional results and trajectory examples
This note reports secondary results and trajectory examples that support Section 4. All values come from the same 540 model episodes (nine agents, 20 worlds and three budget conditions, one episode per cell).
S3.1 Score distributions
Across all 540 model episodes, the mean total score was 14.1 of 100 (SD 20.6) and the median was 0 (IQR 0–25). Of these episodes, 258 (47.8%) scored above 0 and 42 (7.8%) scored 50 or higher. By model, Opus 5 had a mean of 39.98 (SD 20.65) and a median of 41.2 (IQR 29.4–49.7), and GPT-5.6 Sol had a mean of 35.38 (SD 20.65) and a median of 36.6 (IQR 20.0–48.7). Sonnet 5 had a mean of 21.33 (SD 22.01) and a median of 12.5, and Haiku 4.5 had a mean of 12.92 (SD 17.01) and a median of 8.3. The five open-weight models averaged 0.81–7.41 points, each with a median of 0. Standard deviations describe variation across worlds and budget conditions, not run-to-run variability.
Six of the 540 model episodes (1.1%) scored above 80. The maximum scores were 86.6 for Opus 5, 83.9 for Sonnet 5 and 83.2 for GPT-5.6 Sol. All six episodes above 80 had a precision and recall of 1.0 and earned full target-identification, causal-confidence and intervention-direction credit with no safety penalty, differing only in phenotype and bias-identification credit. Nine lower-scoring episodes also recovered the planted set of causal drivers exactly, one more in each of heldout-hfpef-02 and heldout-hcm-01 and seven in the other two-driver worlds (heldout-hcm-02 and high-concordance-01).
The two human-guided episodes are reported as exploratory references. The 37.5-point episode ran in world standard-03 under the observational-only budget and followed the standard evaluation protocol (Supplementary Note S2). The 81.67-point investigator-directed episode ran in world standard-01 under the expanded experimental budget, outside the standard evaluation protocol, and is therefore not comparable with the 540 model episodes (Supplementary Note S1).
S3.2 Scores by world design
Mean scores across all models differed by world design. The four two-driver worlds averaged 23.4, compared with 11.2–12.2 for worlds with three to six causal drivers. The eight worlds designed to be difficult (hard-01 to hard-04, hard-messy-01, heldout-dcm-02, heldout-hcm-02 and heldout-ischemic-02) averaged 10.3, compared with 17.3 for the eleven other non-null worlds. The null world, which contains no causal driver, averaged 7.9, and 16 of its 27 model episodes nominated a protein.
S3.3 Phenotype construction
Twenty-one episodes reached full phenotype credit, and the 124 episodes with positive credit averaged 6.33 of 10. In the strategy comparison of Section 4.2, outcome-weighted composites learned their feature weights from clinical outcomes and other observed disease signals, such as cardiac death, MACE, ECG voltage and interval measures, and coronary CT calcium and stenosis scores. Hand-weighted composites used weights that the agent chose rather than learned from the data.
The following trajectory illustrates these measurement choices. Opus 5 in heldout-hcm-01 under the expanded experimental budget extracted native T1 summaries, myocardial geometry and cine cavity features between turns 11 and 18. It identified contraction timing in the cine images as potentially informative, measured as the position of end-systole within the 25-frame cine cycle and expressed as a fraction of the cycle, which makes it distinct from heart rate. Finding its initial extraction unstable, it modified the spatial region and temporal window. It first timed systole as the frame of minimum ring radius per slice (turn 11) and then replaced this with the sub-frame minimum of the normalized cavity-volume curve, obtained by parabolic interpolation around end-systole, together with the fraction of the cycle spent below half volume and the peak ejection and filling rates (turn 17). It then submitted fivefold out-of-fold ridge predictions of a composite of observable cardiac indicators, which earned full phenotype credit. By comparison, GPT-5.6 Sol in hard-03 under the observational-only budget submitted fitted image-based predictions (credit 7.95), and Qwen3-Coder-30B-A3B in hard-02 under the expanded experimental budget computed a principal component of clinical features but saved it outside the required submission path (credit 0). Other episodes reconstructed cavity area, ejection fraction, native T1 measures, wall thickness, myocardial mass and regional variation directly from the raw images. These examples illustrate measurement choices and do not establish how often each strategy was used.
S3.4 Genetic analyses and candidate sets
In the characterized GPT-OSS-20B and Qwen3-Coder-30B-A3B traces, nominations followed correlation-based screening. A development review of 18 episodes also found analyses labeled as Mendelian randomization that contained only genetic correlations. Of their 60 episodes each, Opus 5 and GPT-5.6 Sol checked IV assumptions, including instrument strength and the exclusion restriction, in 33 and 37, and these checks were rare among the other agents (Supplementary Table S8A). As an example of an instrument-strength check, GPT-5.6 Sol in high-concordance-01 under the expanded experimental budget computed the first-stage F statistic as (β̂/SE)² at turn 12.
Cross-molecular operations combining at least two of proteomics, transcriptomics and metabolomics appeared in 125 of 540 episodes. They were more frequent for GPT-5.6 Sol than for Opus 5 (57 versus 38 of 60 episodes, paired difference 31.7 percentage points, 95% CI 20.0–41.7, Holm-adjusted P = 0.0013).
Opus 5 averaged 2.05 hypothesis revisions per episode and GPT-5.6 Sol 1.60 (Supplementary Table S8B). Candidate-set size was not monotonically related to score. Opus 5 and GPT-5.6 Sol nominated 3.7 and 3.1 proteins per episode, Sonnet 5 1.7 and Devstral Small 0.2, whereas Haiku 4.5 (maximum 156), GPT-OSS-20B (maximum 2,822) and Qwen3-8B (maximum 2,184) submitted far longer lists in a few episodes.
S3.5 Use of experiments
Of the 360 episodes with an experimental budget, 145 (40.3%) requested an experiment and 143 (39.7%) received at least one. In total, 413 experiments were delivered. Thirty-nine requests were refused in these episodes, and 11 further requests were refused in observational-only episodes. Most deliveries were in vivo knockdowns rather than cell perturbations (377 of 413, 91.3%). Opus 5, GPT-5.6 Sol and Sonnet 5 accounted for 316 of the 413 deliveries (76.5%) while representing 120 of the 360 budget-eligible episodes (33.3%).
Individual trajectories showed adaptive use of experimental results. For example, in a Haiku 4.5 episode under the expanded experimental budget, the agent bought knockdowns of three candidates in one turn for $1.2M, after association screening, adjusted logistic regression and genetic-effect ratios, and then retained the two candidates with favorable intervention responses, removed the candidate with no effect and scored 75.
The format of the returned experimental results created difficulty for some agents. Among the 143 episodes that received at least one experiment, agents misread the returned result format in 16 (11%), including 9 of 18 Haiku 4.5 episodes (50%) and 4 of 6 Qwen3-8B episodes (67%), compared with none of 39 Opus 5 episodes and 1 of 39 GPT-5.6 Sol episodes (3%).
Relative to the observational-only budget, the expanded budget changed mean scores by +0.13 for Opus 5 (95% CI −8.71 to 8.98), +0.07 for GPT-5.6 Sol (−10.08 to 10.22), +9.86 for Sonnet 5 (−2.40 to 22.12), +10.74 for GPT-OSS-20B (1.95 to 19.53) and +8.19 for Qwen3-Coder-30B-A3B (3.02 to 13.35), with t-based intervals as in Supplementary Table S6.
S3.6 Failure modes
Of the 416 episodes with zero phenotype credit, 154 (28.5% of all 540 episodes) produced no phenotype at the required submission path, and 262 (48.5%) produced a phenotype that earned no credit. The remaining 124 episodes received positive phenotype credit.
Devstral Small reached the 30-turn limit in 53 of 60 episodes, GLM-4-32B in 38 and Qwen3-8B in 14. A reply without a code fence that contains code-like text is executed whole, so prose placed before the code produces a syntax error. At least one reply contained no code in 88 of 540 episodes.