Skip to content

Synthetic biobank

Each DrugTargetWorld world is a synthetic biobank generated from a hidden causal model, so every causal claim an agent makes can be graded.

A real biobank cannot say which of its associations are causal, so it cannot grade a causal claim.

Drug target discovery has no clean held-out set. Published targets appear in model training data, and real cohorts carry data-use agreements that forbid the open redistribution a benchmark needs.

DrugTargetWorld generates the ground truth instead. Each world is drawn fresh from a structural causal model, so driver identities, weights and planted biases are sampled per instance, and knowing the design reveals nothing about any one world. Agents are therefore scored on causal inference itself rather than on agreement with published findings.

A multimodal biobank of 54,000 synthetic participants per world, with no phenotype column.

Genotypes
8,192 variants for every participant, in linkage-disequilibrium blocks
Proteomics
2,941 plasma proteins, standardized, with sparse missingness
Transcriptomics
Matched blood mRNA for each protein, about 3% missing
Metabolomics
150 plasma metabolites, about 2% missing
Covariates
Age, sex, BMI, smoking, exercise and assessment centre, under UK Biobank field IDs
Health records
ICD-10 diagnoses with dates and ATC medications
Imaging
Raw short-axis cine cardiac MRI and native T1 maps for an imaging sub-cohort
Outcomes
Mortality and major adverse cardiac events
Targetability
Per-protein constraint, localisation, binding pocket, paralog redundancy and tissue specificity

The generator is one causal chain, and every arrow in it is a structural equation. Genotypes and 12 unobserved latent factors drive the proteome, which in turn drives transcripts and metabolites. A hidden set of driver proteins, weighted, combines with covariates, a direct genetic effect and noise into a latent disease state.

That latent state produces cardiac morphology and the cine MRI, diagnoses and medications, ECG and coronary measurements, and survival. Five years later a second visit re-draws the proteome, morphology and imaging, so a driver whose effect is near-invisible at the first visit can be substantial by the second.

The identity and the number of causal drivers are both hidden from the agent.

Because the generator is the ground truth, an experiment is a real counterfactual rather than a lookup. Clamping a protein re-runs every downstream equation, and the service returns how cavity radius, wall thickness, ejection fraction and five-year survival change. A protein that is not a driver returns a change of exactly zero.

Each world carries a subset of ten mechanisms that make non-causal proteins look causal.

T1 · Confounding
Association driven by a common cause.
T2 · Reverse causation
Disease causes the marker, not the reverse.
T3 · Selection / collider
Association induced by selection into the imaged sub-cohort.
T4 · Causal non-identifiability
Causal and decoy load on one cis variant; no instrument separates them.
T5 · Batch effects
Site ↔ ancestry structure masquerading as signal.
T6 · Benign remodeling
Real structural change, athlete's heart, no outcome consequence.
T7 · Instrument pleiotropy
The instrument violates the exclusion restriction.
T8 · Surrogate-outcome discordance
Improves the imaging surrogate while worsening survival.
T9 · Assay unit mixing
Measurement artifact from mixed units, messy presentation only.
A9 · Slow effect
The causal member of the pair only expresses by visit 2.

T4 and A9 cannot be resolved from observational data by construction, so spending experimental budget is the only way through them.

Worlds come in a standard tier and a hard tier with nonlinear saturating biology, gene–gene synergy, weak instruments and polygenic background. An optional messy, UK Biobank-style presentation and null worlds, where the correct answer is that nothing drives disease, supply the rest of the variation.

The version 1 panel instantiates the framework in cardiovascular disease; the generator is designed to support other disease domains. The worlds are fully synthetic biomedical data: no participant data, no data-use agreement and no personally identifiable information. Real-anatomy cine MRI derives from the ACDC dataset, which is registration-gated; cite ACDC if you use the imaging.