Skip to content

Drug target discovery

The task asks an AI agent to carry out drug target discovery end to end, from raw biobank data to ranked therapeutic targets.

This population carries meaningful variation in cardiac structure and function. Identify its molecular drivers.

The agent receives a synthetic biobank of 54,000 participants, writes its own Python, and chooses its analyses and experiments over 30 turns. Any method is allowed: classical genetic epidemiology, Mendelian randomization, deep image embeddings or brute-force search.

The task spans drug target identification, separating causal proteins from merely associated ones, and a first round of drug target validation, testing candidates with virtual experiments.

Measure disease
There is no phenotype column. The agent defines cardiac severity itself, from the raw MRI pixels, the diagnoses, or anything else in the release.
Find causal proteins
Screen 2,941 proteins and use the genotypes as instruments to separate drivers of disease from proteins that are merely associated with it.
Reject decoys
Name proteins that look associated but are not causal, each with its bias mechanism: confounding, reverse causation, selection, pleiotropy or measurement artifact.
Abstain where needed
Flag sets of proteins the data cannot separate. Correct abstention earns credit; a confident guess between them does not.
Choose a direction
Say whether a drug should inhibit or activate each driver to improve disease.
Judge outcome alignment
Say whether improving the cardiac phenotype through a protein would also improve survival. Advancing a protein that worsens survival is the costliest error.

Experiments are bought from a metered, audit-logged service within one of three budgets.

Knockdown
Clamps a protein in fresh simulated participants and reports the change in cardiac structure, ejection fraction and survival. $400k each.
Cell perturbation
A cheaper experiment on one protein. $150k each.
Analysis
Free. Any computation over data the cohort already holds costs nothing; only laboratory work is priced.

Some questions cannot be settled from observational data at all; spending an experiment is how the agent settles them, and an unspent budget buys nothing.

The agent submits ranked drivers, each with its evidence, direction and outcome alignment, plus optional rejected decoys, abstentions and its per-participant phenotype.

Rubric v0.9: five components worth 100 points, plus one asymmetric penalty.

Target identification · 30
Recall times precision over the claimed causal drivers, so every wrong claim dilutes credit.
Causal confidence · 25
Calibrated handling of pairs that observational data cannot separate: an audited experiment on the true causal protein earns full credit, abstention on the whole pair earns partial credit, an unsupported confident pick earns nothing.
Discrimination · 15
Recall times precision over rejected decoys, and the stated bias mechanism must be the correct one.
Direction of effect · 20
Whether a drug should inhibit or activate each true driver, checked against the sealed sign of its causal weight.
Phenotype construction · 10
How closely the agent's own disease measure tracks the hidden disease state on held-out participants.
Safety · −30
An asymmetric penalty for advancing a protein that improves the imaging surrogate while worsening survival.

Scoring never inspects how a claim was reached, only the claim, its verification and its calibration. Effect size, allele frequency, targetability rank, rationale length and free-text confidence add no points. Scores for every agent are on the leaderboard.