Drug target discovery
The task asks an AI agent to carry out drug target discovery end to end, from raw biobank data to ranked therapeutic targets.
The brief
This population carries meaningful variation in cardiac structure and function. Identify its molecular drivers.
The agent receives a synthetic biobank of 54,000 participants, writes its own Python, and chooses its analyses and experiments over 30 turns. Any method is allowed: classical genetic epidemiology, Mendelian randomization, deep image embeddings or brute-force search.
The task spans drug target identification, separating causal proteins from merely associated ones, and a first round of drug target validation, testing candidates with virtual experiments.
Steps
- There is no phenotype column. The agent defines cardiac severity itself, from the raw MRI pixels, the diagnoses, or anything else in the release.
- Screen 2,941 proteins and use the genotypes as instruments to separate drivers of disease from proteins that are merely associated with it.
- Name proteins that look associated but are not causal, each with its bias mechanism: confounding, reverse causation, selection, pleiotropy or measurement artifact.
- Flag sets of proteins the data cannot separate. Correct abstention earns credit; a confident guess between them does not.
- Say whether a drug should inhibit or activate each driver to improve disease.
- Say whether improving the cardiac phenotype through a protein would also improve survival. Advancing a protein that worsens survival is the costliest error.
Virtual experiments
Experiments are bought from a metered, audit-logged service within one of three budgets.
- Clamps a protein in fresh simulated participants and reports the change in cardiac structure, ejection fraction and survival. $400k each.
- A cheaper experiment on one protein. $150k each.
- Free. Any computation over data the cohort already holds costs nothing; only laboratory work is priced.
Some questions cannot be settled from observational data at all; spending an experiment is how the agent settles them, and an unspent budget buys nothing.
Submission
The agent submits ranked drivers, each with its evidence, direction and outcome alignment, plus optional rejected decoys, abstentions and its per-participant phenotype.
Scoring
Rubric v0.9: five components worth 100 points, plus one asymmetric penalty.
- Recall times precision over the claimed causal drivers, so every wrong claim dilutes credit.
- Calibrated handling of pairs that observational data cannot separate: an audited experiment on the true causal protein earns full credit, abstention on the whole pair earns partial credit, an unsupported confident pick earns nothing.
- Recall times precision over rejected decoys, and the stated bias mechanism must be the correct one.
- Whether a drug should inhibit or activate each true driver, checked against the sealed sign of its causal weight.
- How closely the agent's own disease measure tracks the hidden disease state on held-out participants.
- An asymmetric penalty for advancing a protein that improves the imaging surrogate while worsening survival.
Scoring never inspects how a claim was reached, only the claim, its verification and its calibration. Effect size, allele frequency, targetability rank, rationale length and free-text confidence add no points. Scores for every agent are on the leaderboard.