Contamination Disclosure

A score is not a measurement until you say what produced it

When a model passes a benchmark, the number tells you that the answers were right. It does not tell you why they were right.

Two things produce the same score: a system that reasoned its way to each answer, and a system that had encountered the answers before. On the benchmark they are indistinguishable. The difference does not live in the score — it lives in how the evaluation was built, what the system could reach while it ran, and what the person reporting the number chose to say about both.

Why this exists

Most published scores say very little. The model is named and the number is given. The scaffolding that produced it, the population it was averaged over, and its relationship to whatever the model was trained on are usually left out. That is not dishonesty; it is the absence of a convention for saying them.

This project is an attempt at that convention: a short, cheap record that travels with a benchmark score and states what was and was not controlled for — including, explicitly, the parts the reporter could not determine.

Status

Under peer review.

The work behind this page is currently under peer review. The specification, the taxonomy, the schema and the validator are public in the repository under CC BY 4.0.

The paper itself, and the full audit write-up, follow once the review concludes.

In the meantime

If you would like to be told when that happens, or you are working on something related in the meantime, get in touch.

Contamination Disclosure · CC BY 4.0