01 · Benchmark validity & scalable oversight
Memorisation, reward hacking and refusal on a benign item look identical to genuine
capability on a static rubric. I design regimes that make that confusion expensive:
stratified cohort splits, phenotype-similar distractors, dynamic harness elicitation,
held-out regeneration, cluster-bootstrap rank stability, effective sample size.
02 · Contamination control
An evaluation metric is a joint function of model capability, scaffold budget and
harness leakage. I ship the taxonomy, the JSON Schema, the validator and the
pre-registered audit instrument — unknown is a legal field,
because a reporter who cannot determine something should be able to say so.
Taxonomy, schema & validator ↗
03 · Judges, verdicts & grader reliability
Adversarial evaluation of the system under test and of the grader scoring it:
inter-rater variance, LLM-as-judge failure modes, premature verdict commitment, and
sycophancy paired against XSTest and StrongREJECT. A behavioural score is only as good
as the overseer that produced it.
04 · Release-gating infrastructure
AISI Inspect under parameter-parity: same items, same generation parameters, same
grader. Models rank by thresholds violated rather than by a composite, and every
published run states its grader, temperature, seed, harness versions and sample counts.
See the published run ↗
05 · Agentic safety — method, not product
Multi-agent tool use evaluated under annotation-overlap stratification on 1,047
out-of-distribution cases. The domain is genomics; the contribution is isolating
agentic capability from retrieval leakage.
Stratified, not averaged ↓
06 · Multilingual validity
Evaluation across four languages validated against native-speaker judgment, with
translation artifact separated from a real capability or safety gap and harms mapped to
the legal construct they are meant to operationalise. A crucible for construct validity,
not a localisation side-quest.
EN/ES/FR/PT, CLEF 2025 ↗
Domain applications. High-stakes out-of-distribution evaluation, neuro-symbolic
verification and contamination-controlled biomedical benchmarks. The design pattern —
overlap flags, phenotype-similar distractors, publication-clustered inference — is
domain-agnostic; genomics is where it gets stress-tested.
Technical AI governance. The EU AI Act and ISO/IEC 42001 read at article and clause
level, then implemented as pipelines and policy-as-code that emit an audit trail rather
than compliance prose written after the fact — for regulated enterprises.
Evaluation leadership. Built and led a 16-person model-evaluation function at Multiverse
Computing, owning release gating and safety benchmarking across language, vision-language
and compressed models.