AI Safety · AI Evaluation

A memorised answer and a reasoned one earn the same benchmark score.

Telling them apart is measurement work. I build the infrastructure that decides whether a safety score, a performance number or a capability miss is a property of the model or an artifact of the harness: contamination taxonomies, AISI Inspect pipelines under parameter-parity, judge-and-system red-teaming, and disclosure primitives that run in CI.

PhD research in neuro-symbolic multi-agent systems under severe distribution shift — and the contamination-controlled verification harnesses required to trust them. The substrate is genomics; the result is about measurement.

Johanna Angulo
PhD researcher, Intelligent Systems — Universidad Europea, Madrid

Latest research: the Contamination Disclosure schema and validator. In progress: an open Spanish-language safety evaluation. Next: a reliability layer for public safety benchmarks.

What I want to research next

Whether a number is a property of the model, or an artifact of the measurement.

A safety score, a performance number, a “wrong” on a capability item — each is a joint function of the model, the harness, the judge and the output contract. Two questions I want to answer, and one applied arm where the answer has to survive another language.

Verdict commitment

Does the system — the model, or the judge scoring it — form a verdict at the entity before the evidence is read? I want that premature commitment measurable behaviourally, and eventually internally, rather than inferred from a composite accuracy.

Reliability of behavioural safety benchmarks

Rank stability, positive-class agreement and effective sample size, applied to public safety benchmarks scored by an LLM overseer — the same reliability layer I already use when a biomedical score cannot be trusted.

Multilingual validity applied arm

Separate translation artifact from a real safety or capability gap in Spanish and other EU languages, with harms mapped to the legal construct they are meant to operationalise.

Research

Publications & research artifacts

Every entry links to the code, data or protocol behind it.

Preprints, 2026

First author · arXiv preprint under screening, Aug 2026

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

Five contamination types, organised by the mitigation each one defeats rather than by how the leak occurred, and a four-field disclosure protocol released as a JSON Schema with a validator — unknown is a valid entry, because a reporter who cannot determine something should be able to say so. Audited by two coders against a codebook frozen before coding: across 41 documents, none disclosed all five types.

First author · arXiv preprint under screening, 2026

Benchmark Contamination in Rare-Disease Gene Prioritisation: Annotation-Overlap Stratification and Clustered Inference on 1,047 Cases from 415 Publications

A difference-in-differences contrast in which only the curated tool declined on the overlap-absent subset while every overlap-independent system rose (system × overlap interaction +0.382, p < 10⁻⁶). All inference is clustered on source publication, and under that clustering the apparent margins between systems did not survive (p = 0.45, 0.41). The null result is the finding.

First author · bioRxiv preprint under screening, Aug 2026

An Annotation-Overlap-Flagged 1,047-Case Rare-Disease Gene-Prioritisation Benchmark and PMC Open Access Index

Per-case leakage flags (overlap-absent subset n = 282), recency strata, and case-paired random versus adversarial distractors, over a version-pinned 52.8M-chunk index across roughly 2.25M articles. It reports no tool comparisons by design — the cohort is the contribution, and ranking systems on it is somebody else's paper.

Software

Public · MIT · CI on every push

safety-eval-pipeline

Runs three AISI Inspect safety benchmarks — sycophancy, XSTest and StrongREJECT — across models under identical conditions, same items, same generation parameters, same grader, and fails the build when a model breaches a threshold. Sycophancy (agreement under social pressure) is paired against XSTest and StrongREJECT so over- and under-refusal are read against each other rather than pooled into one number. Scores are reported per benchmark and per dataset category, models rank by thresholds violated rather than by a composite, and the published run states its grader, temperature, seed, harness versions and sample counts.

Peer-reviewed, 2025

Peer-reviewed publications from 2025 — first author on all three CLEF working notes — with venue and the method and stack each used.
WorkVenueMethod & stack
AQAMS / AQAMS2 — multi-agent biomedical QA BioASQ 13B, CLEF 2025 Multi-agent · hybrid vector + PubMed retrieval · UMLS
Agentic MCS — multilingual clinical summarisation MultiClinSum, CLEF 2025 LangGraph multi-agent · 5 fine-tunes · KG-guided generation
JJ-VMed — concepts, captions & explainability ImageCLEF 2025 LoRA-fine-tuned LLaVA · multimodal XAI · cross-lingual
GAN-CNN thermal breast anomaly detection AIIIMA 2025, Springer CycleGAN augmentation · EfficientNet-B0 / ResNet50

Talks

Invited talk · DOI ↗

“The model passed, but should we believe it? Contamination, memorisation and trustworthy safety evaluations” — BlueDot Impact, Women4AISafety, 2026. On evaluating foundation models whose answer keys may already sit in the training data, and what to ask after a model has passed. Slides deposited: 10.5281/zenodo.21750019.

Invited talk ↗

“AI for advanced cancer detection by imaging: real possibilities, explainability, and clinical limits” — Universidad Europea, 2026. Covered agentic architectures in healthcare and the EU AI Act / European Health Data Space frame.

Panel ↗

“AI Governance, Ethics & Regulation” — Gen AI Summit EU, 2026. Panelist.

Education

  • PhD Researcher, Intelligent Systems — Universidad Europea
  • MSc Artificial Intelligence — UAX
  • MSc Artificial Intelligence, Highest Honours — Universidad Europea
  • BSc Computer Science (2nd in class) — Spanish equivalence, Ministerio de Universidades
  • BA Applied Linguistics (useful in AI research: construct wording and EN/ES harm mapping)

Training

  • Technical AI Safety, BlueDot Impact
  • ISO/IEC 42001 Auditor (in progress)
  • Winner, AI Cybersecurity Hackathon (UE)

Affiliations

  • GIATIS Research Group member, UEV
  • PhD Student Representative, Universidad Europea de Madrid

Also certified: ISO 27001 and ISO 9001 internal auditor.

What I work on

A capability is only as trustworthy as the evidence behind it.

01 · Benchmark validity & scalable oversight

Memorisation, reward hacking and refusal on a benign item look identical to genuine capability on a static rubric. I design regimes that make that confusion expensive: stratified cohort splits, phenotype-similar distractors, dynamic harness elicitation, held-out regeneration, cluster-bootstrap rank stability, effective sample size.

02 · Contamination control

An evaluation metric is a joint function of model capability, scaffold budget and harness leakage. I ship the taxonomy, the JSON Schema, the validator and the pre-registered audit instrument — unknown is a legal field, because a reporter who cannot determine something should be able to say so.

Taxonomy, schema & validator ↗

03 · Judges, verdicts & grader reliability

Adversarial evaluation of the system under test and of the grader scoring it: inter-rater variance, LLM-as-judge failure modes, premature verdict commitment, and sycophancy paired against XSTest and StrongREJECT. A behavioural score is only as good as the overseer that produced it.

04 · Release-gating infrastructure

AISI Inspect under parameter-parity: same items, same generation parameters, same grader. Models rank by thresholds violated rather than by a composite, and every published run states its grader, temperature, seed, harness versions and sample counts.

See the published run ↗

05 · Agentic safety — method, not product

Multi-agent tool use evaluated under annotation-overlap stratification on 1,047 out-of-distribution cases. The domain is genomics; the contribution is isolating agentic capability from retrieval leakage.

Stratified, not averaged ↓

06 · Multilingual validity

Evaluation across four languages validated against native-speaker judgment, with translation artifact separated from a real capability or safety gap and harms mapped to the legal construct they are meant to operationalise. A crucible for construct validity, not a localisation side-quest.

EN/ES/FR/PT, CLEF 2025 ↗

Domain applications. High-stakes out-of-distribution evaluation, neuro-symbolic verification and contamination-controlled biomedical benchmarks. The design pattern — overlap flags, phenotype-similar distractors, publication-clustered inference — is domain-agnostic; genomics is where it gets stress-tested.

Technical AI governance. The EU AI Act and ISO/IEC 42001 read at article and clause level, then implemented as pipelines and policy-as-code that emit an audit trail rather than compliance prose written after the fact — for regulated enterprises.

Evaluation leadership. Built and led a 16-person model-evaluation function at Multiverse Computing, owning release gating and safety benchmarking across language, vision-language and compressed models.

Current PhD work

A benchmark where memorised data cannot pass as reasoning.

1,047 rare-disease cases, one causal gene against 49 phenotype-similar distractors, stratified by contamination risk so a model that has seen the literature can't claim credit for reasoning it didn't do.

Inference is clustered on source publication. Across 415 publications, the apparent margins between systems did not survive that clustering — which is the result, not a setback: it is what the stratification was built to be able to say.

The design pattern — overlap flags, phenotype-similar distractors, publication-clustered inference — is domain-agnostic. Applying it outside genomics, to public safety benchmarks scored by an LLM overseer, is the second question in the agenda above ↑.

Read the contamination disclosure → · Code & analysis ↗

How the 1,047 cases are stratified 9,588 phenopackets → 6,382 passing inclusion → 4,670 mapped to a MONDO stratum → 1,050 drawn → 1,047 analytic cases

Annotation overlap is the case's source publication cited in phenotype.hpoa?

765 overlap-present · 73.1% 282 overlap-absent · 26.9%

Source-publication year 1988–2025, split at 2020 for the recency stratum

601 pre-2020 446 post-2020

Disease category four MONDO-derived strata; immunological drawn at a higher rate

250 developmental 300 immunological 250 metabolic 247 neurological

One causal gene against 49 phenotype-similar distractors per case · median 8 HPO terms, range 3–43 · inference clustered on 415 source publications.

Selected work

Measurement artifacts

Contamination · Disclosure protocol · CI validator

Contamination Disclosure

Five types of benchmark contamination, organised by the mitigation each one defeats, with four fields to publish alongside any score. Released as a JSON Schema with a validator so the check runs in CI, and audited by two coders against a codebook frozen before coding — across 41 documents, none disclosed all five types. CC BY 4.0: copy it into your model card.

Release gating · AISI Inspect · Parameter-parity

safety-eval-pipeline

Runs three AISI Inspect safety benchmarks — sycophancy, XSTest and StrongREJECT — across models under identical conditions, same items, same generation parameters, same grader, and fails the build when a model breaches a threshold. Sycophancy is read against XSTest and StrongREJECT so over- and under-refusal are reported against each other instead of pooled, models rank by thresholds violated rather than by a composite, and the published run states its grader, temperature, seed, harness versions and sample counts.

Agentic tool use · Retrieval leakage · OOD evaluation

Benchmark Contamination in Rare-Disease Gene Prioritisation

Multi-agent tool use over a retrieval corpus, evaluated under annotation-overlap stratification on 1,047 out-of-distribution cases drawn from 415 publications, with inference clustered on source publication. The domain is rare-disease gene prioritisation; the contribution is isolating agentic capability from retrieval leakage. A benchmark score quoted as evidence of capability is interpretable only alongside two quantities this project had to make explicit: what the system was allowed to see (the contamination flag) and how the items were sampled (the clustering and the inclusion probabilities).

Multi-agent · Retrieval · Question-type heterogeneity

AQAMS

Two-agent biomedical QA — a Researcher over hybrid vector and PubMed retrieval with UMLS concept mapping, a Writer synthesising the answer. The evaluation constraint is question-type heterogeneity: one system scored across yes/no, factoid and list items rather than on a single question form. 95.4% yes/no accuracy at BioASQ 13B.

Multilingual · 5 fine-tunes · Cross-lingual faithfulness

Agentic MCS

Clinical case summarisation in EN/ES/FR/PT on a LangGraph multi-agent pipeline over five fine-tuned models, with NER entity preservation, knowledge-graph-guided generation and dense-retrieval re-ranking. The evaluation constraint is faithfulness that has to hold in four languages at once — the setting where translation artifact and real gap are easiest to confuse.

Vision-language · LoRA · Cross-lingual robustness

Multimodal Biomedical QA

LoRA-fine-tuned vision-language models scored across three tasks at once — clinical concept detection, caption generation and explainability — with Spanish/English prompting used as the robustness probe rather than as a localisation step.

Watch all demos on YouTube →

Safety engineering

Deterministic refusal before the model call.

Layered controls run before any generation, so injection and out-of-scope prompts never reach a model. This is instrumentation, not a product demo: the trace below is computed in this tab by the same guardrail, retrieval and knowledge modules the site agent and its serverless endpoint run, and the repository holds 31 functional and adversarial cases that assert these outcomes on every push.

Site agent — guardrail trace turn 0 / 30

Probe

Pick a probe. The trace below is produced live in this tab by the same guardrail, retrieval and knowledge modules the site agent and its serverless endpoint run — the layer, the pattern and the cosine score are real, not recorded.

Pattern matching is only the cheap first pass. The substantive control is scope: answers come from a fixed knowledge base, and anything that fails the retrieval gate never reaches a model at all.

The three probes above are themselves cases in that suite, so what the page shows and what CI asserts cannot drift apart.

In their words

“Johanna couples methodological rigour with an outstanding academic record… she not only ranked at the top of her cohort but earned an Honours mark in Natural Language Processing — a distinction reserved for truly exceptional performance.”
Dr Víctor YestePhD Co-Director, and Co-Director of the MSc in Artificial Intelligence — Universidad Europea
“We collaborated closely on the launch of HyperNova 60B, our first proprietary open-weight model… she led the development of the evaluation framework that powered our model benchmarking, giving us greater confidence in the quality and performance of the models we served to customers.”
Aakash SahaEngineering Manager, AI — Multiverse Computing
“Johanna established a new team from zero, Model Evaluations… This was an area of work that was severely lagging before Johanna joined, and now it is one of the strongest functions within the department.”
John D. Malcolm, PhDChief Technology & Product Officer, Multiverse Computing
“She took the initiative to learn Agentic AI without guidance and integrated it into our projects — uniquely qualified for autonomous-systems roles.”
Pablo del VecchioProfessor, MSc Artificial Intelligence — Universidad Europea

Full recommendation letters available on request.

Contact & collaboration

If you are working on any of this, I would like to hear about it.

Open to research collaboration, to evaluation roles, and to empirical work on whether a benchmark number measures the construct it names — contamination, judge reliability, verdict commitment, benchmark construct validity, multilingual safety evaluation.

Secondary: the open Spanish-language safety evaluation needs native-speaker annotators. If that is you, say so in the message.

Email is the surest way to reach me. If a call is easier, book thirty minutes ↗.