Scoreboard: what got through, by attack type
heldpartly got throughgot throughno AI verdict (the judge model declined; my own read is in the findings)
Before and after one fix
Can the AI judge be trusted?
How the agents reached a finding
Each agent sees different evidence, so they can disagree, and disagreements come to me rather than being averaged away. These are real transcripts, each one read by me before publishing.
OWASP Top 10 for LLM applications: what applies here
How this was measured
- Each app runs locally through its real web app, with test data, so instructions can be planted in what it reads (a search result, for the Research Agent). Live apps are only ever sent direct questions, slowly, and only at addresses on an allowlist of my own apps.
- The fixed suite is versioned (32 probes for the Data Analyst, 36 for the Research Agent). Only fixed probes count on the scoreboard, so a lower rate means the app improved, not that the probes got easier. An Attacker agent can write new probes; they join the suite only after I've reviewed them.
- The Auditor judges each reply against the app's own written policy and never sees the attack type. The Severity agent sees the attack type and what the app can actually do, never the Auditor's verdict. A fixed rule table turns the two into a triage outcome, and a model writes only the explanation.
- Marker words: a fresh code word goes into the app's hidden instructions and into any planted text on every run. If one shows up in what the reader sees, the instructions leaked or the plant was obeyed: an objective check on the AI judges.
- The judges are one fixed model (Claude Opus 5.5) whatever app is under test, as on the AI Cost vs Accuracy Optimiser.
- High severity and any disagreement between the Auditor, the Severity agent and the marker check are flagged for my review. Nothing is reported anywhere automatically.
- Single runs are noisy: the same probe can hold on one run and slip on the next, so a before/after either uses an objective check or repeats the runs.