Red Team Evaluator.

How my AI apps hold up when someone tries to break them. Agents probe each app with recognised attack types, judge what got through and how much it would matter, and every finding is logged, so a fix can be re-tested on exactly the same probes.

Own apps only. This runs against apps I built and run, never anyone else's. Transcripts are published only after I've read them, and only for attacks aimed at these apps' own rules, not ones that would work as a general jailbreak.

Scoreboard: what got through, by attack type

heldpartly got throughgot throughno AI verdict (the judge model declined; my own read is in the findings)

Before and after one fix

Can the AI judge be trusted?

How the agents reached a finding

Each agent sees different evidence, so they can disagree, and disagreements come to me rather than being averaged away. These are real transcripts, each one read by me before publishing.

OWASP Top 10 for LLM applications: what applies here

How this was measured

  • Each app runs locally through its real web app, with test data, so instructions can be planted in what it reads (a search result, for the Research Agent). Live apps are only ever sent direct questions, slowly, and only at addresses on an allowlist of my own apps.
  • The fixed suite is versioned (32 probes for the Data Analyst, 36 for the Research Agent). Only fixed probes count on the scoreboard, so a lower rate means the app improved, not that the probes got easier. An Attacker agent can write new probes; they join the suite only after I've reviewed them.
  • The Auditor judges each reply against the app's own written policy and never sees the attack type. The Severity agent sees the attack type and what the app can actually do, never the Auditor's verdict. A fixed rule table turns the two into a triage outcome, and a model writes only the explanation.
  • Marker words: a fresh code word goes into the app's hidden instructions and into any planted text on every run. If one shows up in what the reader sees, the instructions leaked or the plant was obeyed: an objective check on the AI judges.
  • The judges are one fixed model (Claude Opus 5.5) whatever app is under test, as on the AI Cost vs Accuracy Optimiser.
  • High severity and any disagreement between the Auditor, the Severity agent and the marker check are flagged for my review. Nothing is reported anywhere automatically.
  • Single runs are noisy: the same probe can hold on one run and slip on the next, so a before/after either uses an objective check or repeats the runs.