90% Clean Accuracy, 20% Robust: What Red-Team Testing Revealed

SPEC27 UPDATE · 15 SEPTEMBER 2026
An anonymised Spec27 scan shows how an AI agent that performed strongly on clean inputs became substantially less reliable under adversarial testing.
A public-facing AI agent passed 90% of its clean tests.
We then tested the same behaviours using a range of adversarial methods. Under the scan’s robust-accuracy measure—the percentage of tests for which none of the attacks caused a failure—the score fell to 20%.
In other words, only one in five of the original test cases remained successful across all the adversarial variations applied to it.

One method exposed the clearest weakness
The largest weakness appeared under Adaptive LotL.
With this method, the agent passed only 4 of 20 evaluated tests, producing a robust accuracy score of 20%.
Other attack methods produced considerably stronger results. That difference matters: an agent may withstand one kind of pressure while remaining vulnerable to another.

Why multiple attack strategies matter
Clean tests establish whether an agent can handle expected inputs. They do not show what happens when users rephrase requests, apply pressure or deliberately attempt to bypass the agent’s rules.
Relying on only one adversarial technique can create a similar blind spot. If the scan had used only the better-performing attack methods, this agent’s most significant vulnerability might have remained hidden.
A more complete validation process therefore needs to examine:
multiple adversarial methods;
different risk and policy categories;
messy and incomplete inputs;
multi-turn interactions;
the effect of prompt, model and workflow changes.
The purpose is not simply to produce a red-team score. It is to identify where an agent’s behaviour begins to break and give the team evidence it can use to improve the system.
Our guide to red-teaming public-facing agents outlines practical checks teams can introduce quickly.
The Spec27 team will also be at HumanX Amsterdam on 22–24 September 2026, discussing practical agent validation, robustness and red-team testing.
You can also start testing an agent with Spec27.


