top of page

What to Look for in an AI Agent Red-Team Run

Writer: Jovanca Garnadi
Jovanca Garnadi
Sep 22
2 min read

SPEC27 UPDATE · 22 SEPTEMBER 2026


A red-team result is most useful when it helps you find a behaviour you can investigate and improve.

In a recent anonymised Spec27 scan, an AI agent passed 90% of its clean tests, while its robust accuracy under adversarial testing was 20%. That difference showed the value of testing beyond expected inputs. It also raised a practical question: what should a team look at once the run is complete?


Start with the failures


An overall score tells you how the agent performed across a set of tests. To act on the result, look more closely:

  • Which expected behaviours failed?

  • What inputs or attack methods exposed the failures?

  • Did the agent break a policy, miss an escalation or respond incorrectly to missing information?

  • Can the team reproduce the behaviour after making a change?


Looking at the individual cases helps turn a broad result into specific work for the team.


Compare different kinds of pressure


An agent may handle one form of adversarial input well and struggle with another. In the scan we shared, one method produced substantially more failures than the others.


That is why a red-team run should help teams inspect results by method and behaviour, rather than relying on a single headline number. The aim is to find where the agent’s behaviour begins to change and what needs further testing.


Our article, A Red-Team Run in Practice with Spec27, takes a closer look at the testing process and its results.


Meet us at HumanX Amsterdam


The Spec27 team is at HumanX Amsterdam on 22–24 September 2026. If you’re attending, come and talk to us about the agents you’re building or deploying, and the behaviours you need to validate.


 
 
bottom of page