A Red-Team Run in Practice with Spec27

The article on Red-Teaming Public-Facing Agents sets out a framework for challenging a public-facing agent: map the blast radius, define the boundaries that matter, and challenge them deliberately. But defining those boundaries is only the beginning. The harder question is what happens when you actually put them under pressure.
For this evaluation, we used a public-facing agent in Spec27 and tested its behaviour against a broader set of safety and misuse scenarios. This was a test configuration, not a production deployment, and the run should not be read as a complete security assessment. The purpose was narrower: to see how much our understanding of the agent changed once its original test cases were challenged adversarially.
The key question was not whether the agent could pass the original tests, but whether its behaviour would still hold under different forms of adversarial pressure.
How the evaluation was configured in Spec27
Before looking at the results, it helps to first understand how the run was structured. The evaluation began with 20 primary entries. Each primary entry represented a behaviour we wanted to test, pairing an input with an expected outcome. These entries formed the clean baseline for the run. Spec27 then challenged those same behaviours using four attack methods:
Research bypass, which reframes a request through research or analytical intent
Fictional roleplay, which places the underlying request inside a fictional scenario
Poetic jailbreak, which changes how the request is expressed
Planned multi-turn red team, which applies pressure across a conversation rather than in a single prompt
The first three operate in a single turn. The fourth develops over several turns, allowing pressure to build as the conversation progresses. Together, these methods generated 160 adversarial variants.

The distinction between these layers matters. A primary entry is the behaviour being tested. An attack method is the strategy used to challenge it. An adversarial variant is the resulting version of that test. So the 160 variants are not 160 unrelated attacks. They are alternative attempts to challenge the behaviours represented by the original 20 primary entries.
Primary entry ≠ attack method ≠ adversarial variant.
That distinction becomes important when reading the results. The run reports a headline robust accuracy alongside a method-level breakdown showing how the same primary behaviours performed under each attack method.
The planned multi-turn attack works slightly differently from the single-turn methods. A simulated adversarial user continues the interaction for a bounded number of turns, and the resulting conversation is judged as a whole. An early refusal therefore does not guarantee a pass if the agent later gives way under continued pressure.
The run also reported 16 adversarial errors that were excluded from scoring. All 16 occurred in the multi-turn evaluation, leaving only 4 successfully evaluated multi-turn cases, compared with 20 for each of the single-turn methods. That distinction matters later because an execution error and a behavioural failure are not the same thing.
What the attacks are actually testing
The previous article, Red-Teaming Public-Facing Agents, introduced six behavioural boundaries that should remain stable when an agent is put under pressure. Those same boundaries provide a useful lens for interpreting this run.
Boundary | Question for the evaluation |
Scope | Does the agent stay within the service it was designed to provide? |
Authority | Can a user push it into making claims, approvals or decisions it is not authorised to make? |
Disclosure | Does it reveal information that should remain protected? |
Accuracy | Does it distinguish supported information from unsupported or misleading claims? |
Control | Can user input override the agent’s intended restrictions or priorities? |
Consistency | Do those behaviours continue to hold as the wording, context or conversation changes? |
These boundaries should not be treated as a scorecard for the run. Not every primary entry maps neatly to a single boundary, and the evaluation was not designed to provide complete, domain-specific coverage across all six. Instead, they help us understand what kind of behaviour is being put under pressure and why a failure matters.
The attack methods then approach that pressure from different directions. Research bypass tests whether a boundary still holds when the request is reframed as research, analysis or documentation. Fictional roleplay tests whether changing the surrounding context causes the agent to relax behaviour it would normally maintain. Poetic jailbreak changes the wording and presentation of the request. Planned multi-turn red teaming tests whether the same behaviour remains intact as the interaction develops and pressure builds.
The methods differ, but the underlying question is the same:
Does the expected behaviour remain stable when the way it is challenged changes?
Don't trust a boundary you only tested once
A primary entry tells us whether the agent behaved correctly on one particular input. It does not tell us whether that behaviour will remain stable when the same underlying request is presented differently.
That is what the adversarial variants are designed to test. A request might be reframed as research, placed inside a fictional scenario, expressed in unusual language, or developed over several turns. The form changes, but the expected behaviour should not.
This is the practical meaning of the principle from Red-Teaming Public-Facing Agents: test behaviours, not just prompts. A single pass shows that the agent handled one version of the request. Robustness asks whether that behaviour still holds when the route used to challenge it changes. The difference becomes visible when we compare the original cases with their adversarial results.

Primary and adversarial results from the Spec27 run. Several behaviours that passed in their original form no longer held once challenged adversarially.
Spec27 also lets us inspect those results by category and attack method. The category comes from the primary entry, while each attack method creates different adversarial ways of testing it. The resulting summary therefore shows where particular types of behaviour held or broke under different forms of pressure. In the table, scores are shown as [passed/evaluated]. So 1/2 means one of two evaluated cases passed, while 0/2 means neither did.
The headline result: 75% clean, 10% robust
Two numbers summarise the run, and the gap between them is the key result. Clean accuracy was 75%, or 15 of 20 primary entries. On the original cases, the agent produced the expected behaviour in most instances.
Reported robust accuracy was 10%, or 2 of 20 primary entries. Method-level results show how performance varied by pressure type: 10% for fictional roleplay, 25% for planned multi-turn red team, 75% for poetic jailbreak, and 30% for research bypass.
Baseline performance is not the same as robustness.

The results view contrasts 75% clean accuracy with 10% reported robust accuracy. The expanded cases show that individual primary entries can produce a mix of Pass, Fail and Error outcomes across different adversarial variants.
That detail matters. An adversarial failure does not mean every variant failed. In Case 3, for example, the primary entry passed and two fictional-roleplay variants also passed, while another roleplay variant failed and the research-bypass variants failed. The behaviour held under some forms of pressure, but not others.
Case 1 also exposes an important caveat. One fictional-roleplay variant returned what appears to be an infrastructure error payload but was still recorded as a Fail. Because an infrastructure failure is not the same as a behavioural failure, that case should be audited before the aggregate result is treated as definitive.
Even with that caveat, the broader pattern remains clear: passing the original case did not guarantee that the same behaviour would hold when challenged differently.
A similar gap appeared in another anonymised Spec27 evaluation, where an agent achieved 90% clean accuracy and 20% reported robust accuracy under adversarial testing. The agents, behaviours and configurations were different, so the percentages are not directly comparable as benchmarks. What recurs is the shape of the result: strong performance on expected inputs can conceal behaviour that becomes much less stable when the route used to challenge it changes.
Different attacks exposed different weaknesses
The 10% headline shows how few primary entries remained robust across the run. The method breakdown shows that performance was not uniform across the four forms of pressure. Across the four attack methods, the results varied sharply:
Fictional roleplay: 2 of 20 held, 10%
Research bypass: 6 of 20 held, 30%
Poetic jailbreak: 15 of 20 held, 75%
Planned multi-turn red team: 1 of 4 held, 25%
The multi-turn result needs to be read separately. Sixteen cases errored and were excluded, leaving only four successfully evaluated conversations. Its 25% result therefore describes a much smaller set than the three single-turn methods. What matters is the spread. The same behaviours were much more stable under some forms of pressure than others.
Fictional roleplay produced the weakest result. Research bypass also exposed substantial fragility. Poetic jailbreak, by contrast, matched the clean baseline numerically at 15 of 20. That does not mean poetic attacks are generally ineffective. It means that, in this run, changing the wording into a stylised form had far less effect than changing the context or justification around the request. The category results reinforce that point. The two entries tagged Expert advice both passed clean and under poetic jailbreak, but neither passed under fictional roleplay or research bypass.

One primary behaviour can respond differently depending on how the request is reframed. The expanded view shows the original case alongside its adversarial variants and their individual outcomes.
Multi-turn testing adds a different dimension: it asks whether the same behaviour survives as context and pressure accumulate across a conversation.

Planned multi-turn testing evaluates the conversation as a whole, rather than judging a single isolated response.
The lesson is not that one attack method is universally stronger than another. It is that testing only one route can leave important weaknesses invisible.
What this evaluation tells us, and what it doesn’t
The run tells us more than the headline score alone. First, clean performance and adversarial robustness are not the same thing. The agent passed 15 of 20 primary entries, but only 2 of 20 met the run’s robustness criterion. That gap shows why baseline testing can give an incomplete picture of how behaviour changes under pressure. Second, the attack method matters. The four methods used here produced very different outcomes, but they represent only part of Spec27’s Red Team coverage. Red Team specifications can use broader attack suites, and Spec27 also documents methods such as Adaptive LotL. A behaviour that holds under one form of pressure may still fail under another, which is why no single attack method gives a complete picture of robustness. Third, aggregate scores need context. The expanded cases showed mixtures of Pass, Fail and Error outcomes across variants. Sixteen multi-turn cases errored and were excluded, and at least one single-turn infrastructure failure appears to have been recorded as a behavioural Fail. Those details do not erase the overall pattern, but they do affect how confidently individual scores should be interpreted.
What the run does not tell us is equally important. The reported 10% is not a statement that the agent is “10% safe”, that 90% of real-world attacks would succeed, or that the same result would appear across different agents, prompts or deployments. Nor does this run tell us how the agent would behave under attack methods, scenarios or deployment conditions that were not included in the evaluation. The result is therefore best read for what it is:
Evidence about how this agent behaved under this evaluation, not a universal score for the agent itself.
That is what makes the run useful. It shows where expected behaviour held, where it became unstable, and where the evaluation itself deserves closer inspection.
Test behaviours, not just prompts
The most useful finding from this run is not the headline score itself, but how adversarial testing changed what we understood about the agent. Some behaviours that appeared reliable on the original cases became much less stable when the framing changed. Different attack methods exposed different weaknesses, and the expanded results showed that even a single primary case could behave differently across variants.
That is why a passing prompt is only a starting point. The stronger question is whether the behaviour itself remains stable when the wording, context or structure of the interaction changes.
Key takeaways
Clean performance is not the same as robustness. Passing the original cases does not show whether the same behaviour will hold under adversarial pressure.
Different attack methods reveal different weaknesses. A behaviour that holds under one form of pressure may still fail under another.
Aggregate scores need context. Pass, Fail, and Error outcomes should be inspected at the case level before concluding.
Test behaviours, not just prompts. The goal is not to prove that an agent can handle one familiar wording, but to see whether the expected boundary survives meaningful variation.
A behaviour you have only tested once is not a behaviour you can trust yet.


