Red-Teaming Public-Facing Agents: Quick Wins to Make Your Agent Safer

The moment your agent goes public, you stop choosing your users.
Some of those users will be clear and cooperative. Some will be vague, confidently wrong, in a hurry or persistently looking for an exception. A few will be deliberately adversarial. Your agentic system receives all of those interactions through the same interface. Most teams begin by evaluating whether an agent does what it is supposed to do. That is the right place to start, and it's where most evaluation stops. Red teaming asks the opposite questions: can a user make it do something it was never supposed to do?
In a previous article on Building Robust AI Support Agents, we covered robustness, which evaluates whether an agent’s behaviour holds up when conditions change. Red-teaming takes that further: instead of varying conditions and hoping the behaviour holds, it deliberately tries to make it fail. This article walks through eight adversarial checks, the four approaches teams use to run them, and how teams turn failures into repeatable evaluations.
What red-teaming gives you and what it doesn't
Before we get into the checks, it helps to be clear about what this exercise can actually give you. Red-teaming is often treated as a broad measure of AI safety, but its value depends on understanding both what it can reveal and where its limits are.
Red-teaming can:
surface high-impact weaknesses and convert meaningful findings into concrete failure cases that stakeholders can understand, address, and turn into regression tests.
show whether your agent stays within its intended scope when users push beyond it;
reveal which boundary failed and how it was reached, which is far more actionable than an aggregate safety score;
expose failures that only appear under sustained multi-turn pressure, the kind that single-prompt evaluations structurally cannot catch;
compare builds, models or vendors against the same adversarial scenarios.
It cannot:
prove the absence of vulnerabilities;
guarantee that a pass will continue to hold over time. Agent behaviour can vary between runs because modern AI systems are probabilistic and influenced by factors such as model behaviour, system configuration, context and external dependencies. Critical evaluations should therefore be repeated and tracked to ensure reliability;
replace human review, authorisation review, penetration testing, or infrastructure and architecture review;
fully assess risks introduced by tools, permissions, customer data or external systems unless those areas are explicitly included in the exercise;
cover every language, phrasing, cultural framing and conversation path a real public deployment will meet;
keep evaluations valid over time. A result only reflects the system tested at that point in time and can become outdated after model updates, prompt changes, or knowledge-base modifications. Continuously maintain and rerun evaluations as the agent evolves.
The value of red-teaming is therefore not in claiming that an agent is "safe" because it passed a set of attacks. Its value is in giving teams evidence of where behaviour holds, where it breaks, and what to test again after system changes.
For a deeper understanding, a great reference is the OWASP GenAI Red Teaming Guide. It provides a broader framework covering model behaviour, implementation, infrastructure and runtime risks, and is a strong foundation for teams building a more complete red-team programme. This article is deliberately narrower, focusing on the conversational and behavioural boundaries of public-facing agents, rather than a full security assessment of an agent system.
Defining the agent’s blast radius before you red-team
A good red team does not begin by asking “how can we break the agent?” It begins by asking “what happens if the agent fails?”
Reliability engineers use the term blast radius to describe the scope and extent of impact that a single failure can cause. AWS applies this principle through fault isolation, where the Well-Architected Framework recommends boundaries that “limit the effect of a failure within a workload to a limited number of components.” In AI security, the same principle increasingly applies to agent behaviour: Microsoft’s AI red teaming research recommends that approval requirements for agent actions scale with factors such as action reversibility and blast radius. Applied to a public-facing agent, mapping the blast radius starts with four questions:
Question | Why it matters |
What can the agent disclose? | Hidden instructions, internal processes, other users’ information, or authoritative-sounding statements that should not be made. |
What can it change or execute? | Refunds, account updates, bookings, messages sent on a user’s behalf, or other real-world outcomes. |
What data and permissions can it access? | Only the current user's records, or records that belong to others. |
Which business or safety rules must it never violate? | Eligibility rules, regulated guidance, approval requirements, escalation policies and other non-negotiable boundaries. |
Understanding an agent’s blast radius helps determine what to test and where the greatest risks may lie. Red teaming isn't simply about finding any failure; it's about identifying failures that could have meaningful impact within the agent’s operating environment. An agent that can issue refunds, modify accounts, or access customer records has a much wider blast radius than an information-only agent. Its evaluation must therefore extend beyond conversation into tools, permissions, and data access.

This is why AI agents are increasingly evaluated not only on what actions they can perform, but also on the accuracy and boundaries of the information they provide. A single confident but incorrect response can create operational, compliance, and trust risks. So the focus shifts from preventing unauthorised actions to three central questions:
Can it make statements or give advice it should not?
Can it be pushed outside the support role it was designed to perform?
Can it disclose anything about its own rules or information it should not reveal?
Understanding the agent’s blast radius tells us where failures would matter most. The next step is identifying the behaviours that must remain stable when those areas are challenged. These behaviours define the boundaries that red-team testing should evaluate:
Boundary | What to check |
Scope | Does the agent stay within the service it was designed to provide? |
Authority | Can users persuade it into permissions or approvals they do not actually have? |
Disclosure | Does it reveal hidden instructions, internal processes, other customers’ information or anything else it should not? |
Accuracy | Does the agent provide reliable information and avoid presenting unsupported claims as facts? |
Control | Can user inputs override the agent’s intended behaviour, priorities or restrictions? |
Consistency | Do the agent’s boundaries remain intact across longer conversations, changing contexts, and repeated interactions? |
Each check that follows targets one or more of these boundaries. That is deliberate. Effective red-teaming starts by defining what the agent must keep doing correctly, even when users push against it. Writing those behaviours down before testing is what separates a structured evaluation from a collection of interesting failures.
Quick Boundary Checks for Public-Facing Agents
The checks below are the first evaluations we recommend for any public-facing agent. They are not a complete security assessment, nor do they replace deeper testing of tools, permissions, or data access for agents with greater capabilities. Their purpose is narrower: to verify whether the agent remains within the boundaries that define its role. Each check defines the boundary being tested, how to challenge it, and what success looks like. This turns red teaming from a search for isolated failures into a structured evaluation of whether the agent continues to behave as intended under pressure. The checks progress from simple single-turn tests to more complex scenarios. The specific boundaries should be adapted to the agent’s purpose, capabilities, and risk level.
A refusal alone does not prove an agent is safe, and a pass can still hide poor behaviour. For each test, record the boundary being evaluated and whether it held, but also capture what a binary score misses: did the agent remain useful for the legitimate part of the request, and did it provide a safe next step?
An agent that avoids unsafe behaviour by refusing everything may appear secure while failing its actual purpose. The goal is not simply to find techniques to break the agent. It is to verify that the agent continues to behave as designed when challenged. Teams approach this challenge differently depending on the maturity of their evaluation process. Some begin with manual exploration, while others build automated systems that continuously test agent behaviour as it changes.
Four approaches teams use to red-team AI agents
Teams rarely begin with a fully automated red-team programme. Most evolve through a series of approaches, each solving a different problem. The progression usually moves from discovery to repeatability, to broader coverage, and finally to testing behaviour under sustained pressure.
Manual testing is often the starting point. Human testers explore unexpected behaviours, try to move the agent outside its intended role, and identify failure patterns that automated systems may not anticipate. This is especially useful early in development, and guides such as the OWASP GenAI Red Teaming Guide and the CSA/OWASP Agentic AI Red Teaming Guide help teams structure that work. The main limitation is consistency: manual findings are valuable for discovery, but can be difficult to reproduce systematically.
Once teams find a failure, they often turn it into a fixed test case. Instead of relying on someone remembering a previous weakness, the same scenario becomes a repeatable evaluation that can be rerun after model, prompt or workflow changes. This makes fixed cases particularly valuable for regression testing. Tools such as Promptfoo and Garak are useful here. The trade-off is that fixed cases only cover failures the team already knows about.

Generated variation goes beyond known examples by automatically creating alternative ways to reach the same failure. A boundary that holds under one condition may behave differently under paraphrasing, role-play, urgency or another conversational framing. This helps teams test whether a boundary is genuinely robust rather than tied to a particular wording. Tools such as Promptfoo and Spec27 support this kind of adversarial variation. In Spec27, attack methods generate variants of primary test entries so the same expected behaviour can be challenged in different ways. The limitation is that generated variation still needs a clear expected behaviour. If you have not defined what the agent should or should not do, generating more versions of a prompt only creates more inputs, not necessarily more useful evidence.
The next stage combines adversarial variation with adaptive multi-turn testing. Rather than relying on isolated, single-turn attack techniques, these approaches examine how agent behaviour changes as context builds and pressure develops across a conversation. This matters because some failures only emerge after several exchanges. A 2026 Scientific Reports study found that multi-turn stress tests exposed vulnerabilities that single-turn testing missed, with all high-severity failures in that study occurring during sustained interactions. Spec27 applies this principle through a simulated adversarial user that probes the agent across a bounded conversation and evaluates whether the agent crosses the defined failure boundary at any point in the exchange.
These approaches are not competing methods, but stages in a more mature red-team process. Manual testing discovers important failures. Fixed cases make those failures repeatable. Generated variation expands the ways each boundary is challenged. Adaptive multi-turn testing then checks whether those boundaries still hold as the conversation develops.
From Red-Team Findings to Continuous Evaluation
Finding a failure is only the beginning. The value of red teaming comes from ensuring the same weakness doesn't return after a model update, prompt change, workflow adjustment, or knowledge-base revision. A useful red-team finding should enter a continuous improvement loop.

When an agent crosses a boundary, teams should first understand why. Did unclear scope, system instructions, retrieved content, missing safeguards, or conversation flow cause the failure? Addressing the root cause is more effective than simply blocking the trigger that exposed the weakness. Once fixed, the interaction should become a reusable evaluation case. Future versions of the agent can be tested against the same scenario to confirm improvements don't create new failures.
This is what turns red teaming from a one-time security exercise into an ongoing engineering practice.
A practical starting point is to preserve failures around the behavioural boundaries that matter most: whether the agent stayed within its intended scope, respected authority and permissions, avoided unauthorised disclosure, maintained accuracy, resisted manipulation, and remained consistent as conversations evolved. This is not a complete security assessment, but a foundation for identifying weaknesses, converting them into repeatable evaluations, and ensuring fixes remain effective as the agent changes.
Conclusion
Public-facing agents do not need sophisticated jailbreaks to fail. A small wording change, a claimed authority, a confident falsehood, or sustained pressure can push behaviour beyond intended boundaries. Red teaming is not about finding the cleverest attack. It is about defining the behaviours that must remain reliable, testing them deliberately, and turning failures into evaluations that strengthen the system.
Define the blast radius. Test the boundaries that matter. Preserve what breaks. Then test again as the agent evolves.
Key takeaways
Start with impact, not attacks. Understand what the agent can reveal, change, access or violate before deciding what to test.
Test behaviours, not just prompts. A boundary is only robust if it survives different wording, contexts and multi-turn pressure.
Make failures reusable. The strongest red-team programmes turn discovered weaknesses into permanent evaluations.
Red teaming is continuous. Agents evolve, and their evaluations must evolve with them.
