top of page

How to Compare Two AI Support Agents Fairly

Writer: Jovanca Garnadi
Jovanca Garnadi
Sep 8
1 min read

A useful comparison requires more than reviewing a few polished answers. Both agents need to face the same scenarios, business rules and scoring criteria.


Putting two support agents through the same tests


We evaluated Intercom Fin and Zendesk’s AI support agent using the same customer-service scenarios.


The tests examined whether each agent could continue to follow the relevant business rules when customer requests became messy, varied or incomplete.


This matters because a convincing answer to one prompt tells us very little about how dependable an agent will be across a wider range of interactions.


For a comparison to be meaningful, the agents should be tested using:

  • the same customer scenarios;

  • the same datasets and available information;

  • the same business rules;

  • the same input variations;

  • the same scoring criteria.


This creates a more consistent basis for identifying where each agent performs well and where its behaviour begins to break down.



Making validation repeatable


A one-off evaluation is useful, but agent behaviour can change whenever a team updates a model, prompt, knowledge source, tool or workflow.


Our practical guide to automating AI agent validation explains how expected behaviour can be converted into tests that run repeatedly as the underlying system changes.


This update was originally shared during AI in Financial Services London on 8–9 September 2026, where the Spec27 team was discussing agent testing, robustness and deployment evidence.


bottom of page