top of page
Spec27 Blog
Insights on AI agent validation, LLM evaluation, regression testing, and safer deployment of AI systems in production.
We write about how teams can move beyond manual prompt testing and vibes-based reviews toward repeatable, evidence-driven validation for AI agents and applications.


Building Robust AI Support Agents: Practical Lessons Across the Four Routes
Choosing a route into AI support is only the beginning. This article explains how robustness changes across AI-agent-first platforms, service platforms with AI, CRM/contact-centre ecosystems, and in-house builds -- and why every route needs repeatable testing across behaviour, policy, tools, workflows, safety, and change.
Brain John Aboze
2 days ago10 min read


Spec27 Evals: How Fin Intercom and Zendesk Customer Support Agents Performed on the Same Knowledge Base and Policy Rules
Intercom Fin and Zendesk were tested against the same knowledge base, policy rules and realistic customer-input variations. Here’s where their reliability held—and where it began to break.
Brain John Aboze
Aug 119 min read


Why AI Agents Fail After Model Upgrades
Model upgrades are not simple component swaps. They can change how AI agents escalate, refuse, call tools, format outputs, and respond to edge cases. This post explains where drift hides and how teams can re-validate agents before switching models.

Jovanca Garnadi
Aug 46 min read


Third-Party AI Support Agents: How to Navigate the Vendor Market
AI support agents are moving from novelty to infrastructure. This guide maps the third-party vendor market, including what these agents do, why adoption is accelerating, where they perform best, and what risks buyers should consider before deployment.
Brain John Aboze
Jul 286 min read


Can Reasoning Make Sentiment Models More Robust?
Sentiment classification looks like a simple task: take a financial statement and return a positive or negative label. It is used in research and production for financial-news monitoring, reporting workflows and downstream analysis. Clean accuracy, however, tells us only how a model behaves on the wording it was given. If the same financial relationship is expressed with different words, a different sentence structure or an added concessive clause, the classification should r

Brian Formento
Jul 237 min read


From Test Prompts to Validation Specs: Uplevelling AI Agent Testing
From Test Prompts to Validation Specs: Uplevelling AI Agent Testing

Jovanca Garnadi
Jul 217 min read


How to Automate AI Agent Validation
In a hurry? There's a preloaded example you can run in CI today, jump to how you can try this yourself.
Michael Wagstaff
Jun 235 min read


Spec27 at Google DeepMind Startup Session
On 11 June, we joined the Google DeepMind Startup Sessions at the Ministry of Sound in London – a morning of talks, demos, and live pitches bringing together AI startups and the Google DeepMind community. The lineup included a talk by Dr Raia Hadsell, a recap of Google’s annual developer conference, and a pitch showcase featuring around ten selected startups. We were glad to be one of them, presenting Spec27, our product for validating how AI agents behave before deployment a

Jovanca Garnadi
Jun 162 min read


BDR, ADR, PRD, WTF: Capturing Decisions for Humans and AI Alike
At AI Engineer Europe 2026, Spec27 Principal Engineer Michal Cichra explored why engineering decisions need to be captured clearly for both human teams and AI systems. Watch the talk and learn how BDD, ADRs, and PRDs connect to specification-driven validation.

Jovanca Garnadi
Jun 51 min read


How to Evaluate Third-Party AI Agents Without Code Access
Third-party AI agents can help teams deploy faster, but they also create a validation challenge: the buyer owns the operational risk without full visibility into the internals. This guide explains how to evaluate vendor-built AI agents from the outside using specifications, realistic test cases, pass criteria, re-validation, and evidence.

Jovanca Garnadi
Jun 38 min read


The Five Types of Enterprise AI Agents
How many agents do we have? And when do they start hiring their own agents? AI Agents are taking the enterprise by storm (at least on executive roadmaps!) and many leading AI tech companies are rolling out frameworks and platforms for AI Agents. Navigating through what works and what doesn't can be hard to keep track of. There's huge value in a lot of the deployments, but they also each present their own challenges. As the number of agents increases the interactions between t
Steven Willmott
May 138 min read
bottom of page