Home / Blogs / Agentic AI
Agentic AI · 17 Sep 2026 · 2 min read
Agent evaluation becomes a separate line item in RFPs
Updated 16 Sep 2026: added Aulric Labs opening its harness to non-platform customers
key facts
- Buyers have started scoping the agent evaluation harness separately from the agent itself.
- Previously it was assumed to be part of the build, and frequently was not delivered.
- Two vendors now sell evaluation as a standalone product that works with someone else’s agent.
- Only 6 of the 19 firms in our Agentic AI area met our production bar, and evaluation was the common gap.
An agent evaluation harness is the test suite that catches a bad run before a person does. For two years it was assumed to be part of whatever a vendor built. It usually was not, and buyers have now started pricing it separately.
Why it was missed
Evaluation is invisible in a demo. A pilot with twenty happy-path runs looks identical whether or not a regression suite exists behind it. The gap only appears at the first production change, when nobody can say what a new prompt or model version broke. By then the harness is an emergency project rather than a line item.
What is changing in the documents
The RFPs we have seen this quarter now ask three things directly: how many test cases exist at handover, who writes new ones, and who owns them afterwards. That last question is the one to watch. Evaluation test cases encode your business rules, and a vendor that owns them owns a large part of your switching cost.
The vendor-agnostic option
Halvern Systems has built its practice on making someone else’s agent defensible, and on 16 September Aulric Labs opened its evaluation harness to customers who are not on its platform. That is a meaningful shift: it turns evaluation from a bundled assumption into a component you can buy from a different supplier than the agent.
What to do if you are mid-project
Ask your current vendor for the test case count today, in writing. If the answer is a number under fifty for anything touching money or customers, treat that as the finding and scope it before the next release rather than after the next incident.
Corrections: none issued for this article.
Read this next
How to tell a real agent from a wrapped prompt
Keep reading
Agentic AI
How to tell a real agent from a wrapped prompt
Almost every firm in our Agentic AI area now uses the word agent. Some are shipping…
Data Science
How leading enterprises are using supply chain analytics to stay ahead of disruption
Discover how leading enterprises use supply chain analytics, predictive forecasting, and real-time visibility to stay ahead…
AI Consulting
How healthcare leaders are choosing AI consulting partners that actually deliver
Learn the five signals that separate a real healthcare AI consulting partner from an expensive pilot,…