Agentic AI  ·  17 Sep 2026  ·  2 min read

Agent evaluation becomes a separate line item in RFPs

Updated 16 Sep 2026: added Aulric Labs opening its harness to non-platform customers

key facts
  • Buyers have started scoping the agent evaluation harness separately from the agent itself.
  • Previously it was assumed to be part of the build, and frequently was not delivered.
  • Two vendors now sell evaluation as a standalone product that works with someone else’s agent.
  • Only 6 of the 19 firms in our Agentic AI area met our production bar, and evaluation was the common gap.

An agent evaluation harness is the test suite that catches a bad run before a person does. For two years it was assumed to be part of whatever a vendor built. It usually was not, and buyers have now started pricing it separately.

Why it was missed

Evaluation is invisible in a demo. A pilot with twenty happy-path runs looks identical whether or not a regression suite exists behind it. The gap only appears at the first production change, when nobody can say what a new prompt or model version broke. By then the harness is an emergency project rather than a line item.

What is changing in the documents

The RFPs we have seen this quarter now ask three things directly: how many test cases exist at handover, who writes new ones, and who owns them afterwards. That last question is the one to watch. Evaluation test cases encode your business rules, and a vendor that owns them owns a large part of your switching cost.

6 of 19firms met our production bar
2vendors now sell evaluation alone
3questions now standard in RFPs

The vendor-agnostic option

Halvern Systems has built its practice on making someone else’s agent defensible, and on 16 September Aulric Labs opened its evaluation harness to customers who are not on its platform. That is a meaningful shift: it turns evaluation from a bundled assumption into a component you can buy from a different supplier than the agent.

What to do if you are mid-project

Ask your current vendor for the test case count today, in writing. If the answer is a number under fifty for anything touching money or customers, treat that as the finding and scope it before the next release rather than after the next incident.

Agentic AIevaluation harnessAulric Labs
Share

Corrections: none issued for this article.

Keep reading

All blog posts
© 2026 Top AI Firms. Independent research \u{2014} we take no payment for placement.How we assess a firm · About ·