In a survey of 157 enterprises, many are giving AI agents increased autonomy without fully trusting the tests designed to ensure their reliability. Half of these companies have deployed AI features that passed internal evaluations but later failed in real-world customer interactions. Despite this, two-thirds allow AI agents to be deployed to production automatically, without human oversight. The biggest issue is the “evaluation gap”—a disconnect between the autonomy granted to AI agents and the trust in the evaluations meant to oversee them. Only 5% of organizations fully trust automated evaluations, citing poor alignment with real-world outcomes as the primary concern. Moreover, monitoring in production often checks only if the AI is functioning rather than verifying the correctness of its outputs. The market for evaluation tools remains fragmented, with many companies relying on provider-native tools or none at all, though investment in human review and observability is growing. While enterprises are moving towards greater automation, they continue to invest in oversight, highlighting a cautious approach amidst burgeoning autonomy. This points to a critical need for evaluations that better reflect actual performance to ensure AI agents can be trusted as their independence grows.
Back