Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

The Evaluation Discrepancy in Enterprise AI: Autonomy Outpaces Trust Yet Deployment Advances

Across 157 enterprises surveyed, AI agents are being granted increasing autonomy while confidence in the evaluations that govern this autonomy remains low. Half of the organizations admitted to deploying AI agents that passed internal tests but then failed in real-world customer scenarios. Only 5% fully trust automated evaluations, mainly due to poor alignment with actual outcomes. Yet, two-thirds of enterprises either already allow or are engineering towards fully automated deployments without human oversight for low-risk AI agents. This gap between granted autonomy and trust in evaluations is widening, raising concerns about scaling failures. The evaluation tools landscape is fragmented and mostly provider-driven, with many organizations lacking dedicated evaluation tooling or real-time quality checks on live outputs. Investment priorities show a paradox: companies are reducing human-in-the-loop deployments while increasing spending on human review and production observability. The AI evaluation market is still emerging, with most enterprises planning to adopt new or additional platforms soon. Ultimately, enterprises face a critical challenge where autonomy outpaces assurance, necessitating evaluations that more accurately reflect real-world performance to avoid costly failures.

Venturebeat
Venturebeat