Enterprise AI teams are granting agents increasing autonomy just as confidence in automated testing methods is waning. According to a June 2026 survey of 157 enterprise respondents, half of enterprises have deployed AI agents or large language model features that passed internal checks but still caused customer-facing failures, with a quarter experiencing multiple such incidents. Despite this, 66% of companies allow some deployments without human review or plan to within the year, even though only 5% fully trust these automated evaluations. This gap—where agent autonomy grows faster than verification capabilities—is leading enterprises to ship agents first and develop control systems later.
Traditional software testing falls short for AI agents, which take varied actions and outcomes can differ each time. The biggest reason enterprises distrust automated evaluation is poor alignment with real-world results, followed by bias, explainability, and privacy concerns. Experts emphasize the importance of repeatability, running multiple tests and evolving evaluation sets based on incidents. While some low-risk tasks can tolerate more autonomous agents, high-risk actions require stricter oversight. Larger companies tend to push for zero-human review faster but also face more failures. The key takeaway for leaders is that removing humans from the loop increases uncertainty unless strong evaluation and regression testing become a priority. The market drives increasing autonomy, but success will favor those balancing speed with reliability.