Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
info@treatmybrand.com

Support


Monday to Friday
8AM to 8PM
support@treatmybrand.com
Back

At Waymo, AI Projects Are Defined by Rigorous Evaluations, Not Just Performance Milestones

Waymo, Alphabet’s self-driving car division, faces unique challenges where AI performance directly impacts safety on the streets. Their approach to AI development prioritizes continuous and comprehensive evaluation over simply achieving strong model results. Manasi Joshi, Waymo’s director of engineering, shared insights on this at VB Transform 2026, highlighting their strategy of “eval-centric development.”

This means that a project’s readiness is judged by the maturity of its evaluations, not just the model’s output. Waymo conducts ongoing assessments during and after model training, combining real-world driving data with extensive simulations to ensure safety and effectiveness. Evaluations continue post-launch to adapt to evolving conditions and maintain high standards linked directly to key business outcomes.

Safety remains paramount, with detailed attention to rare and hazardous scenarios like construction zones and vulnerable road users. Human oversight remains crucial; automated systems do not make release decisions alone. Efficiency in computing resources is balanced carefully against reliability to meet growing demands without sacrificing quality.

Internally, AI agents assist engineering teams, and these agents themselves undergo rigorous testing to confirm accuracy and trustworthiness. Joshi emphasized that successful deployment of agentic AI depends on clear objectives, representative data, continuous evaluation, scalable infrastructure, and accountable human leadership to build and maintain trust.

Summary:
Waymo sets an industry benchmark by embedding continuous, rigorous evaluation into AI development, focusing on safety and trustworthiness beyond mere model performance. Their approach offers valuable lessons for enterprises deploying AI, emphasizing ongoing testing, human oversight, and linking evaluations to real-world outcomes.

Venturebeat
Venturebeat