Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

GPT-5.5 Triumphs Over Claude Fable 5 on Challenging Agents’ Last Exam Benchmark

A team from UC Berkeley’s Center for Responsible, Decentralized Intelligence (RDI), with over 300 industry experts, has introduced the Agents’ Last Exam (ALE), a demanding benchmark designed to evaluate AI’s capability to handle complex, economically valuable professional workflows. In a surprising result, OpenAI’s GPT-5.5, using the Codex agent harness, achieved the highest pass rate of 24.0% on ALE, outperforming Anthropic’s newly released Claude Fable 5, which placed third with 22.0%. ALE is groundbreaking in its approach, requiring AI agents to perform multi-step tasks across five functional areas—reasoning, visual perception, orchestration, tool use, and runtime execution—within realistic software environments. Unlike previous benchmarks that allowed “cheating” through hidden answers or unreliable grading, ALE employs a rigorous evaluation protocol, relying mostly on deterministic code-based verification. Covering 1,490 tasks from 55 industry sectors, ALE reflects real-world professional demands and reveals that even the best AI models today struggle with long-horizon workflows, as demonstrated by low pass rates on the most difficult tasks. The benchmark maintains integrity through a carefully managed data release, preventing models from memorizing test questions. This new standard offers a clear-eyed view of current AI limits and challenges developers to build agents truly ready for professional work.

Venturebeat
Venturebeat