Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark
Source:
ventureBeat
June 10, 2026 · 16:16
Researchers from the University of California, Berkeley's Center for Responsible, Decentralized Intelligence (RDI), alongside an advisory committee of over 300 domain experts, have launched Agents’ Last Exam (ALE)—a grueling new benchmark built to measure whether artificial intelligence can actually execute economically valuable, long-horizon professional workflows.In a shocking upset, OpenAI’s GPT-5.5 from April, operating through the Codex harness, secured the absolute top spot on the new ALE Leaderboard with a 24.0% pass rate, beating Anthropic's highly anticipated, brand new Mythos-class Claude Fable 5 model released just yesterday, which came in third with a score of 22.0%.Rather than testing models on isolated coding puzzles, ALE is explicitly designed as an instrument to close the g…
The original article opens on the publisher's website.
More from Automotive
View topic →Larry Page’s flying car company Pivotal loses its CEO
techcrunch
Sep 1, 2026 · 16:59
SEC proposes transfer agent rule, sets event to figure out round-the-clock U.S. trading
coinDesk
Sep 1, 2026 · 16:55
China dissented from G20 statement opposing 'cheap exports' flooding market, Bessent says
cnbc
Sep 1, 2026 · 16:52
The Range Rover Electric: Specs, Price, Availability
wired
Sep 1, 2026 · 16:01
North Carolina Rep. Chuck Edwards is formally censured over harassment allegations
nprNews
Sep 1, 2026 · 15:18