Claude Fable 5 Dominates MirrorCode; GPT-5.5 Scores Drop
Summary
Claude Fable 5 has achieved the highest solve rate on the MirrorCode benchmark, scoring 64%. This leads its closest competitor by 44 percentage points. What's interesting is how other models performed. GPT-5.6 Sol scored 20%, and GPT-5.4 was at 16%. GPT-5.5 landed in last place with just 10%. This 10% score for GPT-5.5 is a significant change. In the June 2026 paper introducing MirrorCode, GPT-5.5 scored 44%. The leaderboard configuration is much harder than the paper's original test. MirrorCode is a long-horizon autonomous coding benchmark where AI agents reimplement software programs from scratch. Agents cannot access the internet or original source code. Solutions must pass 100% of both visible and hidden tests. The updated leaderboard uses harder programs and different implementation languages, like Go and Ada. Epoch AI states that these scores are not directly comparable to the paper's results. The bottom line is that a model strong on easier tasks in popular languages may not perform the same on these more challenging configurations. This shows how crucial testing configurations are for evaluating AI models.
This is an AI-generated audio summary. Always check the original source for complete reporting.