OpenAI: 2 Settings Tripled AI Benchmark Scores

5d ago·0:00 listen·Source: OpenAI

Summary

OpenAI has found that two specific settings tripled their AI model's scores on the ARC-AGI-3 benchmark. The model, GPT-5.6 Sol, initially scored only 7.8% on this benchmark of 2D puzzle games. Here's the thing: GPT-5.6 Sol had previously solved complex math problems and beaten games like Pokémon FireRed. What's interesting is that the low scores on ARC-AGI-3 were not due to the model's inherent ability. Instead, they were caused by how the benchmark's API settings were configured. By enabling "retained reasoning" and "compaction," which are used in ChatGPT and Codex, the model's scores significantly improved. These two settings tripled scores and cut output tokens by six times on the public task set. The bottom line is that how AI models are tested can dramatically impact their apparent performance.

Read the full article on OpenAI

This is an AI-generated audio summary. Always check the original source for complete reporting.

Share
Keep Listening