AI Visual Perception: Models Struggle, GPT-5.6 Sol Leads
Summary
New research shows that AI models still struggle significantly with visual perception. A new test called PerceptionBench, developed by Moonshot AI, evaluates how well multimodal models "see," separate from logical reasoning or external knowledge. What's interesting is that all leading models, including GPT-5.6 Sol, Kimi K3, and Claude Fable 5, showed major weaknesses. The test reveals that many supposed logical errors in AI are actually caused by poor visual perception. No frontier model reached 60 percent accuracy, with GPT-5.6 Sol leading with 59.7 percent. The PerceptionBench breaks visual perception into ten specific sub-skills, like counting, depth, and object recognition, based on real-world errors. This matters because it highlights a fundamental limitation in current AI, suggesting that improving their visual understanding is crucial for more reliable performance.
This is an AI-generated audio summary. Always check the original source for complete reporting.