BDH-CQ AI: 29.5% on ARC-AGI-1, Near-Zero Cost

1h ago·0:00 listen·Source: ScienceBlog.com

Summary

A new model called BDH-CQ achieved a 29.5 percent success rate on the ARC-AGI-1 evaluation set. What's interesting is this model, with 150 million parameters, performed this at an estimated cost of about $0.0007 per puzzle. This highlights a focus on cost-efficiency rather than just accuracy. For comparison, another model, OpenAI’s GPT-5.6 Luna, scored 34.2 percent on the same benchmark but at a higher cost. BDH-CQ is estimated to be about 11 times cheaper. The ARC-AGI-1 benchmark tests rules through colored grids, requiring systems to infer transformations from input-output examples. BDH-CQ solved 118 out of 400 public evaluation tasks. Here's the thing: BDH-CQ uses a different approach. Instead of generating long chains of intermediate tokens, it reasons iteratively within a continuous latent state. This means its intermediate work is not decoded into language. The bottom line is this approach suggests a potentially more cost-effective way for AI models to tackle complex reasoning tasks.

Read the full article on ScienceBlog.com

This is an AI-generated audio summary. Always check the original source for complete reporting.

Share
Keep Listening