Chinese AI Outperforms Claude Code in Autonomous Research

Jul 21·0:00 listen·Source: South China Morning Post

Summary

A Chinese AI agent has outperformed Anthropic’s Claude Code in autonomous research. The Qiushi Engine, from Zhejiang University, topped the ResearchClawBench leaderboard. This benchmark tests AI agents' ability to independently conduct research. It compares their findings against human-written papers to see if they can match or exceed original conclusions. The Shanghai Artificial Intelligence Laboratory created this benchmark. The Qiushi Engine is a large language model-based agent. It's designed for scientific research in real physical environments. Its developers state it can perform "end-to-end autonomous scientific discovery," unlike other systems that might be task-limited. What's interesting is that while it performs well, this AI can't yet reliably make new discoveries. This shows the ongoing progress and current limitations in AI's research capabilities.

Read the full article on South China Morning Post

This is an AI-generated audio summary. Always check the original source for complete reporting.

Share
Keep Listening