AI-Generated Patches: Over Half Are Broken, New Bugs Emerge
Summary
More than half of AI-generated patches are broken. This comes as concerns grow about an expanding attack surface for malicious hackers. New research tested the patching capabilities of two popular commercial models: OpenAI’s ChatGPT 5.5 and Anthropic’s Claude Opus 4.8. They found that generative AI is more likely to create an exploitable patch or introduce entirely new bugs than to fix a vulnerability. The overall success rate for fully patching a vulnerability without new problems was less than 47%. Researchers found that models often addressed only some vulnerable code paths or added fragile code. They sometimes even introduced subtle changes while patching. Another report found that while large language models have made strides in crafting workable code, security is a different story. The average security "pass rate" for AI-generated code is around 56%. In 44% of tests, models introduced a detectable vulnerability. This suggests that largely autonomous vulnerability discovery and patching may not yet effectively fix the explosion of vulnerable code in the AI era.
This is an AI-generated audio summary. Always check the original source for complete reporting.