AI Struggles to Fix Vulnerabilities Without Human Oversight
AI models struggle with security flaw remediation, achieving only a 26% success rate. Human oversight remains crucial for effective patching.
Research indicates that AI models may not be effective in addressing security flaws. An analysis by 1Password's Off-by-1 Labs found that patches generated by models like ChatGPT 5.5 and Claude Opus 4.8 succeeded only 26% of the time, with many either failing to fully remediate issues or introducing new problems.