Researchers followed 6,774 merged AI coding agent pull requests authored by OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code in the AIDev-pop dataset of open-source repositories with at least 500 stars, comparing them with 5,044 contemporaneous human pull requests from the same repositories. They report that merged agent PRs had roughly 1.62 times the odds of requiring a verified follow-up fix, that agents authored most of the fixing work themselves, and that a merged agent PR should not be assumed to be finished work.
Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests
Researchers followed 6,774 merged AI coding agent pull requests authored by OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code in the AIDev-pop dataset of open-source repositories with at least 500 stars, comparing them with 5,044 contemporaneous human pull requests from the same repositories. They report that merged agent PRs had roughly 1.62 times the odds of requiring a verified follow-up fix, that agents authored most of the fixing work themselves, and that a merged agent PR should not be assumed to be finished work.
Source: Arxiv