Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: July 13, 2026
Key Takeaways
- Published studies on GitHub Copilot show task-completion gains ranging from 55% faster on controlled tasks to 19% slower on real-world complex work. Results depend heavily on methodology and context.
- High AI adoption increases PR volume but also creates a clear review tax, with review times rising up to 91% and no corresponding improvement in DORA metrics.
- AI-generated code introduces persistent technical debt, with 22.7% of issues surviving review and complexity increases that reduce future velocity after initial gains fade.
- Self-reported productivity claims often diverge from objective outcomes, as developers perceive speed improvements even when actual task completion slows.
- Exceeds AI provides the only provenance layer that delivers line-level attribution, longitudinal outcome tracking, and multi-tool visibility. Start your free pilot today to measure real AI impact in your organization.
What Copilot Productivity Studies Really Show
Peng et al. (2023) remains the most-cited figure in AI coding productivity research, finding that developers completed a controlled HTTP server task 55.8% faster on average with GitHub Copilot. That number has anchored vendor marketing ever since. A more recent controlled experiment from METR (Becker et al., 2025) recruited 16 experienced open-source maintainers working on real repositories averaging more than 1 million lines of code and found the opposite pattern. Developers using Cursor Pro with Claude 3.5 Sonnet completed tasks more slowly than they expected while believing they were 20% faster.
This difference in findings reflects methodology, not a contradiction. The Peng et al. study used a bounded, well-specified task on a clean codebase. The METR study used real-world complexity. Neither study tracked what happened to that code 30, 60, or 90 days later.
Adoption data shows that developers are committing to AI tools regardless of the measurement debate. The 2025 Stack Overflow Developer Survey found that 51% of professional developers use AI tools daily. In the same survey, few reported “great productivity gains,” and 66% reported spending more time fixing “almost-right” AI code.
The core caveat is clear. Every headline velocity figure is a task-level measurement on a controlled input. None of the widely cited studies attribute outcomes at the commit or PR level across a real multi-tool engineering organization over a sustained period.
Task Completion Time Across 2025–2026 Studies
Three studies from 2025–2026 represent the strongest available evidence on task completion time, and each carries a distinct methodology caveat.
Peng et al. (arXiv 2023, widely cited through 2025): As noted earlier, this study found 55.8% faster completion on a single, well-scoped HTTP server task. The task was isolated, the codebase was clean, and no downstream quality outcomes were tracked, so the result does not generalize to production engineering work.
METR (Becker et al., July 2025): Experienced developers ran slower than expected on real-world issues from large open-source repositories. A February 2026 METR post stated that its new experiment data was unreliable due to bias and did not report any revised productivity estimate. The perception gap persisted in both rounds, as developers still estimated they were faster despite the data.
DX Q1 2026: A longitudinal analysis of more than 400 engineering organizations found that a 65% average increase in AI tool usage produced a median PR throughput increase of just under 8%, with most organizations landing in the 5–15% range. The same-engineer longitudinal methodology, which tracks individuals against their own pre-AI baseline, is more rigorous than cross-sectional comparisons. PR throughput still remains a volume metric, not a quality or outcome metric. Like the Peng and METR studies, DX’s findings confirm that measurement scope determines the result.

The pattern across all three studies is consistent. Task-level gains appear on bounded work, while organizational-level gains remain modest and depend heavily on what happens after the code is written.
Connect my repo and start my free pilot to measure task completion and PR outcomes at the commit level across every AI tool your team uses.
PR Throughput, Review Speed, and the Review Tax
PR volume is the metric where AI’s impact is most visible, yet it becomes misleading when used as a standalone signal.
Faros AI telemetry analysis of more than 10,000 developers across 1,255 teams found that high-AI-adoption teams completed 21% more tasks and merged 98% more PRs. PR sizes grew 51% and review times increased 91% or 441% in high-AI-adoption teams. DORA metrics remained flat. This pattern shows a clear chain: higher volume produced larger PRs, which demanded longer reviews, while system-level delivery speed did not improve.

Swarmia’s analysis extended this picture by examining batch size over time. Median batch size increased between Q1 2025 and Q1 2026 as agent adoption became mainstream. Swarmia concluded that enabling engineers to code faster produces only modest increases in organizational output when review and deployment processes remain unchanged.
DX Q1 2026 data showed daily AI users merged 2.3 PRs per week (median) versus 1.4 PRs per week for non-users, which represents a roughly 60% higher PR rate. The same dataset showed that some organizations experienced a nearly 2 percentage-point rise in Change Failure Rate after AI adoption, equivalent to shipping up to 50% more defects.
The higher PR volume paired with longer reviews in high-adoption cohorts provides the clearest quantification of the review tax. Vendor dashboards that report only acceptance rates or lines suggested do not surface this effect.
Connect my repo and start my free pilot to see PR throughput and review cycle data broken down by AI-touched versus human-only contributions.
Technical Debt Signals Beyond the Merge Button
Most studies that engineering leaders cite stop at merge, while the studies that matter for long-term AI ROI begin at that point.
He et al. (Carnegie Mellon University, MSR 2026) analyzed 806 GitHub repositories adopting Cursor using a difference-in-differences quasi-experimental design with propensity-score-matched controls. Cursor adoption produced a statistically significant 3–5x increase in lines added during month 1 that dissipated after two months. Increases in code complexity and static analysis warnings did not dissipate. Panel GMM models showed that accumulated complexity subsequently reduced future development velocity, creating a self-reinforcing cycle.
Liu et al. (arXiv 2026, Singapore Management University and Huazhong University of Science and Technology) analyzed 302,579 verified AI-authored commits across 6,299 GitHub repositories and five coding assistants. More than 15% of AI-authored commits from every assistant introduced at least one detectable issue. Code smells accounted for 89.3% of all 484,366 identified issues. Critically, 22.7% of those issues survived at the latest repository revision. AI-generated technical debt is settling permanently into production architectures instead of being caught in review.
GitClear’s analysis of 211 million changed lines across Google, Microsoft, Meta, and enterprise repositories found that between 2021 and 2024, the share of moved lines, a proxy for refactoring activity, collapsed from 25% to less than 10% (or 9.5%). Copy-pasted lines rose from 8.3% to 12.3%, marking the first time in software development history that cloned lines exceeded refactored lines.
SonarSource’s 2026 State of Code Developer Survey found that 53% of developers attributed negative technical debt impact to AI-generated code. The survey also reported that 61% agreed AI often produces code that looks correct but is not reliable, which creates hidden bugs and encourages teams to skip thorough review. The most frequent AI users experience a toil shift into managing technical debt and correcting AI-generated code.
None of these signals appear in metadata-only tools. Capturing them requires tracking AI-attributed code from commit through production outcomes over more than 30 days.
Connect my repo and start my free pilot to track longitudinal outcomes on AI-touched code before technical debt compounds.
How to Validate AI Productivity Claims in Your Org
The studies above show that velocity claims from AI coding tools function as testable hypotheses, not settled facts. Verifying them in your organization requires three specific capabilities that no metadata-only tool provides.

1. Line-level AI versus human attribution. PR cycle time dropping 20% means something very different if 15% of those PRs were AI-assisted versus 60%. Without line-level attribution, you cannot separate AI’s contribution from tenure effects, seasonality, or team composition changes. Heuristic and watermark-based detection tops out around 20–25% accuracy. The only authoritative method uses client-level capture at the moment the work is done, observing what the engineer typed into the AI tool, how long they iterated, and which lines the tool produced.
2. Longitudinal outcome tracking over more than 30 days. The He et al. and Liu et al. findings both show that AI-generated issues survive review and accumulate in production. A measurement framework that stops at merge cannot detect this pattern. Effective tracking connects the AI-attributed commit to downstream signals such as incident rates, follow-on edits, rework rates, and test coverage changes over a minimum 30-day window.
3. Multi-tool visibility. Engineering teams in 2026 rarely use a single AI coding tool. The research literature covers GitHub Copilot, Cursor, Claude Code, Codex, and Windsurf as distinct tools with distinct acceptance rate profiles, complexity footprints, and security issue rates. A single-vendor dashboard remains structurally blind to the majority of AI activity on most teams.
Exceeds AI is the only provenance layer that writes portable, auditable Git Notes at commit time. It captures which tool, which model, which interaction mode, and which lines across Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf. Those Git Notes live in your own repository, travel across forks and mirrors, and feed the Exceeds AI platform’s longitudinal outcome tracking and multi-tool comparison without requiring always-on daemons or global git configuration changes.
A practical verification checklist for engineering leaders:
- Establish a pre-AI baseline for PR throughput, cycle time, and defect density using same-engineer longitudinal methodology, not cross-sectional comparisons.
- Deploy line-level AI attribution across every tool in use, not just the tool your vendor dashboard covers.
- Track AI-attributed code for 30, 60, and 90 days post-merge, including incident rates, rework patterns, static analysis warning trends, and test coverage changes.
- Separate volume metrics such as PRs merged and lines committed from outcome metrics such as defect density, change failure rate, and review cycle time per AI-touched PR.
- Identify interaction-mode patterns, including plan, ask, agent, edit, and headless, to distinguish coachable adoption patterns from structural quality risks.
Frequently Asked Questions
What is the difference between self-reported and code-level velocity measurement?
Self-reported measurement relies on developer surveys that ask how much time AI saved or how productive developers feel. The METR 2025 study illustrates the problem directly. Developers predicted they would be 24% faster, reported feeling 20% faster during the study, and were actually 19% slower on objective task completion time. Self-reported data captures perception, not outcome.
Code-level measurement attributes specific lines, commits, and PRs to AI tools and then tracks what happens to that code over time, including cycle time, defect density, rework rates, and incident rates. It requires repo access and a provenance layer that captures AI authorship at the moment of creation, not inferred after the fact from metadata patterns. This difference matters for board reporting because a perception-based productivity claim cannot withstand a CFO’s follow-up question about defect rates or rework costs.
How long do AI-generated code quality issues typically take to surface?
Research points to two distinct timelines. Immediate issues such as logic errors, security vulnerabilities, and readability problems are detectable at review but frequently missed because reviewer bandwidth is overwhelmed by higher PR volume. The Liu et al. 2026 study, discussed earlier, found that a significant share of AI-introduced issues survived to the latest repository revision, which shows that many problems are not caught even over the full project history.
Structural issues such as complexity accumulation, architectural misalignment, and duplicate code blocks surface over 30 to 90 days as teams attempt to build on AI-generated foundations. The He et al. CMU study found that velocity gains from Cursor adoption dissipated within two months while complexity increases persisted through month 6 and beyond. Exceeds AI’s longitudinal outcome tracking monitors AI-attributed code over more than 30 days specifically to catch this second category before it becomes a production crisis.
Can metadata-only tools prove Copilot ROI?
No. Metadata-only tools, including traditional developer analytics platforms that track PR cycle time, commit volume, and review latency, cannot distinguish AI-generated lines from human-authored lines. These tools can show that PR throughput increased after Copilot deployment, but they cannot prove causation, cannot identify whether the increase came from AI assistance or from other factors, and cannot connect the throughput change to downstream quality outcomes. The Faros AI finding discussed earlier, where dramatic PR volume increases coincided with longer reviews yet flat DORA metrics, remains invisible to a metadata dashboard that reports only merge volume. Proving ROI requires code-level attribution connected to outcome tracking, which requires repo access and a provenance layer.
Which 2025–2026 studies used commit-level attribution?
Liu et al. (arXiv 2026) built a dataset of 302,579 verified AI-authored commits by mining explicit Git metadata signals such as actor logins, author emails, author names, and Co-authored-by trailers. They then ran commit-level differential static analysis on the parent and post-commit revisions to attribute introduced issues directly to individual AI-authored commits. He et al. (CMU, MSR 2026) used a difference-in-differences design on 806 repositories identified via .cursorrules files, tracking static analysis warnings and complexity metrics at the commit level over 20 months. DX’s Q1 2026 longitudinal analysis combined Git analytics with AI tool telemetry at the commit and PR level to track same-engineer performance over time. None of these studies had access to line-level interaction-mode data, which would distinguish whether a developer used agent mode without a plan phase versus a structured ask-then-verify workflow. That signal is only available through client-level capture at commit time.
Conclusion: Why Code-Level Proof Now Matters
Published evidence on GitHub Copilot’s impact on developer velocity spans a 74-percentage-point range, from a 19% slowdown for experienced developers on complex tasks to a 55% speedup on bounded, well-specified work. That range reflects genuine variation in task type, codebase complexity, team maturity, and measurement methodology. Studies that show the largest gains use the narrowest measurement windows. Studies that track outcomes over more than 30 days consistently find that velocity gains are transient and quality costs are persistent.
Engineering leaders evaluating AI spend cannot rely on vendor dashboards that report acceptance rates or on metadata tools that report PR volume without attribution. A verification framework that produces defensible board-level answers requires line-level AI authorship attribution across every tool in use, longitudinal outcome tracking over a minimum of 30 days, and multi-tool visibility that covers the full reality of how engineers work in 2026.
Exceeds AI is the only platform that delivers all three capabilities. Exceeds Ink writes a portable, auditable, line-level attestation alongside every commit, capturing tool, model, interaction mode, and session across Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf. The Exceeds platform connects that provenance to outcome analytics, longitudinal debt tracking, and prescriptive coaching surfaces so engineering leaders can answer the board with evidence rather than estimates.
Connect my repo and start my free pilot and get your first commit-level AI attribution insights within the hour.