How to Measure Percentage of Code Written by AI Tools

How to Measure AI Code Contribution in Engineering Teams

Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: September 1, 2026

Key Takeaways

  • AI code contribution percentage measures the share of committed code generated or assisted by AI tools, distinct from adoption or acceptance rates.
  • Accurate measurement is difficult because metadata tools cannot distinguish AI-generated code, and heuristic detectors reach only 20–25% accuracy.
  • Three primary methods exist: repository diff analysis, IDE telemetry, and Git metadata, with client-level capture (origin telemetry) delivering the highest accuracy.
  • Always pair the percentage with quality and delivery metrics to avoid vanity metrics and to surface real ROI.
  • Exceeds AI provides line-level, tool-agnostic provenance across every AI coding assistant your team uses. Connect your repo and start your free pilot today.

Why Measuring AI Code Contribution Is Hard

Most tools in the market cannot answer “what percentage of our code is AI-generated?” with confidence. Metadata-only platforms such as Jellyfish, LinearB, and Swarmia track PR cycle times, commit volumes, and review latency. LinearB’s 2026 Software Engineering Benchmarks Report, based on more than 8.1 million pull requests from 4,800 teams, found that 44.7% of organizations do not formally measure AI’s impact, yet 76.1% of leaders report productivity gains based on adoption signals rather than delivery data. These tools can show that PR throughput increased after you rolled out Cursor. They cannot show whether that increase came from AI or from engineers simply working faster.

AI detection tools fall into two main camps. The first relies on heuristics and watermarks, looking for patterns like large volumes of code written in a short window or markers left in AI tool output. By Exceeds AI’s assessment, these signals reach only about 20–25% accuracy. They are fallible, gameable, and cannot explain how the engineer actually worked. Code detection is fundamentally harder than prose detection because idiomatic solutions converge, boilerplate is identical regardless of author, and formatters erase most stylistic signals that stylometry relies on.

The second camp uses client-level capture and observes what happens on the engineer’s machine at the moment the work occurs. Origin-capture telemetry observes rather than infers, which makes it the accurate ongoing measurement method, while diff-churn analysis remains inference and cannot attribute code to a person. Only a small number of platforms support this capability. Exceeds AI, powered by Exceeds Ink, is one of them.

Actionable insights to improve AI impact in a team.
Actionable insights to improve AI impact in a team.

Three Ways to Measure AI Code Contribution (and Their Trade-offs)

Teams use three primary methods to measure AI code contribution percentage. Each method carries specific accuracy characteristics and operational trade-offs.

1. Repository Diff Analysis (High Accuracy)

Repository diff analysis connects directly to your repositories and analyzes actual code diffs at the commit and PR level. It examines structural patterns, formatting habits, and syntax anomalies to distinguish AI-generated code from human-written code across tools such as Cursor, GitHub Copilot, Claude Code, and others.

Accuracy is high relative to the alternatives. This method offers the closest available approximation to a gold standard for post-hoc measurement. The limitation is that the analysis is inferential. It reads patterns in finished code rather than observing the authorship event directly. You can start with basic Git metadata to identify explicitly tagged commits:

git log --all --grep="Co-Authored-By: .*AI" --pretty=format:"%h %an %s"

This command captures only a small subset of actual AI usage. True repo diff analysis goes deeper and examines the code itself. AI-detection tools are trained on stylistic quirks that both sides can defeat, produce false positives on clean human code, and cannot see anything that happened before the file reached its current state.

2. IDE Telemetry (Medium-High Accuracy)

IDE extensions track keystrokes, completion acceptance rates, and active window time directly from developer tools like GitHub Copilot or Cursor. Accuracy is medium-high. A critical blind spot remains when developers copy-paste code from web browsers or use AI tools outside the instrumented IDE.

The most reliable approach for measuring AI code share combines editor telemetry with commit-level metadata for cross-validation, but reported AI code share percentages still vary by team and domain. If your team uses Cursor, Claude Code, and Codex, you receive fragmented data from each vendor’s telemetry with no aggregate view.

3. Git Metadata (Low-Medium Accuracy)

Git metadata analysis filters commit logs for specific hooks, tags, or patterns such as “Co-Authored-By” trailers. Accuracy is low to medium because it depends on developers manually and consistently appending tags. AI-Assisted Commits % depends on tooling configuration and developer compliance. Without enforcement, developers may forget to tag AI-assisted commits, which results in undercounting.

Client-level capture provides the most authoritative approach because it observes the AI tool’s output at the moment it is generated. Exceeds Ink captures AI authorship across Cursor, Claude Code, Codex, GitHub Copilot, and Windsurf with line-level fidelity and writes a portable attestation as a Git Note at refs/notes/exceeds-ink. This process is deterministic and synchronous. It uses only short-lived hook processes and avoids any PATH-shimmed git binary or global git config mutation.

The Balanced Scorecard for AI Code Contribution

AI code contribution percentage becomes a vanity metric when you remove quality and delivery context. A team with 70% AI-assisted lines and high code turnover above 7% has a quality problem masked by a volume metric, while a team with 50% AI-assisted lines and low code turnover below 4% is performing well.

Exceeds AI Impact Report shows AI code contributions, productivity lift, and AI code quality
Exceeds AI Impact Report shows AI code contributions, productivity lift, and AI code quality

GetDX’s research across 500+ engineering organizations found that while AI-authored code share climbed from 24% to 52.7% over three quarters and pull request size nearly doubled, code maintainability improved only 3.8% and developers’ change confidence fell 6.1%. Volume and value diverged in those environments.

A balanced scorecard uses four layers:

  • Adoption: Percentage of developers using AI tools weekly. Industry average is 30–40% WAU, and top-quartile organizations reach 60–70%.
  • Contribution: The AI code contribution percentage itself. Industry average sits at 15–25% of committed lines, and top-quartile teams reach 40–60%.
  • Quality: Rework rate, defect density, and incident rate for AI-touched code versus human code. GitClear’s analysis of 211 million lines found code churn rose from 5.5% to 7.9% and duplicated blocks increased eightfold during 2024, with cloned blocks linked to 15–50% more defects.
  • Delivery: Cycle time, throughput, and PR review efficiency for AI-assisted versus human-authored work.

Always report the AI code contribution percentage alongside at least one quality metric. A team at 25% AI contribution that outperforms on quality and delivery tells a stronger story than a team at 70% with rising churn.

Four-Step Framework to Implement AI Contribution Measurement

This four-step playbook shows how to measure AI code contribution percentage in your organization.

Step 1: Establish a Baseline

Best practice is to collect 60–90 days of pre-AI telemetry. Baseline data should include delivery metrics such as lead time for changes, deployment frequency, change failure rate, PR cycle time, review turnaround time, rework rate, and escaped defect rate. If AI is already rolled out, analyze historical repository data to establish a comparable starting point. Freeze the baseline by team and by work type so later comparisons remain fair.

Step 2: Implement Tool-Agnostic Intelligence

Relying on a single vendor’s dashboard leaves gaps. Your team likely uses multiple AI tools, such as Cursor for feature work, Claude Code for refactoring, and Codex for batch tasks. GetDX’s analysis confirmed that developers without enterprise AI tool telemetry still report meaningful time savings and AI-authored code, which shows that shadow AI is widespread and leads to undercounting when you track only enterprise telemetry. That is why Exceeds AI captures AI contributions across all tools with line-level fidelity through Exceeds Ink’s per-tool checkpoint materializers.

Step 3: Conduct Same-Engineer Analysis

Track each developer’s performance against their own historical baseline rather than against other engineers. This approach removes confounding variables like skill level and task complexity. Using same-engineer methodology, one major financial services company found a 30% increase in pull request throughput year-over-year among AI tool users, compared to 5% among non-adopters.

Step 4: Set Context-Specific Targets

Set targets that reflect the type of work and its risk level. Aim for 50–70% AI composition for boilerplate and test files. Use a tighter 10–25% range for critical business logic or security-sensitive modules. AI code share varies by role: frontend developers see 40–50% AI contribution, while security engineers see 10–15%; boilerplate and tests are 50–70% AI-written, while critical infrastructure is 5–15%. Derive targets from your own baseline first, then adjust by team, work type, and risk profile. Start your free pilot to see your AI contribution baseline.

Benchmarks and Targets for AI Code Share

Benchmarks help you calibrate the targets you set in the previous step. The 2025 Stack Overflow Developer Survey found that 84% of developers use or plan to use AI tools, with 51% of professional developers using them daily. The global average for AI code contribution percentage sits at roughly 27–42% in active environments depending on measurement methodology. The optimal range for balanced adoption is 20–40%. Beyond that threshold, teams face higher risk of technical debt accumulation.

Exceeds AI Repo Leaderboard shows top contributing engineers with trends for AI lift and quality
Exceeds AI Repo Leaderboard shows top contributing engineers with trends for AI lift and quality

Benchmarks by team performance tier:

  • Bottom quartile: Less than 10% AI-assisted lines
  • Industry average: 15–25% AI-assisted lines (as noted earlier)
  • Top quartile: 40–60% AI-assisted lines
  • Elite teams: Above 75% AI-assisted lines

What is the 10/20/70 rule in AI? The 10/20/70 rule describes a governance framework for AI adoption. Ten percent of effort goes to model development, 20% to infrastructure and tooling, and 70% to organizational change, process redesign, and adoption. The rule highlights that AI code contribution percentage is only one piece of the puzzle. Most effort should focus on making AI work effectively for your teams.

What percentage of code should be AI-generated? Context determines the right percentage. AI-generated code share in high-adoption organizations is 30–70%, reflecting production data from teams with mature AI coding workflows; teams working on greenfield features tend toward the higher end, while teams maintaining complex legacy systems trend lower. Set targets that reflect your team’s actual work type and risk profile.

Common Measurement Pitfalls with AI Code Contribution

Vanity metrics. AI code contribution percentage without quality context lacks meaning. Common measurement mistakes include relying on vanity metrics such as lines of code, measuring only speed, skipping a baseline, drawing conclusions too early, and focusing on individuals instead of team or system outcomes. Always pair the percentage with rework rate, defect density, and review depth.

Gaming lines of code. AI tools can generate massive volumes of code that appear productive but create technical debt. GitClear’s January 2026 report found that regular AI users averaged 9.4x higher code churn than their non-AI counterparts, more than double the productivity gains the tools provided. Incentives that reward lines of code encourage volume rather than outcomes.

Surveillance concerns. Measurement should support coaching and enablement. According to the 2025 JetBrains Developer Ecosystem survey, 66% of developers do not believe that current metrics reflect their contributions. Engineers who feel monitored will game the metrics or resist adoption. Exceeds AI builds trust by giving engineers personal insights and AI-powered coaching through ink-prompting-coach, which provides value directly in their workflow.

Drawing conclusions too early. AI adoption takes 4–8 weeks to stabilize, and organizations are advised to allow at least 3 to 6 months before drawing strong conclusions about AI impact. Wait for at least one full quarter of data before judging the AI code contribution percentage.

Tools Comparison: Exceeds AI vs. Metadata-Only Platforms

Most developer analytics platforms such as Jellyfish, LinearB, and Swarmia were built for the pre-AI era. They track metadata like PR cycle times, commit volumes, and review latency. As noted earlier, they track metadata and cannot distinguish AI from human code, which makes them unable to measure the AI code contribution percentage with confidence. A 2026 study by Multitudes of more than 700 engineering professionals found that 75% of participants struggle to measure AI’s impact, and 60% cited a lack of clear metrics to evaluate it as their most common challenge.

Exceeds AI Impact Report with Exceeds Assistant providing custom insights
Exceeds AI Impact Report with PR and commit-level insights

GetDX and Git AI are the only other players with real client-level capture technology. Both create operational problems that Exceeds Ink was deliberately designed to avoid.

Git AI ships a long-lived per-user daemon and a PATH-shimmed git binary. On Windows, git.exe is literally a copy of git-ai.exe. It also destructively overwrites global git config. Your CISO cannot audit it easily, and your fleet operations team gains another always-on process to manage. Git AI’s async daemon reconciliation also creates a race window, because a fast push can land before attribution completes.

GetDX ships a closed-source CLI daemon that runs continuously and transmits aggregates to DX Data Cloud. All attribution lives in their cloud, and nothing portable lives in your repo. Security teams must trust the mechanism without local evidence.

Exceeds AI, powered by Exceeds Ink, follows a different architecture:

  • Line-level provenance via Git Notes at refs/notes/exceeds-ink, which is portable, auditable, and stored in your own repo
  • Per-tool checkpoint materializers for Claude Code, Cursor, and Codex that resolve edit evidence at commit finalization and behave deterministically
  • Short-lived hook processes only, which avoids any daemon, PATH shim, or global git config mutation
  • HMAC-SHA256-signed remote ingest with revocable per-machine tokens
  • LLM-based prompt redaction before persistence
  • In-agent coaching via ink-prompting-coach, installed directly into Claude Code or Cursor

Setup takes hours instead of months. First insights arrive within 60 minutes. Board-ready ROI reports follow within weeks. Connect your repo and get your first insights in 60 minutes.

Conclusion: Turning AI Contribution Data into Outcomes

Measuring the AI code contribution percentage is the first step in a longer journey. The real goal is using that data to improve outcomes such as higher quality, faster delivery, and smarter adoption decisions. Metadata-only platforms and heuristic-based detectors cannot measure the AI code contribution percentage accurately. They will leave you guessing. Exceeds AI, powered by Exceeds Ink, provides authoritative, line-level provenance across every AI tool your team uses and pairs it with actionable coaching that turns measurement into improvement.


Frequently Asked Questions

What is the difference between AI code contribution percentage and AI adoption rate?

AI adoption rate measures how many developers on your team use AI tools and how frequently they use them. AI code contribution percentage measures how much of the code your team actually ships originated from those tools. A team can have 90% adoption and 10% code contribution, which means most engineers have licenses but use them only for minor autocomplete. Another team can have 30% adoption and 50% code contribution, which means a smaller group of power users generates the majority of committed code with AI assistance. Adoption rate shows who uses AI. Contribution percentage shows what AI produces. Both matter, and conflating them is one of the most common measurement mistakes engineering leaders make. Contribution percentage also provides the denominator that makes other metrics like velocity and quality interpretable. Without it, a 40% increase in PR throughput could mean almost anything.

How does Exceeds AI differ from GitHub Copilot’s built-in analytics?

GitHub Copilot Analytics shows usage statistics such as acceptance rates and lines suggested. It does not prove business outcomes. It does not show whether Copilot-touched code has higher or lower defect rates than human-written code, how Copilot-assisted PRs perform over time compared to unassisted PRs, which engineers use Copilot effectively versus struggling with it, or what happens to that code 30, 60, or 90 days after it merges. Copilot Analytics is also blind to every other AI tool. If your team uses Cursor, Claude Code, Windsurf, or Codex alongside Copilot, those contributions remain invisible. Exceeds AI provides tool-agnostic AI detection and outcome tracking across your entire AI toolchain, with line-level attribution that lives in your own repository as a portable Git Note rather than inside a vendor’s cloud dashboard.

Why does accurate AI code measurement require repo access?

Metadata cannot distinguish AI-generated code from human-written code. Without repo access, a tool can only see that PR #1523 merged in four hours with 847 lines changed and two review iterations. With repo access and Exceeds Ink’s line-level provenance, you can see that 623 of those 847 lines were AI-generated by Cursor, that those AI lines required one additional review iteration compared to human lines, that the AI-touched module had higher test coverage, and that 30 days later the AI-touched code had zero production incidents. That difference separates a metadata dashboard from code-level truth. Repo access provides the only way to prove and improve AI ROI at the code level. Exceeds AI is designed to pass enterprise security review, with minimal code exposure, no permanent source code storage, encryption at rest and in transit, SSO/SAML support, and an in-SCM deployment option for the highest-security requirements.

What is a realistic timeline for seeing meaningful AI code contribution data?

With Exceeds AI, first insights appear within 60 minutes of connecting your repo, and complete historical analysis finishes within four hours. Drawing meaningful conclusions about AI impact requires more time than the data ingestion step. The first four to eight weeks after AI rollout usually form an adoption and workflow-adjustment period. Engineers learn new tools and habits, and the metrics reflect that instability. Quality signals lag speed signals by eight to twelve weeks. The recommendation is to read leading indicators like utilization and cycle time weekly but to draw conclusions quarterly. Wait for at least one full quarter of data before judging the AI code contribution percentage, and compare each team against its own pre-AI baseline using same-engineer methodology rather than comparing teams against each other.

How should engineering leaders present AI code contribution data to executives and boards?

AI code contribution percentage alone does not qualify as a board-ready metric. It serves as a starting point. Executives and boards need to see the percentage in context. They want to know what share of committed code is AI-generated, what the quality of that code looks like relative to human-written code (rework rate, defect density, incident rate), and what the delivery impact has been (cycle time, throughput, PR review efficiency). The most compelling board presentation pairs a contribution percentage with a quality signal and a delivery outcome. For example, AI-assisted code now represents 38% of committed lines, AI-touched PRs have a rework rate 12% lower than the pre-AI baseline, and cycle time has improved 18% for the teams with the highest AI adoption. That narrative of contribution, quality, and delivery turns a vanity metric into proof of ROI. Exceeds AI generates board-ready reports that connect these three layers, anchored by Exceeds Ink’s per-commit, per-tool attestation.

Discover more from Exceeds AI Blog

Subscribe now to keep reading and get access to the full archive.

Continue reading