Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: July 6, 2026
Key Takeaways
- Engineering teams running multiple AI coding tools need commit-level attribution to separate AI-generated code from human work and avoid hidden technical debt.
- Pre-AI baselines across PR throughput, rework rate, change failure rate, and maintainability make post-AI ROI claims credible.
- Tracking 30-, 60-, and 90-day outcomes shows whether AI code creates persistent issues that snapshot metrics never reveal.
- Override and rework patterns by tool and interaction mode give the clearest signals for targeted coaching and process changes.
- Exceeds AI provides line-level provenance across major AI coding tools plus direct coaching actions—start your free pilot today.
Before You Begin: Four Foundations for Reliable AI Measurement
Four prerequisites determine whether your measurement program produces actionable signal or noise.
Repo access. Metadata-only tools see PR cycle time and commit volume. They cannot see which 623 of 847 changed lines were AI-generated, which tool produced them, or what happened to those lines 45 days later. Read-only repo access unlocks code-level attribution and makes those questions answerable.
Baseline data. Pre-deployment baselines using actuals, not estimates, for PR throughput, change fail percentage, and time allocation across SDLC phases are required to substantiate post-deployment ROI claims. Capture at least 90 days of pre-AI history before drawing comparisons.
Stakeholder alignment. Agree in advance on which metrics connect to business outcomes. Metrics for evaluating AI in engineering must be owned at the leadership level and connect directly to business outcomes, or measurement becomes noise.
Scope and time expectations. Attributable impact on throughput and quality requires six months after workflow integration and proficiency ramp; measuring ROI at 30 days captures adoption friction rather than return. Plan for 60-to-90-day checkpoints as early signal, and a six-month review as the primary evaluation gate.
Step-by-Step Tutorial: Building a Complete AI Measurement System
With these foundations in place, your team can build a measurement program that produces reliable, actionable signal. The seven steps below form a complete system, and each step builds on the one before it.
Step 1: Instrument Every AI Coding Tool in Your Stack
Purpose: Most engineering teams in 2026 run multiple AI coding tools simultaneously. AI tool usage has increased while median PR throughput gains have been more modest, which creates a gap that you cannot diagnose without knowing which tool produced which code.
What to configure: Deploy a provenance layer that captures AI authorship at the line level across every tool in use. Exceeds Ink achieves this through a lightweight binary that integrates into your existing Git workflow. It fires from standard Git hooks at commit finalization, so engineers keep their normal process. At that moment, it writes a structured attestation as a Git Note alongside every commit, which creates a permanent, portable record of AI involvement. This architecture enables coverage across Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf with dedicated per-tool checkpoint materializers, plus lighter-weight detection across up to approximately 50 additional tools.
Success indicator: Every commit in your target repositories carries a machine-readable attestation identifying AI-generated lines by tool, model, session, and interaction mode. Lines that cannot be confidently attributed are recorded as unknown rather than silently assigned to human or AI.
Common mistake: Teams often rely on heuristics or watermark scanning instead of client-level capture. By Exceeds’ own assessment, heuristic and watermark-based detection tops out around 20–25% accuracy. Authoritative measurement requires observing what happens on the engineer’s machine at the moment the work is done.
Step 2: Establish Pre-AI Baselines Across Five Core Indicators
Purpose: Pre-AI baselines turn raw post-adoption numbers into meaningful comparisons. A two-year longitudinal study tracking 39 developers across 703 repositories found that GitHub Copilot users were already more productive before adoption; after controlling for pre-existing differences, there were no statistically significant changes in commit-based activity. Selection bias is real and must be controlled.
What to document: For each team in scope, record PR throughput per engineer per week, rework rate (lines rewritten within 30 days), change failure rate, mean time to restore, and code maintainability scores. Use each engineer as their own control by capturing individual pre-adoption baselines rather than team averages, which obscure high variance.
Success indicator: A baseline report segmented by team, repository, and individual contributor, using at least one full quarter of pre-AI history already captured in your prerequisites.
Pro tip: In months one and two of a measurement program, run developer experience surveys alongside objective metrics to capture perceived cognitive load alongside throughput data. The two signals diverge in ways that matter: surveys of GitHub Copilot users often report being significantly more productive even though objective metrics show more modest gains in sprint velocity and bug rates.
Step 3: Isolate AI vs. Human Contributions at Commit and PR Level
Purpose: Aggregate metrics hide the signal that matters. AI suggestion acceptance rates vary significantly across developers, which makes aggregate productivity metrics misleading.
What to review: Use AI Usage Diff Mapping to identify which specific commits and PRs are AI-touched, down to the line level. Segment by tool (Cursor vs. Claude Code vs. Copilot), by interaction mode (agent vs. plan vs. edit), and by team. This segmentation creates the foundation for the longitudinal tracking in the next step.
Success indicator: A per-PR view showing the exact count of AI-generated versus human-authored lines, attributed to the specific tool and session that produced them.

Watchout: Task type is a dominant factor in PR acceptance rates, producing a 29 percentage-point gap between the highest and lowest task categories, which is larger than typical inter-agent variance. Always segment AI vs. human comparisons by task type (feature, fix, refactor, documentation) before drawing conclusions about tool effectiveness.
Step 4: Track Decision-Quality Indicators Over 30, 60, and 90 Days
Purpose: Decision quality shows up over weeks and months, not just at merge time. AI-generated code that passes review today can fail in production weeks later. 22.7% of AI-introduced issues tracked across 302,579 commits survived to the latest repository revision, which demonstrates persistent technical debt rather than transient problems. Snapshot metrics miss this entirely.
What to monitor: For every AI-attributed commit, track three longitudinal outcomes. First, incident rate at 30, 60, and 90 days post-merge. Second, follow-on edit rate, meaning lines rewritten after initial merge. Third, maintainability score trajectory. Compare these against the same indicators for human-authored commits in the same repositories and time periods.
Success indicator: A cohort view showing AI-touched code versus human-authored code across 30-day, 60-day, and 90-day outcome windows, with clear statistical separation between the two populations.

Troubleshooting: If AI-touched and human-authored code show similar incident rates but AI code shows higher rework rates, prompting patterns likely drive the problem more than tool quality. GitClear’s analysis of 211 million lines found duplicated code blocks rose eightfold in 2024 while refactoring activity dropped to historic lows, which is a pattern that appears in rework data before it appears in incident data.
Step 5: Map Override and Rework Patterns by Team and Tool
Purpose: Override and rework patterns provide a direct observable proxy for AI decision quality. High override rates indicate engineers are correcting AI output before it ships. High rework rates indicate AI output that passed review but required correction afterward.
What to analyze: Segment override rates and rework rates by team, by AI tool, and by interaction mode. Agent-mode sessions without a preceding plan phase represent a coachable pattern, and Exceeds Ink’s interaction-mode classification captures exactly this signal. Compare teams with low rework rates against teams with high rework rates to identify what the high-performing teams are doing differently.
Success indicator: A ranked list of teams by AI decision-quality indicators, with the specific interaction-mode and tool patterns that correlate with better outcomes identified and ready for distribution.

Start surfacing override and rework patterns across your organization with a free pilot today.
Step 6: Quantify AI ROI at Commit and PR Level
Purpose: Engineering leaders need board-ready proof, not sentiment. 94% of engineering leaders say the metrics that matter most are entirely missing from their current measurement frameworks.
What to calculate: Connect AI attribution data to outcome data to produce per-tool, per-team ROI figures. A useful model: for a 200-person engineering organization with 75-plus percent AI usage and an 8% throughput gain, output equals roughly 16 additional engineers; at $200K fully-loaded cost per engineer this yields approximately $3.2M equivalent capacity versus $800K all-in annual AI spend, for 300% ROI. This calculation depends on knowing which throughput gains are attributable to AI, which in turn depends on commit-level attribution rather than metadata.
Success indicator: A per-tool ROI figure that connects AI-attributed commits to productivity outcomes and quality outcomes, with a before-and-after comparison against the pre-AI baseline established in Step 2.

Pro tip: Token spend data paired with outcome data produces an Agentic ROI signal that finance teams can act on directly. Exceeds Ink captures cost and token usage per agent and model, so the dollars and the outcomes appear in the same view.
Step 7: Turn Measurement Into Repeatable Coaching
Purpose: Measurement only matters when it changes behavior at scale. Organizational-complexity metrics including team size and management span are among the strongest predictors of defect-proneness; as spans widen and bandwidth for mentorship and code review shrinks, quality suffers. Manager-to-IC ratios have stretched from a typical 1:5 toward 1:8 or higher, so coaching must travel through systems, not only through managers.
What to distribute: Use Best Practices Insights to identify the top interaction-mode and prompting patterns that correlate with better decision-quality outcomes. Distribute these as versioned skills directly into engineers’ Claude Code or Cursor agents via Exceeds Ink’s ink-prompting-coach. Because skills are versioned and centrally tracked, you can monitor adoption rates in real time and roll back any skill that does not improve outcomes. This rapid iteration cycle makes the system powerful: when one team’s AI PRs show 3x lower rework than another team’s, you can capture that pattern, version it as a skill, and replicate it across the organization within days rather than waiting quarters for organic knowledge transfer.
Success indicator: Measurable improvement in rework rate and override rate for teams that received distributed coaching skills, within two sprint cycles of distribution.
Validation and Success Criteria for Your AI Measurement Program
A measurement program is working when several observable indicators appear together. AI-attributed commits carry complete provenance attestations covering tool, model, session, and interaction mode. Rework rates for AI-touched code trend toward or below rework rates for human-authored code in the same repositories. The 30-day incident rate for AI-attributed commits does not sit materially higher than the baseline established before AI adoption. Override rates decline as coaching skills are distributed. Engineering leaders can produce a per-tool, per-team ROI figure that connects AI usage to throughput and quality outcomes without relying on surveys or estimates.
Before-and-after checks should compare change failure rate, rework rate, and PR throughput against the 90-day pre-AI baseline captured in Step 2. Change failure rate, change confidence, and code maintainability should be treated as core metrics in any AI measurement program because throughput gains without quality tracking can mask up to 50% increases in defects.

Get your first decision-quality scorecard within the hour with a free pilot.
Advanced Considerations for Scale, Governance, and Risk
Scaling across teams. Once the measurement framework is validated on one team, the same provenance layer and coaching skills can roll out across the organization without proportional increases in manager overhead. Skill transfer and rollback capabilities mean that patterns identified in high-performing teams can be versioned, distributed, and tracked centrally.
Governance improvements. Exceeds Ink’s structured JSON attestation in Git Notes feeds directly into policy engines. Organizations can express policies such as requiring additional review on commits where agent mode produced more than a defined percentage of the diff, or blocking deploys when AI authorship exceeds a threshold in sensitive paths. These policies become practical because the attestation is structured and lives in the repository, close to the code and the existing review process.
Cognitive debt monitoring. Margaret-Anne Storey defines cognitive debt as the project-level erosion of shared understanding among developers about what a program does and how it can be changed, which is distinct from technical debt that resides in the code itself. Warning signs include team members hesitating to make changes for fear of unintended consequences and a growing sense that the system is becoming a black box. Longitudinal outcome tracking combined with interaction-mode data surfaces these patterns before they become production crises.
Adjacent topics. As measurement matures, the same commit-level provenance that powers ROI reporting also supports patent examiner inquiries about AI’s role in specific files, legal counsel questions about AI authorship percentages, and auditor requests for machine-readable records of AI involvement in production code.
FAQ
How is measuring AI decision quality different from tracking DORA metrics?
DORA metrics, which include Deployment Frequency, Lead Time for Changes, Change Failure Rate, and Mean Time to Restore, measure delivery system performance at the organizational level. They provide valuable baselines but cannot tell you whether a specific PR’s change failure was caused by AI-generated code, which tool produced it, or what interaction mode the engineer used. Measuring AI decision quality requires commit and PR-level attribution that connects specific lines of code to specific tools, sessions, and outcomes over time. DORA metrics act as an input to that analysis, not a substitute for it.
What if my team uses three or four different AI coding tools?
Multi-tool environments are the norm in 2026, and they represent the scenario where commit-level attribution matters most. Without tool-level provenance, you cannot determine whether a rework spike comes from one tool’s output quality, a specific interaction mode, or a team’s prompting patterns. Exceeds Ink uses per-tool checkpoint materializers for Claude Code, Cursor, and Codex, with adapters covering GitHub Copilot, Windsurf, and lighter-weight detection across approximately 50 additional tools. The result is a single aggregate view of AI impact across your entire toolchain, with tool-by-tool breakdowns available for comparison.
How long does it take before the data is meaningful?
First insights, such as AI attribution on new commits, adoption maps, and tool-level breakdowns, are available within the first hour of connecting a repository. Historical analysis covering up to 12 months of prior commits completes within four hours. However, longitudinal decision-quality indicators require time to accumulate. Thirty-day rework rates need 30 days of post-merge observation, and 90-day incident cohorts need 90 days. The six-month review timeline mentioned in the prerequisites reflects this reality and sets expectations for credible ROI claims.
Will engineers view this as surveillance?
The distinction between surveillance and coaching depends on structure, not slogans. Exceeds AI is built so that engineers receive something valuable in return: personal AI-powered coaching delivered directly into their own Claude Code or Cursor agent, performance review support grounded in their actual contribution data, and insights that help them adopt more effective AI patterns. The provenance layer captures interaction-mode and session data to enable coaching, not to produce punitive scorecards. Privacy is configurable along four rungs, from local-only capture where nothing leaves the machine, to aggregate-only mode, to full identified replay by explicit approval, and different teams in the same organization can operate at different rungs.
Can this replace our existing developer analytics platform?
Exceeds AI does not replace tools like Jellyfish, LinearB, or Swarmia. Those platforms track metadata such as PR cycle times, deployment frequency, and review latency, and they remain useful for traditional delivery metrics. Exceeds AI acts as the AI intelligence layer that sits on top of your existing stack, providing the commit-level attribution and longitudinal outcome tracking that metadata-only tools cannot deliver. Most customers run Exceeds AI alongside their existing tools and integrate it with GitHub, GitLab, Azure DevOps, JIRA, and Linear within their current workflows.
Conclusion: Close the AI Attribution Gap in Engineering
The measurement gap in AI-era engineering stems from attribution, not from a lack of data. Engineering teams are generating more commit and PR data than ever. Without knowing which lines are AI-generated, by which tool, in which mode, and what happened to those lines 30 or 90 days later, every dashboard describes activity rather than decision quality.
The seven steps in this playbook, which include instrumenting every tool, establishing baselines, isolating contributions, tracking longitudinal outcomes, identifying override and rework patterns, quantifying ROI, and converting measurement into coaching, form a complete system. Each step depends on the one before it, and the entire system depends on commit-level provenance that lives in your repository and travels with your code.
Exceeds AI with Exceeds Ink delivers authoritative, portable, line-level attestation across the full multi-tool AI coding landscape, pairs that provenance with 30-plus-day outcome tracking, and converts the resulting signal into coaching actions distributed directly into engineers’ own AI agents. Setup takes hours. First insights arrive within the hour. Board-ready ROI reports are available within weeks.
Stop guessing and start measuring AI’s impact on your engineering decisions with a free pilot.