Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: July 8, 2026
Key Takeaways
- DORA metrics were built for human-written code and become misleading when AI generates 30–70% of committed code across many tools.
- The four metrics miss developer experience degradation, review overload, and burnout that hide behind green delivery numbers.
- DORA acts as a lagging indicator and misses the 60–90 day quality cliff where velocity improves while technical debt and incidents rise later.
- Teams can game DORA scores by shipping AI-generated boilerplate and configuration changes without increasing meaningful engineering output.
- Exceeds AI closes these gaps by adding commit-level AI attribution; connect your repo and start your free pilot today.
1. Developer Experience Gaps Behind Green DORA Scores
Engineering leaders see a consistent pattern. DORA scores look green while engineers report exhaustion, rising cognitive load, and eroding trust in the codebase. The metrics say the team is performing, yet the team experiences mounting strain.
- The Harness State of Engineering Excellence 2026 report found that 89% of leaders say their current metrics accurately reflect AI’s impact, yet 94% say key factors including tech debt, validation time, and developer burnout are missing from those same metrics, which exposes an experience gap DORA cannot close.
- 81% of respondents in the same report say developers spend more time in code review since adopting AI coding tools, yet this review burden never appears in DORA’s four metrics.
- DORA metrics cannot reveal review overload on senior engineers, trust erosion in AI-generated code, rising cognitive load, or burnout hidden behind acceptable delivery numbers, according to analysis of the 2025 DORA State of AI-assisted Software Development report.
The solution starts with understanding how engineers interact with AI tools during real work. Exceeds Ink’s interaction-mode classification tracks five distinct patterns: plan for structured workflows, ask for query-driven usage, agent for autonomous execution, edit for human-guided refinement, and headless for batch automation. These modes reveal which patterns correlate with review burden and cognitive load. This session-level signal explains why some teams experience burnout while their DORA scores remain green.
2. Quality Problems That DORA Sees Too Late
AI-era teams experience quality issues long before DORA metrics reflect trouble. By the time DORA shows a problem, technical debt has already compounded across several sprints. The metrics confirm what went wrong instead of predicting risk while work is still in motion.
- AI has made code production nearly free while leaving ownership costs untouched, creating a 90-day divergence where velocity metrics improve immediately but quality failures appear 60–90 days later.
- Code churn, the percentage of code revised within two weeks of being written, has more than doubled since widespread AI adoption, rising from roughly 3.3% in 2021 to 5.7–7.1% in recent years, and rising churn at day 30 predicts incident rate jumps by day 90.
- DORA metrics function as lagging indicators that become known only after a client receives a deliverable, whereas leading indicators predict deliverable quality while work remains in progress, and management intervention has the greatest leverage on leading indicators.
Exceeds AI addresses this with longitudinal outcome tracking anchored to Exceeds Ink’s per-commit attestation. The platform monitors AI-touched code over 30, 60, and 90 days for incident rates, rework patterns, and maintainability signals before these issues surface as DORA failures.
3. Inflated DORA Scores from AI-Generated Throughput
AI coding tools make it easy to inflate throughput without increasing real value. Teams can hit elite DORA classifications while shipping mostly low-impact AI-generated changes. The metrics reward volume, and AI produces volume cheaply.
- Teams can move from “medium” to “elite” DORA classification by shipping AI-generated boilerplate and configuration changes, increasing deployments from five to twenty per week without a corresponding increase in meaningful output or team capability.
- Faros AI telemetry from over 10,000 developers showed individual developers merged 98% more pull requests with AI assistance, yet organizational DORA metrics remained essentially flat.
- DORA metrics are diagnostic tools, not competition rankings, and treating them as a leaderboard encourages teams to game the numbers rather than reflect true value delivery.
Exceeds AI counters this with AI Usage Diff Mapping. The system identifies exactly which commits and pull requests are AI-touched at the line level. Leaders can separate AI-generated boilerplate from genuine engineering output when reporting to stakeholders.
4. Hidden Technical Debt Behind Strong DORA Metrics
AI accelerates technical debt in ways DORA cannot see. Teams maintain healthy DORA scores for months, then a senior engineer leaves or a critical module needs refactoring and the codebase proves far less understood than the metrics implied.
- AI accelerates technical debt accumulation 10–50 times faster than traditional human coding because shortcuts become invisible to the merging developer and are not recognized as debt, causing the time spent debugging and maintaining AI code to eventually exceed initial productivity gains.
- GitClear found a 4x increase in code cloning, with an eightfold spike in copy-paste patterns in 2024, in AI-assisted codebases, yet this metric is absent from sprint dashboards, OKRs, and engineering scorecards.
- High DORA scores can mask critical technical debt, brittle architectures, and unsustainable practices, creating a “DORA paradox” where strong aggregate metrics hide deteriorating system health.
Exceeds AI surfaces this risk through AI vs. Non-AI Outcome Analytics. The platform tracks whether AI-touched code requires more follow-on edits, generates higher incident rates, or shows lower test coverage over time. Leaders gain a quantified, board-ready signal for invisible debt.
5. Multi-Tool AI Usage That DORA Cannot Attribute
Modern engineering teams rely on several AI tools at once. A single sprint can involve Cursor for feature work, Claude Code for large refactors, Codex for batch transforms, and GitHub Copilot for autocomplete. DORA aggregates all of this into a single stream of activity without attribution.
- DORA metrics have an attribution gap and cannot distinguish AI-assisted code from human-authored code, making it impossible to determine whether improvements in deployment frequency come at the cost of higher change failure rates due to harder-to-review or maintain code.
- A 2026 arXiv paper found that DORA metrics capture only part of the productivity factors relevant to AI coding assistants, making them incomplete for evaluating AI ROI.
- Organizations estimate approximately 31% of developer time is now consumed by invisible work such as reviewing AI-generated code, fixing bugs, and context switching between tools, and DORA does not surface this time cost.
Exceeds AI closes this attribution gap through Exceeds Ink’s first-class adapters for Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf, with lighter-weight detection across roughly 50 additional AI tools. Leaders see aggregate AI impact across the full toolchain instead of one vendor’s slice.
6. Missing Link Between DORA Metrics and Business Value
Executives need a clear connection between AI spend and business outcomes, which DORA cannot provide. Deployment frequency and lead time describe engineering output, not revenue, margin, or risk reduction.
- Only 6% of enterprises succeed with AI at scale.
- As average AI tool usage rose 65% across studied organizations, median pull request throughput increased by just under 8%, a modest gain that DORA cannot connect to revenue or cost savings.
- Only a small minority of engineering leaders believe existing frameworks can address the gaps in measuring AI’s true business impact, according to the Harness State of Engineering Excellence 2026 report.
Exceeds AI measures business linkage through AI vs. Non-AI Outcome Analytics paired with token cost capture from Exceeds Ink. The platform correlates raw token spend with shipped output, cycle time, and incident rates. Finance and engineering receive a shared Agentic ROI signal they can act on together.
7. Metric Gaming and Misuse at the Team Level
The business-value gap creates a second problem. When leaders lack meaningful ROI signals, they default to treating DORA scores as the primary performance indicator. When DORA scores become the main signal, teams optimize for the metric instead of the outcome, and AI amplifies this behavior.
- A team of eight engineers increased from five to twenty weekly deploys after adopting AI coding tools, with fifteen of the additional deploys consisting of AI-generated test files, configuration changes, boilerplate scaffolding, and small refactors rather than new features.
- DORA’s aggregation at the team or organization level obscures concentration patterns, including knowledge silos and single points of failure that make “high-performing” teams fragile to the departure of key individuals.
- In the AI era, output is abundant while judgment, direction, and quality control become the production constraints, which makes volume-based metrics such as DORA deployment frequency actively misleading.
Exceeds AI addresses this with Best Practices Insights. A LangGraph-backed analysis pipeline distills real AI-coding patterns into the top skills worth scaling, sorted by confidence instead of vanity. Managers coach toward durable outcomes instead of metric targets.
Strategic Implications for AI-Era Engineering Leaders
DORA still provides a necessary baseline. Deployment Frequency, Lead Time for Changes, Mean Time to Recovery, and Change Failure Rate continue to describe the delivery pipeline accurately. The issue is not that DORA is wrong. The issue is that DORA is incomplete in ways that become actively misleading when AI generates a large share of committed code.
The 2025 DORA State of AI-assisted Software Development report found that higher AI adoption correlates with increases in both software delivery throughput and software delivery instability, which creates a tension that aggregate DORA metrics cannot resolve. These metrics carry no information about which code was AI-generated, by which tool, in which mode, or what happened to that code 90 days later.
AI provenance forms the missing layer. Without commit-level attribution across every tool an engineering team uses, DORA improvements remain unverifiable. A 20% reduction in lead time could reflect genuine engineering efficiency, AI-generated boilerplate inflating deploy counts, or a senior engineer absorbing review burden that will eventually cause burnout. DORA cannot distinguish these scenarios because it lacks authorship data. Exceeds Ink solves this by writing a portable, line-level attestation alongside every commit, recording the tool, model, session, interaction mode, and timestamp for every AI-touched line. This provenance data, stored as a Git Note in the team’s own repository, makes it possible to segment every DORA improvement by AI contribution and reveal whether velocity gains are sustainable or masking hidden costs.
Engineering leaders who treat DORA as sufficient are not measuring AI ROI. They are measuring a proxy that AI has learned to inflate. The strategic move is to keep DORA as the delivery foundation and add an AI provenance layer that turns those four metrics into trustworthy intelligence.
Connect my repo and start my free pilot
Frequently Asked Questions
What are the biggest limitations of DORA metrics for AI-era engineering teams?
DORA metrics were designed when humans wrote all production code. The four core metrics, Deployment Frequency, Lead Time for Changes, Mean Time to Recovery, and Change Failure Rate, measure delivery pipeline performance but carry no information about code authorship, AI tool usage, or long-term quality outcomes. When AI generates 30–70% of committed code, teams can achieve elite DORA classifications by shipping AI-generated boilerplate and configuration changes without a corresponding increase in meaningful output. The metrics also cannot detect the 60–90 day quality cliff that commonly follows rapid AI adoption, where velocity metrics improve immediately while technical debt and incident rates rise weeks later.
Can DORA metrics be gamed with AI coding tools?
Yes, and often unintentionally, as detailed in Section 3. The core mechanism is straightforward. AI coding assistants compress Lead Time for Changes by writing code in seconds, even as review time increases due to larger and more complex AI-generated diffs. This pattern creates a gap between apparent velocity and actual team capacity.
What metrics should engineering leaders add to DORA to measure AI impact accurately?
The 2025 DORA framework itself added AI Code Share, Code Durability, and Complexity-Adjusted Throughput as AI-era supplements. Practitioners recommend layering five additional signals on top of the core four: AI code share as a percentage of merged code, AI versus human pull request cycle time, AI code churn rate, AI suggestion acceptance trend over time, and pull request review load per senior engineer. Beyond these, longitudinal outcome tracking, which monitors AI-touched code for incident rates, rework patterns, and maintainability issues at 30, 60, and 90 days, provides the only reliable way to detect technical debt before it surfaces as a production crisis. Token cost per shipped outcome is also emerging as a critical signal for connecting AI investment to business value.
Why does multi-tool AI usage make DORA metrics less reliable?
Most engineering teams in 2026 use several AI coding tools simultaneously, such as Cursor for feature development, Claude Code for large refactors, Codex for batch tasks, and GitHub Copilot for autocomplete. DORA metrics aggregate all code contributions without distinguishing tool, model, or interaction mode. A team could see identical DORA scores whether their AI usage is concentrated in one well-governed tool or fragmented across five tools with inconsistent quality outcomes. Without per-tool attribution, leaders cannot identify which tools drive the best results, which teams use AI effectively, or where AI-generated code introduces disproportionate risk, all of which matter for governing AI investment at scale.
How does Exceeds AI address DORA metrics limitations without replacing existing tools?
Exceeds AI sits alongside existing developer analytics platforms instead of replacing them. DORA metrics from tools like LinearB or Swarmia remain useful as delivery baselines. Exceeds AI adds the AI provenance layer those tools cannot provide, using Exceeds Ink’s line-level attestation stored as a Git Note in the team’s own repository. Every DORA metric becomes segmentable by AI code share, tool, and interaction mode. Leaders can answer board questions about AI ROI with specific commit and pull request level evidence rather than adoption statistics or sentiment surveys.
Conclusion: Making DORA Trustworthy in the AI Era
DORA metrics are incomplete rather than broken. The four core metrics still describe delivery pipeline performance accurately for the inputs they were designed to measure. AI coding tools have changed those inputs in fundamental ways.
The seven limitations described here, from missing developer experience signals to multi-tool attribution gaps and structural gaming incentives, share a common root cause. Teams lack line-level AI code attribution. Without knowing which lines are AI-generated, by which tool, in which mode, and what happened to them over the following 90 days, DORA improvements remain unverifiable and potentially misleading.
Exceeds Ink supplies the missing provenance layer. Every commit carries a portable, auditable attestation. Every DORA metric becomes segmentable by AI authorship. Every board question about AI ROI becomes answerable with commit-level evidence rather than adoption statistics. DORA stays as the delivery foundation it was always meant to be and finally becomes trustworthy intelligence for the AI era.