Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: July 2, 2026
Key Takeaways
- Engineering leaders at 100–1,000-engineer companies need code-level AI contribution analysis to prove ROI on tools like Cursor, Claude Code, Codex, GitHub Copilot, and Windsurf.
- Traditional metadata dashboards and developer surveys fall short because they cannot deliver line-level attribution, token-cost tracking, or longitudinal outcome correlation.
- Client-level capture with portable Git Notes attestations outperforms heuristic detection and closed-source daemons for accuracy, auditability, and security.
- Platforms that combine provenance with prescriptive coaching and skill-transfer pipelines deliver faster adoption and measurable behavior change across teams.
- Exceeds AI stands out by delivering authoritative, multi-tool attribution and actionable insights, so you can start your free pilot today.
What AI Contribution Analysis Covers
AI contribution analysis identifies, attributes, and measures the business impact of AI-generated code within an engineering organization. A complete solution captures which lines of code came from which AI tool, in which interaction mode, and at what token cost. It then connects those attributions with downstream outcomes such as cycle time, defect density, rework rates, and long-term incident rates. Leaders can prove ROI, and managers can coach teams toward more effective adoption patterns.
Evaluation Framework for AI Contribution Platforms
Eight dimensions separate meaningful platforms from dashboard noise. Implementation model determines whether attribution is authoritative through client-level capture or estimated through heuristics and watermarks. Data sources range from Git metadata and survey responses to actual code diffs and on-machine session data. Visibility depth spans aggregate adoption stats at one end and line-level, turn-level attribution at the other.
Actionability distinguishes platforms that deliver prescriptive coaching from those that stop at descriptive charts. Security and privacy covers how prompt data is handled, whether attestations live in the customer’s own repo, and the operational footprint on developer machines. Multi-tool support matters because 84% of developers are now using or planning to use AI tools, and 51% of professional developers use them daily, and most teams use more than one.
Setup speed determines how quickly leaders can answer board questions. Team-size fit reflects where each platform delivers the most value.
Platform Comparison Overview
The following sections examine five solution categories in depth: Exceeds AI’s client-level capture with portable attestations, DX’s survey-based developer experience platform, Git AI’s open-source Git Notes approach, metadata-only platforms built for pre-AI SDLC tracking, and native vendor analytics limited to single-tool visibility. Each platform review follows a consistent structure covering deployment, analytics depth, AI-specific capabilities, governance, workflow support, best fit, and limitations.
Connect my repo and start my free pilot.
1. Exceeds AI
Exceeds AI is an AI-impact analytics platform built for the multi-tool AI coding era. Its core differentiator is Exceeds Ink, an on-machine provenance layer that captures AI authorship across Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf at the line level. Ink writes a portable, machine-readable attestation as a Git Note at refs/notes/exceeds-ink alongside every commit.

The attestation lives in the customer’s own repository, so anyone with repo access can audit it. The record is portable across forks and mirrors and never locked inside a vendor’s cloud. Lines that cannot be confidently attributed are recorded as unknown_lines rather than silently assigned, which supports governance and legal review.
The platform pairs that provenance with outcome analytics such as AI vs. non-AI cycle time, rework rates, defect density, and longitudinal incident tracking on AI-touched code 30 or more days after merge. A LangGraph-backed Best Practices Insights pipeline distills real team patterns into the top three skills worth scaling. The ink-prompting-coach skill installs directly into a developer’s own Claude Code or Cursor agent, which closes the loop between measurement and behavior change.

Interaction-mode classification covers plan, ask, agent, edit, and headless modes per session. This detail enables coaching that matches how engineers actually work instead of generic tips. Exceeds AI founder Mark Hull used Claude Code to develop three workflow tools totaling around 300,000 lines of code at a token cost of approximately $2,000, which illustrates the token-spend-to-output correlation the platform surfaces at scale.
Ink’s architecture avoids the operational drag common to competing capture approaches. There is no long-lived daemon, no PATH-shimmed git binary, and no global git config mutation. Capture fires from standard Git hooks and exits immediately, so the developer’s machine carries no always-on process.
HMAC-SHA256-signed remote ingest with revocable per-machine tokens and LLM-based prompt redaction before persistence address the two most common CISO objections. Privacy is configurable along four rungs: local only, aggregate only, abstracted replay, and full identified replay. Different teams in the same organization can run at different rungs.
Deployment: GitHub, GitLab, and Azure DevOps authorization in minutes, per-machine Ink install, first insights within 60 minutes, and complete historical analysis within four hours. A self-hosted collector option is available.
Analytics depth: Line-level AI vs. human attribution, per-tool and per-model breakdowns, interaction-mode detail, longitudinal outcome tracking beyond 30 days, and token spend correlated with shipped output.

AI-specific capabilities: Five first-class adapters with dedicated checkpoint materializers for Claude Code, Cursor, and Codex, plus lighter-weight detection across up to approximately 50 AI tools and cross-tool outcome comparison.
Governance: Portable Git Notes attestation in the customer’s own repo, policy-expressible structured JSON, token spend per agent and model, and aggregate-only mode controlled by a single environment variable.
Workflow support: Integrations with JIRA, Linear, and Slack (beta), ink-prompting-coach in Claude Code and Cursor, skill transfer and rollback for org-wide pattern distribution, and AI-powered performance review support.

Best fit: Engineering leaders at 50–1,000-engineer companies who need board-ready AI ROI proof and managers who need prescriptive coaching tools, not just another dashboard.
Limitations: Repo access is required for code-level analysis, so organizations that cannot grant scoped read-only access will need the in-SCM deployment option. The strongest value begins at approximately 50 engineers, and very small teams may not yet face the governance and adoption-scaling problems Exceeds is built to solve.
2. DX (GetDX)
DX, now an Atlassian company, is an engineering intelligence platform centered on developer experience measurement. Its AI Code Insights and Agent Experience modules extend that survey-and-sentiment foundation toward AI usage tracking. DX offers broad engineering-intelligence coverage, Atlassian distribution through Jira and Bitbucket, and a deep compliance posture including SOC 2, ISO 27001, ISO 27701, and the Data Privacy Framework.
The capture mechanism for AI attribution is a closed-source CLI daemon that runs continuously on developer machines and transmits aggregates to DX Data Cloud. The weakest tier of its three-tier capture model falls back to filesystem-change heuristics. All attribution data lives in DX Data Cloud, and nothing is written to the customer’s own repository.
This architecture creates two practical problems. Security teams must trust a closed-source binary without direct inspection, and the provenance record does not survive outside the DX platform. There is no Git Notes attestation, no portable audit trail, and no in-agent behavior-change layer. DX’s Agent Experience module relies on self-grading surveys rather than session-level interaction data.
Deployment: Enterprise sales-led with median ARR reported at approximately $51,520 (Vendr), weeks to months for full integration, and SaaS-only with no self-host option.
Analytics depth: Developer experience surveys and AI usage aggregates, with no code-level diff analysis and no longitudinal outcome tracking on AI-touched code.
AI-specific capabilities: AI Code Insights module and Agent Experience self-grading, limited to tools with available telemetry, with filesystem heuristics as a fallback.
Governance: All data resides in DX Data Cloud, with no portable attestation and a closed-source capture mechanism.
Workflow support: Jira, Bitbucket, and the broader Atlassian ecosystem, with no in-agent coaching distribution.
Best fit: Organizations already deep in the Atlassian ecosystem that prioritize developer experience measurement and accept a SaaS-only, survey-anchored approach to AI transformation.
Limitations: DX cannot prove AI ROI at the code level, offers no portable provenance, and uses a closed-source daemon that creates audit friction for regulated buyers. There is no self-host option, and an enterprise sales gate with significant ARR commitment is required before value is visible.
3. Git AI
Git AI is the closest architectural peer to Exceeds Ink in terms of producing line-level AI authorship via Git Notes. It uses an open-source Apache 2.0 core, so security teams can inspect the capture code, which is a meaningful advantage over DX’s closed-source daemon. Git AI writes attestations at refs/notes/ai, and the OSS install is fast for individual developers.
The deployment model creates operational friction at scale. Git AI runs a long-lived per-user daemon with a file lock and dual Unix sockets, a PATH-shimmed git binary, and a destructive global Trace2 mutation that overwrites any existing Trace2 tooling silently. On Windows, the installer writes git.exe into ~/.git-ai/bin/ as a copy of the git-ai.exe binary and prepends that directory to PATH, so EDR, AppLocker, and signing-certificate checks see the wrong binary and require re-pointing at every upgrade.
Attribution is reconciled asynchronously by the daemon after the commit, which means a fast git push immediately after git commit can race ahead of the attestation. The OSS local server has no authentication on the Unix socket beyond owner-only file permissions, and secret redaction is entropy-only, which is weak on low-entropy short tokens and prefixed PATs embedded in URLs. Teams and Enterprise tiers are sales-led.
Deployment: Fast OSS install for individuals, with Teams and Enterprise requiring sales engagement and a self-hosted Enterprise option available.
Analytics depth: Line-level Git Notes attribution with recently added Individual Prompt Analysis and generic prompt tips, but no interaction-mode classification at the session level and no longitudinal outcome tracking.
AI-specific capabilities: Multi-tool coverage via daemon, without dedicated checkpoint materializers equivalent to Ink’s per-tool model.
Governance: Git Notes live in the customer’s own repo and the OSS core is auditable, but there is no HMAC-signed ingest and no LLM-based prompt redaction.
Workflow support: Dashboard and basic analytics, with no in-agent coaching distribution and no skill transfer or rollback pipeline.
Best fit: Developer-led teams that want open-source Git Notes provenance and feel comfortable managing a daemon-based deployment without strict enterprise security requirements.
Limitations: The always-on daemon, PATH shim, and global git config mutation create operational drag and CISO friction. Async attribution creates a race window on fast pushes, and there is no in-agent behavior-change layer or outcome analytics beyond attribution.
4. Metadata-Only Platforms (Jellyfish, LinearB, Swarmia)
Jellyfish, LinearB, and Swarmia were built for the pre-AI era of software development. They measure what happened in the development workflow, such as PR cycle times, commit volumes, review latency, and deployment frequency. These tools remain blind to AI’s code-level impact.
None of these platforms can tell a leader which specific lines in a pull request were AI-generated, whether AI-touched code is higher or lower quality than human-authored code, or which adoption patterns are worth scaling across the organization. That gap limits their usefulness for AI ROI questions.
Jellyfish is commonly positioned as an executive-facing DevFinOps tool for engineering resource allocation and financial reporting. Its time-to-value is a known liability, with setup often extending to approximately nine months before ROI is visible, which conflicts with the pace of AI coding decisions. LinearB focuses on workflow automation and SDLC process metrics, and some users report that its data collection approach raises surveillance concerns, with onboarding friction as a recurring theme.
Swarmia targets DORA metrics and developer engagement via Slack notifications. It offers limited AI-specific context and no mechanism for connecting AI usage to business outcomes. These strengths and weaknesses reflect the era in which the products were designed.
The core limitation is architectural rather than a simple feature gap. Organizational-complexity metrics including team size and management span are among the strongest predictors of defect-proneness. As manager-to-IC ratios stretch from the typical 1:5 toward 1:8 or higher, bandwidth for code review and mentorship shrinks while AI-generated code volumes rise.
Metadata tools cannot detect whether AI code that passes review today will cause incidents 30, 60, or 90 days later. That longitudinal blind spot is where AI technical debt accumulates.
Deployment: Cloud-side aggregation with no on-machine capture. Jellyfish commonly takes approximately nine months to ROI, LinearB requires weeks to months with significant onboarding effort, and Swarmia setup is faster but delivers limited depth.
Analytics depth: PR cycle time, commit volume, review latency, and DORA metrics, with no AI vs. human attribution and no code-level diff analysis.
AI-specific capabilities: No capabilities that connect AI usage to code-level outcomes, and at best, adoption statistics from a single vendor’s telemetry.
Governance: No token spend visibility, no AI authorship attestation, and no longitudinal outcome tracking on AI-touched code.
Workflow support: Jira, Linear, GitHub, and GitLab integrations with executive dashboards, but no AI coaching layer.
Best fit: Teams that need traditional SDLC process metrics, DORA tracking, or engineering resource allocation reporting and are not yet focused on proving AI ROI at the code level.
Limitations: These platforms cannot distinguish AI from human code contributions, cannot prove AI ROI, and cannot identify AI technical debt. They were built for the pre-AI era and are not designed to answer the questions boards are asking in 2026.
Having examined each platform’s capabilities individually, the next step is to understand the fundamental architectural tradeoffs that separate these approaches. These decisions matter more than any single feature.
Cross-Platform Tradeoff Analysis
Metadata vs. code-level. The most consequential divide in this category sits between platforms that read Git metadata and platforms that analyze actual code diffs. Metadata tells you that PR #1523 merged in four hours with 847 lines changed. Code-level analysis tells you that 623 of those lines were AI-generated by Cursor, that those lines required one additional review iteration compared to human-authored lines, and that the AI-touched module had zero incidents in the 30 days after merge.
The first answer acts as a lagging indicator. The second answer provides the proof a board requires. Organizations like Zapier are already tracking token usage at the individual level and drawing conclusions about golden patterns worth multiplying versus anti-patterns to coach away, which requires code-level visibility rather than metadata aggregates.
Single-tool vs. multi-tool. GitHub Copilot Analytics provides acceptance rates and lines suggested for Copilot and then goes dark when an engineer opens Cursor or Claude Code. Given the widespread AI tool adoption documented earlier, most engineering organizations already run three or more AI coding tools simultaneously. A platform that covers only one vendor’s telemetry produces a systematically incomplete picture of AI’s impact.
Descriptive vs. actionable. Dashboards that show adoption rates without prescribing next steps leave managers to guess. The gap between knowing that Team A’s AI PRs have three times lower rework than Team B’s and knowing what to do about it is where most platforms stop. Platforms with in-agent coaching distribution and skill transfer pipelines close that gap.
Lightweight vs. heavy footprint. Always-on daemons, PATH-shimmed git binaries, and global git config mutations create operational drag, fleet ops overhead, EDR friction on Windows, and CISO review cycles that can block deployment for months. Hook-direct architectures that fire only at commit time and exit immediately avoid that overhead.
By Exceeds’ own assessment, heuristic and watermark-based AI detection tops out around 20–25% accuracy, with higher accuracy for some tools and lower for others. These approaches cannot capture interaction mode, token cost, or session context. Client-level capture remains the only path to authoritative attribution.
Connect my repo and start my free pilot.
Selection Guidance for Your Organization
By company size. Teams of 50–1,000 engineers with active AI tool adoption face the sharpest version of the ROI proof and adoption-scaling problem, and Exceeds AI is purpose-built for this range. Teams under 50 engineers may find the governance and multi-tool challenges less acute. Organizations above 5,000 engineers face additional compliance and procurement complexity that warrants a direct conversation about timing and deployment architecture.
By AI adoption stage. Early-stage adopters running a single AI tool can extract initial value from that tool’s native analytics. Teams running two or more AI tools simultaneously, which describes the majority of engineering organizations in 2026, need a tool-agnostic platform with per-tool fidelity. Organizations that have already deployed AI tools broadly and now need to prove ROI to the board require code-level provenance rather than adoption statistics.
By security requirements. Organizations that can grant scoped read-only repo access unlock the full value of code-level analysis. Those with stricter requirements should evaluate in-SCM deployment options and self-hosted collector configurations. Regulated buyers should prioritize platforms where the capture code is auditable and the attestation lives in the customer’s own repository instead of a vendor’s proprietary cloud.
By stakeholder needs. CFOs and CTOs focused on engineering resource allocation may find Jellyfish’s financial reporting useful alongside a code-level AI platform. Engineering managers who need prescriptive coaching tools, not just dashboards, require a platform with actionable insights and in-agent coaching distribution. Individual contributors who feel skeptical of monitoring tools respond better to platforms that deliver personal value, such as AI-powered coaching and performance review support, rather than surveillance-style tracking.
Connect my repo and start my free pilot.
Implementation Considerations for Rollout
Repo access. Code-level AI attribution requires read access to the repository. The security conversation is easier when the platform addresses the three most common CISO objections directly. Minimal code exposure comes from transient analysis, where repos are analyzed in seconds and not stored permanently. Prompt privacy comes from LLM-based redaction before any data leaves the organization. Control comes from HMAC-signed ingest with revocable tokens plus a self-host option for the most restrictive environments. Platforms that have passed formal enterprise security reviews, including Fortune 500 processes spanning two months, provide the most credible evidence for CISO approval.
Rollout complexity. On-machine capture requires installation on developer machines, so the operational footprint matters. A lightweight binary that fires from Git hooks and exits immediately is far easier to deploy and maintain across a fleet than an always-on daemon with a PATH shim and global git config mutation. Per-repo opt-in with no global configuration changes reduces the blast radius of any rollout issue.
Stakeholder alignment. Engineering leaders need board-ready ROI reports. Managers need actionable coaching tools. Individual contributors need to see that the platform delivers personal value through coaching, performance review support, and skill development rather than punitive monitoring. Platforms that address all three audiences reduce adoption friction and build the organizational trust required for sustained use.
Privacy. Prompt content is sensitive, so privacy controls matter. Platforms that offer configurable privacy rungs, from local-only capture with nothing leaving the machine to full identified replay with explicit approval, allow organizations to start conservatively and expand access as trust grows. Different teams in the same organization may require different privacy configurations.
Value validation. The time from setup to first meaningful insight strongly predicts whether a platform survives its pilot. Platforms that deliver first insights within 60 minutes and complete historical analysis within four hours allow leaders to validate value before committing to a full deployment. Platforms that require months of integration work before any signal appears carry significant adoption risk.
Connect my repo and start my free pilot.
Frequently Asked Questions
How do heuristic AI detection and client-level capture differ for ROI proof?
Heuristic detection infers AI authorship after the fact by looking for patterns such as large volumes of code written in a short window, watermarks left by certain tools, or commit message keywords. These signals are fallible and, by Exceeds’ own assessment, top out around 20–25% accuracy across tools. Client-level capture observes what actually happens on the developer’s machine at the moment the work is done, including which tool was used, which interaction mode was active, how many tokens were spent, and which lines were produced by the AI versus typed by the engineer. Only client-level capture produces the authoritative, line-level attribution required to prove AI ROI to a board, answer a patent examiner’s questions about AI’s role in a given file, or trace an incident back to the exact session and tool that produced the relevant code.
Can a single platform cover Cursor, Claude Code, Codex, GitHub Copilot, and Windsurf simultaneously?
Yes, although the depth of coverage varies significantly by platform. Platforms built around a single vendor’s telemetry, such as GitHub Copilot Analytics, go dark when engineers use other tools. Platforms with tool-agnostic capture architectures and dedicated per-tool adapters can attribute code across all five tools in a single view. The most capable implementations use dedicated checkpoint materializers for the highest-volume tools, including Claude Code, Cursor, and Codex, and lighter-weight detection for the broader ecosystem, covering up to approximately 50 AI tools. Cross-tool outcome comparison, such as whether Cursor or Copilot drives better results for a specific team or task type, becomes possible only when attribution is consistent across tools.
How should engineering leaders handle security and privacy when a platform needs repo access?
CISO reviews focus on a few core questions. Leaders need to know whether source code persists on the vendor’s servers or is analyzed transiently. They also need to know whether the capture mechanism is open to inspection or a closed-source binary. Prompt content handling matters, including whether prompts leave the developer’s machine and whether they are redacted before persistence. The location of the attestation also matters, specifically whether it lives in the customer’s own repository or only in the vendor’s cloud, and whether the platform offers a self-host option.
Platforms that analyze code transiently without permanent storage, use auditable capture code, apply LLM-based prompt redaction, write portable attestations to the customer’s own repo, and offer configurable privacy rungs address the most common objections. Evidence of completed enterprise security reviews, including formal processes at regulated organizations, provides additional assurance.
How long does it take to go from signup to board-ready AI ROI data?
Time-to-value varies dramatically by platform. Metadata-only platforms that require deep integrations with project management tools, HR systems, and financial data can take months to produce meaningful signal, and Jellyfish is commonly cited as taking approximately nine months to ROI. Platforms built on lightweight repo authorization and on-machine capture can deliver first insights within 60 minutes of setup, complete historical analysis within four hours, and board-ready ROI reports within weeks. Platforms that analyze code directly from the repository without extensive data pipeline construction reach value faster than those that depend on aggregating metadata from multiple upstream systems.
When is an AI contribution analysis platform not the right investment?
Several scenarios indicate a poor fit. Teams under 50 engineers typically face less acute governance and adoption-scaling challenges, and the platform’s leverage grows with team size. Organizations whose primary need is traditional SDLC process metrics, such as DORA, deployment frequency, and change failure rate, without AI-specific context are better served by LinearB or Swarmia for that use case. Companies that cannot grant read-only repo access due to compliance constraints, even with in-SCM deployment options, will not be able to unlock code-level attribution. Organizations seeking punitive monitoring tools rather than coaching and enablement platforms will also struggle, because the most effective AI contribution analysis tools are designed around trust and two-sided value, not surveillance.
Making the Right Choice for AI Contribution Analysis
The decision comes down to three lenses. First, leaders must decide what level of attribution they require. Adoption statistics are sufficient for early pilots, but board-level ROI proof requires code-level provenance. Second, they must consider the operational and security context. Always-on daemons and closed-source binaries create friction that hook-direct architectures avoid, and attestations that live in the customer’s own repository remain more durable than those locked in a vendor’s cloud.
Third, leaders must decide what happens after the measurement. Platforms that translate attribution data into prescriptive coaching, skill transfer, and in-agent guidance produce compounding returns that descriptive dashboards cannot match.
Metadata-only platforms remain useful for traditional SDLC metrics. Client-level capture platforms are the only category that can answer the questions engineering leaders face in 2026: which lines are AI-generated, by which tool, at what cost, and with what effect on quality and delivery over time.