Written by: Mark Hull, Co-Founder and CEO, Exceeds AI | Last updated: June 27, 2026
Key Takeaways
- AI engineering analytics platforms split into metadata-only tools that estimate impact and code-level platforms that deliver commit-level provenance proof.
- Exceeds AI stands out with its lightweight Rust binary, Exceeds Ink, which captures line-level AI attribution via portable Git Notes without long-lived daemons.
- Only client-level capture provides authoritative multi-tool attribution across Claude Code, Cursor, Codex, Copilot, and Windsurf, unlike heuristic or single-vendor approaches.
- Longitudinal tracking of AI technical debt, incident rates, and rework over 30-plus days requires repo-resident attestation that metadata platforms cannot deliver.
- Engineering leaders seeking board-ready AI ROI proof should start a free pilot with Exceeds AI to measure real commit-level outcomes.
Exceeds AI: Commit-Level AI ROI With In-Agent Coaching
Exceeds AI is the only platform in this comparison built specifically for the multi-tool AI coding era. Its provenance layer, Exceeds Ink, is a lightweight Rust binary that installs on developer machines, fires from standard Git hooks, and writes a structured line-level attestation as a Git Note at refs/notes/exceeds-ink alongside every commit. That attestation records the AI tool, model, session, interaction mode, and token cost for every attributed line, without running a long-lived daemon, without shimming the git binary, and without mutating global Git configuration.
The platform layer converts Ink attestations into board-ready ROI reports and AI versus human outcome analytics across cycle time, rework rates, and 30-plus-day incident rates. It also provides an AI Adoption Map across teams and tools, Best Practices Insights powered by a LangGraph analysis pipeline, and Coaching Surfaces that send guidance directly into the developer’s own Claude Code or Cursor agent through ink-prompting-coach. Exceeds AI founder Mark Hull used Claude Code to build three workflow tools totaling roughly 300,000 lines of code at a token cost of approximately $2,000. That example illustrates the kind of spend-to-output signal the platform surfaces across an entire organization.

Five first-class adapters cover Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf, with lighter-weight detection across up to approximately 50 AI tools. Setup completes in hours. GitHub or GitLab OAuth authorization takes minutes, and first insights are available within 60 minutes. Pricing is outcome-aligned at $49 per manager seat per month (Early Partner Pricing) with no per-contributor data tax. Exceeds Ink is available as a standalone add-on or paired with the platform.

Strengths: Only platform combining portable Git Notes attestation with in-agent coaching, no long-lived daemon, deterministic attribution at commit finalization, longitudinal AI technical debt tracking, multi-tool fidelity across the full AI coding landscape, self-host option available.

Limitations: Repo access is required for code-level analysis, sweet spot is 50–1,000 engineers, SOC 2 Type II certification is in progress.
Best fit: Engineering leaders at 50–1,000 engineer companies who need board-ready AI ROI proof and managers who need prescriptive guidance to scale adoption across teams using multiple AI tools.
Start measuring AI ROI across your team’s tools
DX: Developer Sentiment And Workflow Intelligence
DX is an engineering intelligence platform centered on developer experience surveys, workflow analytics, and two AI-specific modules called AI Code Insights and Agent Experience. Its compliance posture is strong with SOC 2, ISO 27001, ISO 27701, and Data Privacy Framework, and Atlassian distribution through Jira and Bitbucket gives it broad enterprise reach.
DX AI capture relies on an always-on, closed-source CLI daemon that transmits aggregates to DX Data Cloud. Attribution lives in DX Data Cloud only, and nothing is written to the repository as a portable attestation. The weakest tier of its three-tier capture model falls back to filesystem-change heuristics. Security teams must trust the closed-source binary on faith, and the SaaS-only architecture means there is no self-host option.
Strengths: Broad engineering intelligence coverage, strong compliance certifications, Atlassian ecosystem integration, established enterprise sales motion.
Limitations: Closed-source daemon, no portable Git Notes attestation, all attribution locked in DX Data Cloud, SaaS-only with no self-host, enterprise sales gate with a median ARR around $51,520 (Vendr), measures developer sentiment with AI rather than code-level business impact.
Best fit: Enterprises already invested in the Atlassian ecosystem that prioritize developer experience measurement and compliance certifications over code-level AI provenance.
Git AI: Open-Source Provenance With Heavier Footprint
Git AI is the closest architectural peer to Exceeds Ink in the market. It also produces line-level AI authorship via Git Notes, uses an open-source core under Apache 2.0, and supports a self-hosted enterprise tier. For teams that want portable, repo-resident attestation and are comfortable with the operational footprint, Git AI is a genuine option.
The deployment model creates friction for regulated buyers and fleet operations teams. Git AI runs a long-lived per-user daemon with a file lock and dual Unix sockets, installs a PATH-shimmed Git binary, and on Windows git.exe is literally a copy of git-ai.exe, which causes EDR and AppLocker checks to see the wrong binary. It also destructively overwrites trace2.eventTarget globally, which silently clobbers any existing Trace2 tooling. Attribution reconciliation is asynchronous, so a fast git push immediately after git commit can race ahead of the daemon attribution write. Coaching is limited to recently added generic prompt tips called Individual Prompt Analysis, without session-level interaction-mode calibration or the skill-transfer and rollback capabilities Exceeds Ink provides.
Strengths: Open-source core, portable Git Notes attestation, self-hosted enterprise option, genuine client-level capture.
Limitations: Long-lived daemon, PATH-shimmed Git binary creates EDR and AppLocker conflicts on Windows, destructive global Git config mutation, async attribution creates a race window, no in-agent behavior-change layer, Teams and Enterprise tiers are sales-led.
Best fit: Teams comfortable with open-source tooling and a heavier on-machine footprint that do not have strict EDR or Trace2 constraints and do not need in-agent coaching distribution.
Jellyfish: DevFinOps Without AI Provenance
Jellyfish is a DevFinOps platform designed to help CFOs and CTOs understand engineering resource allocation and financial alignment. It aggregates Jira and Git metadata to produce high-level investment reporting. For organizations that need to connect engineering spend to business units, Jellyfish has an established track record.
Jellyfish has no provenance layer and no ability to distinguish AI-generated from human-authored code at the line or commit level. It cannot prove whether AI investments are improving productivity or degrading quality. Setup and time-to-value are the slowest in this comparison. Onboarding commonly extends to two months, and customers report an average of approximately nine months before meaningful ROI is visible.
Strengths: Financial reporting and resource allocation, established enterprise relationships, Jira and Git metadata aggregation.
Limitations: No AI provenance, metadata-only, cannot distinguish AI versus human contributions, nine-month average time to ROI, opaque per-seat enterprise pricing, no actionable guidance for engineering managers.
Best fit: CFOs and CTOs at large enterprises who need engineering investment reporting and are not yet asking code-level questions about AI ROI.
LinearB: Delivery Workflow Metrics Without AI Insight
LinearB measures software delivery workflow such as cycle time, review latency, and deployment frequency, and offers workflow automations called WorkerB to reduce friction in the PR process. It is effective at surfacing where handoffs slow down delivery and at automating routine review assignments.
LinearB operates on metadata only. It cannot distinguish AI-generated from human-authored code, cannot prove AI ROI at the commit level, and cannot track longitudinal outcomes of AI-touched code. Some users have reported that its data collection approach raises surveillance concerns. Onboarding requires significant effort and clean repository data before value is realized.
Strengths: Workflow automation, cycle time visibility, PR process optimization, established mid-market presence.
Limitations: No AI provenance, metadata-only, cannot prove AI ROI, per-contributor pricing penalizes team growth, reported onboarding friction, surveillance concerns raised by some users.
Best fit: Teams optimizing traditional software delivery workflows that are not yet measuring AI-specific outcomes.
Swarmia: Lightweight DORA Metrics For Smaller Teams
Swarmia is a developer productivity platform focused on DORA metrics, team habits, and Slack-based nudges. It is lightweight to set up and surfaces delivery metrics clearly. For teams that want a low-friction way to track traditional productivity signals, Swarmia is accessible.
Swarmia was built for the pre-AI era. It has limited AI-specific context, no provenance layer, and no ability to connect AI tool usage to code-level outcomes. It cannot identify which lines are AI-generated, cannot compare outcomes across Cursor versus Copilot versus Claude Code, and cannot track AI technical debt over time.
Strengths: Fast setup, clean DORA metric dashboards, Slack integration, straightforward per-seat pricing.
Limitations: No AI provenance, metadata-only, no multi-tool AI analytics, no longitudinal outcome tracking, no actionable coaching beyond Slack notifications.
Best fit: Smaller teams that want traditional delivery metrics and are not yet measuring AI-specific productivity or quality outcomes.
Synthesis: Metadata Guesses Versus Code-Level Provenance
The core split in this category is not feature depth but what each platform can actually know. Metadata-only platforms such as Jellyfish, LinearB, and Swarmia observe what happened in the delivery pipeline. Code-level platforms observe what was written and by whom.
Heuristic and watermark-based AI detection, the approach underlying most tools that claim any AI awareness, tops out at around 20–25% accuracy by Exceeds’ own assessment. This accuracy ceiling becomes critical at current adoption levels. Eighty-four percent of developers are now using or planning to use AI tools, and 51% of professional developers use them daily. At that scale, guessing at 20–25% accuracy produces noise rather than signal.
Client-level capture, which observes what actually happens on the developer’s machine at the moment the work is done, is the only path to authoritative attribution. Only two platforms in this comparison have that technology: Exceeds AI through Exceeds Ink and Git AI. DX captures AI activity via an always-on closed-source daemon, but its attribution lives in DX Data Cloud rather than the repository, and its weakest capture tier falls back to filesystem heuristics.
Single-tool telemetry compounds the accuracy problem. GitHub Copilot Analytics reports acceptance rates for Copilot suggestions and remains blind to Cursor sessions, Claude Code rewrites, and Codex batch transforms happening in the same codebase. Exceeds Ink per-tool checkpoint materializers for Claude Code, Cursor, and Codex, plus adapters for Copilot and Windsurf, produce a unified cross-tool attribution record that single-vendor telemetry cannot replicate.
Descriptive dashboards leave managers with numbers and no direction. Exceeds AI Coaching Surfaces and ink-prompting-coach distribute guidance directly into the developer’s own AI agent, which closes the loop between measurement and behavior change. Neither Git AI nor DX delivers in-agent coaching calibrated to session-level interaction-mode data.
The Hidden Cost Of AI Technical Debt
AI tools generate code that often passes initial review cleanly. The risk surfaces 30, 60, or 90 days later as higher incident rates, elevated rework, reduced test coverage, and architectural drift. Metadata-only tools cannot detect this pattern because they observe PR merge status rather than the long-term fate of the merged code.
Wider management spans amplify this risk in AI-heavy environments. Managers who oversee more engineers have less time for detailed code review, and that bandwidth constraint matters when AI tools are generating more code than ever. Organizational-complexity metrics including team size and management span are among the strongest predictors of defect-proneness. As spans widen and bandwidth for mentorship and code review shrinks, quality suffers, and AI-generated code that looks clean at merge time can hide defects that surface weeks later.
Longitudinal outcome tracking requires repo access and per-commit attestation. Exceeds AI tracks AI-touched code over time, monitoring incident rates, rework patterns, and maintainability signals anchored to Ink per-commit Git Notes, and surfaces early warnings before hidden debt becomes a production crisis. No metadata-only platform in this comparison offers equivalent capability.
Track AI technical debt before it becomes a crisis
Build Versus Buy: Provenance Decision Guide For Engineering Leaders
Engineering teams that consider building internal AI attribution tooling need a clear decision framework. The calculus depends on three variables that shape the tradeoffs: internal engineering capacity, time-to-value requirements, and governance obligations.
Internal capacity sets the ceiling for how many tools a homegrown system can support. Building a provenance layer from scratch requires maintaining per-tool adapters for every AI coding tool the team uses, and those tools update frequently. Claude Code, Cursor, and Codex each have distinct session models, checkpoint formats, and attribution edge cases. A homegrown solution that handles one tool accurately will drift as others are adopted. Keeping pace with five first-class adapters plus lighter-weight detection across approximately 50 tools is a sustained engineering investment rather than a one-time project.
Time-to-value is the second variable and often the most visible to executives. A board that expects AI ROI evidence in the next quarter cannot wait for an internal build that takes six months to reach baseline accuracy. Exceeds AI delivers first insights within 60 minutes of authorization and complete historical analysis within four hours, which compresses the feedback loop for both engineering leaders and finance partners.
Governance is the third variable and becomes decisive for regulated organizations. When auditors, patent examiners, or incident responders ask about AI’s role in specific code, they need auditable, machine-readable attestation in a stable versioned schema, and that attestation must survive outside the analytics platform. Exceeds Ink Git Notes attestation lives in the repository and is readable by any Git client. A proprietary internal database does not provide the same portability.
Exceeds AI functions as the AI-intelligence layer that sits atop existing stacks, alongside LinearB, Jellyfish, or Swarmia for teams that use them, rather than as a replacement for the entire analytics toolchain. For most teams, the ongoing adapter maintenance cost, the time-to-value gap, and the governance portability requirement all point toward buying the provenance layer and directing internal engineering capacity toward the product.
Frequently Asked Questions
How accurate is client-level AI provenance compared to heuristic detection?
Heuristic and watermark-based detection, the approach used by most platforms that claim AI awareness, tops out at roughly 20–25% accuracy. These methods look for patterns such as large code volumes written in short windows or watermarks that some tools occasionally embed in output. Both signals are fallible because engineers iterate, rewrite, and mix AI and human edits in ways that defeat pattern matching, and watermarks are inconsistent across tools and versions.
Client-level capture observes what actually happens on the developer’s machine at the moment the work is done, including which tool was active, what the engineer typed, how long the session ran, and which interaction mode was used. Exceeds Ink per-tool checkpoint materializers resolve edit evidence against the actual working tree at commit finalization, so multi-edit Cursor sessions correctly retain human-typed lines and Claude Code rewrites are attributed to Claude. Lines that cannot be confidently attributed are recorded as unknown rather than silently assigned to either category. The result is an attestation that is auditable, conservative, and grounded in direct observation rather than inference.
Which AI coding tools does Exceeds AI support?
Exceeds Ink ships five first-class adapters with deep per-tool fidelity across Claude Code, Cursor, Codex, GitHub Copilot, and Windsurf. Each adapter has a dedicated checkpoint materializer that resolves attribution at commit finalization. Beyond those five, Ink provides lighter-weight detection across approximately 50 AI tools, capturing the underlying model such as Claude or Opus, GPT, and Gemini, and token cost per session. The platform supports the multi-tool reality of 2026, where engineers switch between tools depending on the task and leaders need aggregate AI impact visibility across the entire toolchain, not just the tool a vendor happens to have telemetry for.
What does Exceeds AI security posture look like for teams that need repo access approval?
Exceeds AI has passed enterprise security reviews including a formal two-month evaluation at a Fortune 500 retailer. Key security properties include HMAC-SHA256-signed remote ingest with revocable per-machine tokens, LLM-based prompt redaction before any prompt content is persisted, and an aggregate-only mode that keeps transcripts off the wire entirely through a single environment variable. The platform also supports per-repo opt-in with no global Git config mutation and no PATH-shimmed Git binary, minimal code exposure where repositories exist on servers for seconds and are then permanently deleted, and encryption at rest and in transit. SSO and SAML support, data residency options for US-only or EU-only hosting, an in-SCM deployment option for teams that cannot transfer code externally, and a self-host option where the remote ingest URL is fully configurable round out the posture. SOC 2 Type II certification is currently in progress, and security whitepapers and detailed documentation are available as part of any evaluation.
How quickly does Exceeds AI deliver value compared to other platforms?
Exceeds AI reaches useful insights on a much shorter timeline than metadata-only platforms. GitHub and GitLab OAuth authorization takes approximately five minutes, and repository scoping takes another fifteen. The 60-minute time-to-first-insight window covers initial ingest, attestation processing, and baseline analytics across AI versus human contribution patterns. Complete historical analysis typically completes within four hours, and real-time updates appear within five minutes of new commits. By contrast, Jellyfish onboarding commonly extends to two months with an average of approximately nine months before ROI is visible. LinearB typically requires two to four weeks of setup with significant onboarding friction, and DX enterprise onboarding runs four to six weeks and is sales-led, with a median ARR commitment around $51,520. Exceeds AI Pro plan starts at $49 per manager seat per month with a free seven-day pilot, no per-contributor data tax, and no enterprise sales gate required to get started.
Conclusion: Choosing The Right Lens For 2026 AI ROI
The decision framework for 2026 AI analytics centers on what question leaders need to answer. If the question is “what is our AI investment actually producing, provably, at the commit and PR level, across every tool our engineers use,” only platforms with client-level provenance can provide that answer. Jellyfish, LinearB, and Swarmia were built for the pre-AI era and remain useful for the metadata questions they were designed to address. DX measures developer sentiment toward AI tools. Git AI produces portable Git Notes attestation but carries an operational footprint that creates friction for regulated buyers and fleet operations teams.
Exceeds AI is the only platform that combines portable, auditable, line-level attestation through Exceeds Ink with a behavior-change layer that distributes coaching into the developer’s own AI agent, longitudinal AI technical debt tracking anchored to per-commit provenance, and outcome-aligned pricing that does not penalize team growth. Setup takes hours, and board-ready ROI reports are available within weeks.
Engineering leaders who need to answer “is our AI investment paying off” with hard evidence, rather than sentiment, metadata estimates, or a single vendor’s acceptance rate, have one path to that answer: code-level provenance at the commit and PR level.