The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new, auditable telemetry instrument can compute firms' AI adoption stages from existing production logs and distinguish common stall modes in synthetic tests. Its thresholds and real-world validity, however, are not yet proven and require deployment-level calibration.

Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals
Damon A. Young · August 22, 2026
arxiv descriptive low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Damon A. Young unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Damon A. Young provider ID
The paper proposes 'adoption telemetry' and NANTE — a computable, auditable five-stage model plus open-source implementation — that infers enterprise AI adoption stage-progression from production telemetry and discriminates synthetic stall modes, but its thresholds remain unvalidated on real deployments.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We introduce adoption telemetry: a method for measuring enterprise AI adoption by computing change-management stage-progression directly from production usage signals. We contribute (1) a framework unifying pre-deployment evaluation gates, production telemetry, and change-management staging into one instrumented system; (2) NANTE, a concrete five-stage operationalization with defined telemetry thresholds, published openly so they can be tested and disproven; and (3) an open-source reference implementation that distinguishes a healthy cohort from five characteristic adoption-failure modes on synthetic populations with known ground truth. We are explicit that the thresholds are proposed constructs requiring empirical validation against real outcomes -- a research agenda we outline -- not a calibrated model.

Summary

Main Finding

The paper introduces adoption telemetry: a computable, model-explicit method for measuring enterprise AI adoption by mapping production usage signals to staged change-management milestones. It presents NANTE — a concrete five-stage operationalization (Notice, Attempt, Navigate, Transform, Embed) with openly published telemetry thresholds — and an open-source reference implementation that proves computability by distinguishing a healthy adoption cohort from five synthetic adoption-failure modes. The paper is explicit that the proposed thresholds are unvalidated constructs (demonstrating feasibility, not empirical validity) and outlines a research agenda to calibrate them against real outcomes.

Key Points

  • Measurement gap: current measurement traditions (agent evaluation/observability, enterprise usage dashboards, product analytics, change management) each miss what practitioners call adoption — the population-level progression from access to changed work behavior.
  • Adoption telemetry definition and commitments:
    • Object: measures cohorts’ progression through an organizational change (not agent quality or raw activity).
    • Substrate: uses production telemetry already emitted by deployed systems (invocations, sessions, turns, task outcomes).
    • Interpretation: explicit staged change model with computable predicates and open thresholds.
    • Properties: computable (mechanical rules), model-explicit (open thresholds), population-level (cohort-focused, not individual scoring), intervention-mapped (diagnoses map to intervention classes).
  • NANTE (Notice, Attempt, Navigate, Transform, Embed): a five-stage operationalization that operationalizes stage membership as predicates over telemetry (frequency, consistency, task embedding signals). Threshold values are published as falsifiable proposals.
  • Reference implementation: open-source code demonstrates the method on synthetic populations with known ground truth, separating a healthy cohort from five characteristic modes of adoption failure.
  • Limits emphasized by authors: the work shows computability, not validity. Thresholds are unvalidated and require production deployment data and empirical calibration.
  • Scope: most directly applicable to human-initiated assistants (copilots, chat UIs); extensions for human-in-the-loop agents are discussed but need further work.
  • Privacy/ethics design: cohort-level measurement (minimum-size floors) to avoid per-person surveillance and reduce compliance/privacy concerns.

Data & Methods

  • Literature synthesis: compares four measurement traditions (agent evaluation/observability, enterprise usage dashboards, product analytics, change management) and identifies the structural gap.
  • Formal framework: unifies pre-deployment evaluation gates, production telemetry, and change-management staging into a single instrumented system.
  • NANTE operationalization:
    • Stages defined as computable signatures over production signals (e.g., invocation counts, session structure, recurrence / consistency, task outcome integration).
    • Thresholds are explicitly published so they can be tested and falsified.
  • Reference implementation:
    • Open-source code (no URL in the excerpt) that ingests simulated telemetry and applies NANTE predicates.
    • Synthetic-population experiments: generate cohorts with known “ground-truth” adoption behaviors, including one healthy cohort and five adoption-failure modes; validate that the implementation classifies cohorts correctly (demonstrates discriminative computability).
  • Methods borrowed/related: process-mining style use of event logs as authoritative behavioral record; product-analytics methods (activation funnels, cohort retention) adapted to population-level change staging.
  • Empirical limitations: experiments use synthetic data; no calibrated results from real enterprise deployments presented in this paper.

Implications for AI Economics

  • Better micro-level empirical inputs: Adoption telemetry can provide richer, stage-based behavioral measures (depth and embedding of AI use) beyond simple active-user counts. These measures can improve microeconometric analyses of how AI affects productivity, worker tasks, and firm outcomes by distinguishing shallow sampling from structural workflow change.
  • Improved ROI and diffusion estimates: Economists and valuation analysts can use stage-progression signals (e.g., share of teams at Transform/Embed) as leading indicators for sustained productivity gains and durable P&L impact, rather than relying on headline deployment counts or license activations that overstate economic adoption.
  • Calibration of production-function effects: Stage-based telemetry can be used to parameterize models that map AI adoption to labor substitution/complementarity, task reallocation, and total factor productivity, improving counterfactuals in policy and firm-level studies.
  • Forecasting abandonment and sunk-cost risks: Diagnosing stall types (the paper maps failure modes to interventions) can inform expected abandonment rates, maintenance costs, and renewal risk — useful for investors, procurement decisions, and macro projections of AI diffusion.
  • Policy and market-level measurement: If adopted at scale (with appropriate privacy safeguards), NANTE-like, open thresholds could standardize cross-firm measures of “deep AI adoption,” improving comparability across studies and informing regulatory/competition policy that depends on actual changes in firm behavior.
  • Research agenda relevant to economists:
    • Validate telemetry thresholds against firm-level outcomes (productivity, revenue, worker hours) using field deployments and linked administrative data.
    • Develop causal identification strategies tying stage progression to economic impacts (e.g., difference-in-differences around rollout waves, instrumental variables for exogenous exposure).
    • Aggregate cohort-level signals to industry- or macro-level measures while accounting for firm heterogeneity and selection into pilot/testing phases.
    • Study strategic responses and gaming risks (firms might optimize for measured thresholds); design robust metrics or complementarities (survey/perception measures) to mitigate gaming.
  • Caveats for application:
    • Current thresholds are constructs needing empirical calibration — do not treat NANTE’s threshold values as proven signals of productivity or P&L impact yet.
    • Privacy and governance constraints matter: cohort-level design mitigates some concerns, but cross-firm aggregation and linking to outcomes require careful legal and ethical handling.
    • Representativeness: enterprises that allow instrumentation or data sharing may not be representative, so inference to broader populations requires attention to selection bias.

Suggested next steps for AI economists interested in this work - Collaborate on pilots that link NANTE stage outputs to firm accounting or productivity measures to validate thresholds and estimate stage-specific economic effects. - Explore extensions for human-in-the-loop and agentic systems (as the paper notes) and quantify how stage definitions change when AI initiates actions. - Combine telemetry-derived stages with survey instruments (perceptions, managerial support) to model the complementary roles of observed behavior and reported readiness in generating economic returns.

Assessment

Paper Typedescriptive Evidence Strengthlow — The paper provides a conceptual framework (adoption telemetry), an explicit five-stage operationalization (NANTE), and an open-source reference implementation tested on synthetic populations with known ground truth; however, it contains no validation against real-world production deployments or outcomes and the proposed telemetry thresholds are explicitly uncalibrated and presented as constructs requiring empirical validation. Methods Rigormedium — The paper is methodologically careful in defining computable stages, publishing explicit thresholds, situating the approach relative to four measurement traditions, and providing an open-source implementation and synthetic-ground-truth tests; but it lacks empirical validation on real enterprise telemetry, robustness checks across heterogeneous org contexts, and sensitivity analyses of threshold choice, which limits the practical reliability of its empirical claims. SampleReference implementation evaluated on synthetic populations with known ground truth (synthetic event logs designed to instantiate five adoption-failure modes and a healthy cohort); the paper cites industry measurements (e.g., ActivTrak, Microsoft Copilot analytics) for motivation and context but does not use or analyze real enterprise production deployments to validate thresholds or outcomes. Themesadoption org_design productivity human_ai_collab GeneralizabilityValidated only on synthetic data — thresholds uncalibrated to real enterprises, sectors, or tasks, Applies primarily to systems where humans initiate interactions (copilots, chat assistants); extension to fully autonomous/agentic systems is not validated, Requires access to sufficiently granular production telemetry — organizations lacking such logs cannot apply it, Cohort-level design limits individual-level inference (intentionally), reducing applicability where individual adoption heterogeneity matters, Organizational, cultural, regulatory, and workflow differences may alter stage signatures and threshold behavior, Does not capture perceptual or attitudinal determinants of adoption (surveys remain complementary)

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
MIT's NANDA initiative found that 95% of the generative-AI pilots it analyzed delivered no measurable P&L impact. Firm Revenue negative Measurable profit-and-loss impact from generative-AI pilots
Reading fidelity high
Study strength low
n=300
95%
0.09
The share of companies abandoning most of their enterprise AI initiatives increased from 17% to 42% in one year. Adoption Rate negative Share of companies abandoning most of their AI initiatives
Reading fidelity high
Study strength medium
from 17% to 42%
0.18
Gartner predicts that more than 40% of agentic-AI projects will be cancelled by the end of 2027. Adoption Rate negative Projected cancellation rate of agentic-AI projects
Reading fidelity high
Study strength low
n=3412
over 40%
0.09
In a large monitored population, 57% of AI users spend less than 1% of their working hours in AI tools. Automation Exposure negative Share of working time spent using AI tools
Reading fidelity high
Study strength medium
57% spend under 1% of working hours
0.18
Among more than 120,000 workers tracked quarterly, 82% of AI users sustain usage quarter over quarter, but only about 2% reach a level of use consistently embedded in their workflows. Adoption Rate mixed Sustained AI usage and workflow-embedded AI usage
Reading fidelity high
Study strength medium
n=120000
82% sustain usage; about 2% reach consistently embedded use
0.18
The NANTE operationalization defines enterprise AI adoption as progression through five telemetry-based stages: Notice, Attempt, Navigate, Transform, and Embed. Adoption Rate positive Stage progression in organizational AI adoption
Reading fidelity high
Study strength low
not reported
0.09
The open-source reference implementation distinguishes a healthy cohort from five characteristic adoption-failure modes in synthetic populations with known ground truth. Adoption Rate positive Correct classification of adoption cohorts and failure modes
Reading fidelity high
Study strength low
five characteristic adoption-failure modes
0.09
The proposed telemetry thresholds are unvalidated constructs and are not calibrated against real adoption outcomes. Adoption Rate null_result Empirical validity and calibration of adoption-stage thresholds
Reading fidelity high
Study strength speculative
not reported
0.03

Notes