The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

CIRCLE offers a six-stage protocol that turns stakeholder concerns into measurable signals, closing the gap between model-centric benchmarks and real-world AI impacts; it aims to produce comparable, context-aware evidence to support governance based on how systems materialize over time.

CIRCLE: A Framework for Evaluating AI from a Real-World Lens
Reva Schwartz, Carina Westling, Morgan Briggs, Marzieh Fadaee, Isar Nejadgholi, Matthew Holmes, Fariza Rashid, Maya Carlyle, Afaf Taïk, Kyra Wilson, Peter Douglas, Theodora Skeadas, Gabriella Waters, Rumman Chowdhury, Thiago Lacerda · February 27, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Reva Schwartz unresolved corpus identity
  2. Carina Westling unresolved corpus identity
  3. Morgan Briggs unresolved corpus identity
  4. Marzieh Fadaee unresolved corpus identity
  5. Isar Nejadgholi unresolved corpus identity
  6. Matthew Holmes unresolved corpus identity
  7. Fariza Rashid unresolved corpus identity
  8. Maya Carlyle unresolved corpus identity
  9. Afaf Taïk unresolved corpus identity
  10. Kyra Wilson unresolved corpus identity
  11. Peter Douglas unresolved corpus identity
  12. Theodora Skeadas unresolved corpus identity
  13. Gabriella Waters unresolved corpus identity
  14. Rumman Chowdhury unresolved corpus identity
  15. Thiago Lacerda unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Reva Schwartz provider ID
  2. Carina E. I. Westling provider ID
  3. Morgan Briggs provider ID
  4. Marzieh Fadaee provider ID
  5. Isar Nejadgholi provider ID
  6. Matthew Holmes provider ID
  7. Fariza Rashid provider ID
  8. M. Carlyle provider ID
  9. Afaf Taïk provider ID
  10. Kyra Wilson provider ID
  11. Peter Douglas provider ID
  12. Theodora Skeadas provider ID
  13. Gabriella Waters provider ID
  14. Rumman Chowdhury provider ID
  15. Thiago Lacerda provider ID
CIRCLE is a six-stage lifecycle framework that operationalizes TEVV's Validation phase by translating stakeholder concerns into measurable, context-sensitive signals to produce prospective, comparable evidence about deployed AI systems' real-world effects.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI's materialized outcomes in deployment. Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insights into system stability and model capabilities, but they do not provide decision-makers outside the AI stack with systematic evidence of how these systems actually behave in real-world contexts or affect their organizations over time. CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by formalizing the translation of stakeholder concerns outside the stack into measurable signals. Unlike participatory design, which often remains localized, or algorithmic audits, which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context-sensitive qualitative insights to scalable quantitative metrics. By integrating methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline, CIRCLE produces systematic knowledge: evidence that is comparable across sites yet sensitive to local context. This, in turn, can enable governance based on materialized downstream effects rather than theoretical capabilities.

Summary

Main Finding

CIRCLE (Contextualize, Identify, Represent, Compare, Learn, Extend) is a six-stage, lifecycle framework for evaluating AI systems from a real-world, socio-technical perspective. It operationalizes the “Validation” stage of TEVV by translating stakeholder priorities outside the AI stack into measurable constructs and coordinated evaluation activities (stakeholder elicitation, red teaming, field tests, longitudinal monitoring). The framework closes the “reality gap” between model-centric benchmarks and the material downstream effects of deployed AI, producing systematic, comparable evidence that is sensitive to local context and actionable for deployment, governance, and investment decisions.

Key Points

  • Six-stage lifecycle:
    • Contextualize (elicit stakeholder priorities → context brief)
    • Identify (operationalize constructs → evaluation design plan)
    • Represent (execute tests → execution plan)
    • Compare (synthesize outcomes → findings synthesis report)
    • Learn (translate to stakeholder actions → insights brief)
    • Extend (continuous monitoring → monitoring plan)
  • Treats contextual heterogeneity as signal, not noise; prioritizes ecological, construct, and consequential validity.
  • Bridges qualitative stakeholder insights and scalable quantitative metrics via a construct operationalization schema.
  • Integrates a spectrum of methods (in‑silico benchmarking → usability studies → red teaming → field trials → longitudinal studies) into a traceable pipeline so results are comparable across sites while remaining context-sensitive.
  • Produces deliverables at each stage to create an auditable evaluation trail that supports deployment decisions beyond the AI stack.
  • Designed to be iterative and interoperable with MLOps and TEVV processes; emphasizes stakeholder-centered evidence for real-world outcomes rather than only technical capability metrics.

Data & Methods

  • Mixed-methods, lifecycle approach:
    • Qualitative elicitation: interviews, workshops, process mapping to capture tacit and explicit stakeholder priorities.
    • Construct design: define latent constructs (e.g., “over-reliance”, “cognitive offloading”) and map to observable indicators.
    • Evaluation planning: choose methods along a continuum (benchmarks, usability studies, red teaming, field trials, longitudinal studies); specify sampling (stratified, oversampling non-users), human subjects protocols, incentives, instrumentation.
    • Execution: large, diverse participant groups; capture interaction transcripts, logs, incident reports; augment with automated analysis (e.g., LLM review of documentation) where appropriate.
    • Analysis: synthesize qualitative and quantitative signals; link outcomes to constructs; generate findings report.
    • Monitoring: continuous tracking of post-deployment signals (usage patterns, incidents, productivity measures), with iterative feedback to improve system design and governance.
  • Deliverables aligned to stages: context brief, evaluation design plan, execution plan, findings synthesis report, stakeholder insights brief, continuous monitoring plan.
  • Causal and quasi-experimental techniques are implied for attributing downstream effects: randomized trials, difference-in-differences, interrupted time series, matching, longitudinal panel analysis; plus descriptive and process-oriented analyses to surface mechanisms.
  • Example vignette: EdTech chatbot — stakeholder-elicited concern (“over-reliance”), operationalized into measurable behaviors in classroom field trials and longitudinal monitoring.

Implications for AI Economics

  • Better measurement of realized ROI and productivity:
    • Moves beyond executive surveys to direct, context-sensitive measures of operational outcomes (task completion rates, decision times, error rates, throughput, learning outcomes, retention of skills).
    • Enables tracking of secondary and tertiary effects that shape long-run productivity and profitability (workflow changes, task reallocation, skill atrophy or upskilling).
  • Reduced information asymmetry and investment risk:
    • Comparable, auditable evaluation outputs across deployments help investors, procurement officers, and managers differentiate vendors and make evidence-based adoption decisions.
    • Systematic evidence lowers uncertainty about realized benefits and harms, improving capital allocation and pricing of AI-enabled services.
  • Informing labor-market and human-capital economics:
    • Ability to measure how AI changes job tasks, authority, and skill requirements enables better estimates of displacement, complementarities, and re-skilling needs; informs workforce planning and training investments.
  • Enabling policy and regulation grounded in measured externalities:
    • CIRCLE’s downstream-effect focus supplies regulators and public purchasers with operationally valid data for standards, procurement criteria, liability regimes, and sector-specific rules (e.g., education, healthcare).
  • Valuation of monitoring and governance infrastructure:
    • The framework highlights the organizational and resource costs for robust, continuous evaluation (stakeholder engagement, field studies, instrumentation). Those costs should be treated as part of total cost of ownership and factored into business cases and public budgets.
  • Improved macro-level measurement and forecasting:
    • Aggregating CIRCLE-style evaluations across organizations/sectors can provide inputs for estimating economy-wide productivity impacts of AI, informing growth models and public policy (e.g., investment in complementary capital, education).
  • Market competition and standardization:
    • Shared constructs and metrics enable cross-provider comparisons, fostering competition based on realized, context-specific performance rather than marketing claims or raw model capability benchmarks.
  • Implications for risk management and insurance:
    • Traceable evidence of real-world behavior and monitoring can inform underwriting, pricing of liability insurance, and contractual risk-sharing between suppliers and adopters.

Practical considerations for economists and decision-makers: - Use CIRCLE outputs (context briefs, findings synthesis) when building business cases, cost–benefit analyses, or forecasting adoption impacts. - Combine CIRCLE evaluations with causal inference methods (RCTs, diff‑in‑diff, panel models) to strengthen attribution of economic outcomes to AI deployments. - Budget for ongoing evaluation and monitoring as operating expenses necessary for accurate ROI assessment and regulatory compliance. - Encourage shared infrastructure (sectoral evaluation labs, public datasets of construct mappings) to reduce per-deployment costs and enable aggregation of evidence for macroeconomic analysis.

If you want, I can: (a) produce a one-page checklist for incorporating CIRCLE evidence into ROI/cost–benefit analyses, or (b) map specific economic metrics and econometric designs to each CIRCLE stage for a chosen sector (e.g., healthcare, finance, education). Which would be most useful?

Assessment

Paper Typetheoretical Evidence Strengthn/a — Paper proposes a conceptual framework without presenting empirical tests, causal estimates, or quantitative evaluation of the framework's effectiveness. Methods Rigormedium — The framework integrates established methods (field testing, red teaming, longitudinal measurement) and addresses practical translation of qualitative concerns into metrics, but it lacks formal validation, operational details for measurement, and empirical demonstrations of reliability or validity. SampleNo empirical sample; the paper is a conceptual/methodological proposal (CIRCLE) synthesizing prior practices (MLOps, audits, TEVV) and methodological components (field tests, red teaming, longitudinal studies) rather than analyzing original data. Themesgovernance org_design GeneralizabilityNot empirically validated — effectiveness across settings is untested, May be resource- and expertise-intensive, limiting adoption in smaller or resource-constrained organizations, Relies on quality of stakeholder elicitation; biased or incomplete stakeholder input will limit validity, Measurement infrastructure and data access requirements vary across firms/sectors and constrain implementation, Local context sensitivity may reduce comparability despite intended standardization, Legal, regulatory, and cultural differences across jurisdictions can affect applicability

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
CIRCLE is a six-stage, lifecycle-based framework that bridges the reality gap between model-centric performance metrics and AI's materialized outcomes in deployment. Governance And Regulation positive alignment between model-centric performance metrics and AI materialized outcomes in deployment
Reading fidelity high
Study strength speculative
not reported
0.02
Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insights into system stability and model capabilities. Output Quality positive insights into system stability and model capabilities
Reading fidelity high
Study strength low
not reported
0.06
These current approaches do not provide decision-makers outside the AI stack with systematic evidence of how these systems actually behave in real-world contexts or affect their organizations over time. Decision Quality negative availability of systematic, deployment-context evidence for non-AI decision-makers
Reading fidelity high
Study strength low
not reported
0.06
CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by formalizing the translation of stakeholder concerns outside the stack into measurable signals. Governance And Regulation positive translation of stakeholder concerns into measurable signals (operationalized validation)
Reading fidelity high
Study strength speculative
not reported
0.02
Unlike participatory design, which often remains localized, or algorithmic audits, which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context-sensitive qualitative insights to scalable quantitative metrics. Decision Quality positive ability to prospectively link qualitative stakeholder insights to scalable quantitative metrics
Reading fidelity high
Study strength speculative
not reported
0.02
CIRCLE integrates methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline. Research Productivity positive integration of diverse evaluation methods into a coordinated pipeline
Reading fidelity high
Study strength low
not reported
0.06
By integrating these methods into a coordinated pipeline, CIRCLE produces systematic knowledge that is comparable across sites yet sensitive to local context. Research Productivity positive generation of systematic, cross-site comparable yet locally sensitive evidence
Reading fidelity high
Study strength speculative
not reported
0.02
CIRCLE can enable governance based on materialized downstream effects rather than theoretical capabilities. Governance And Regulation positive use of evidence about materialized downstream effects to inform governance
Reading fidelity high
Study strength speculative
not reported
0.02

Notes