0 cumulative citations
View corpus contextSocial workers co‑design a benchmark to judge whether LLMs can productively challenge practitioners' assumptions; the worker-created rubric aligns closely with practitioners' ratings and distinguishes performance across six leading LLMs.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what "successful" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can "challenge" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their professional and lived expertise. In validations of the worker-created benchmark, we find that there is strong agreement between worker and LLM judge ratings and that the resulting benchmark can differentiate performance across six state-of-the-art LLMs. Based on our case study, we discuss opportunities for future work to support worker-driven AI measurement as a complementary approach to existing top-down AI evaluation approaches.
Summary
Main Finding
Worker-driven AI measurement—where frontline workers collaboratively define what useful AI augmentation looks like and design the evaluation instruments—can produce valid, discriminating benchmarks for LLM augmentation in real work. In a case study with 19 school social workers who co-designed an LLM-as-a-judge rubric and 16 realistic test cases focused on generating reflective questions for teachers, the worker-created benchmark showed strong agreement with worker judgments and could differentiate performance across six state-of-the-art LLMs.
Key Points
- Problem framed: Typical AI evaluations are top-down and often miss what frontline workers value; this can lead to misaligned deployments in occupations like social work and education.
- Proposed approach: Worker-driven AI measurement — a bottom-up, collaborative lifecycle for evaluation that has workers (a) pick use cases, (b) articulate what “good” looks like, and (c) operationalize those concepts into measurement instruments (benchmarks/rubrics).
- Case study context: A local school social work organization exploring LLM support (19 workers across ~30 staff organization).
- Use case chosen by workers: LLM-generated reflective questions to “challenge” social workers’ assumptions about classroom observations and to broaden teachers’ perspectives — characterized as measuring the LLM’s ability to be a “people challenger” rather than a people-pleaser.
- Design principles for the workshops: support expressive participation, enable informed iterative design (inspect instrument behavior on realistic data), and foster meaningful collaboration/co-learning among workers.
- Outcome: Workers jointly created 16 realistic test cases and an LLM-as-a-judge rubric derived from professional and lived expertise. Validation showed strong worker–LLM-judge agreement and the benchmark discriminated among six SOTA LLMs.
- Contributions: (1) articulation of worker-driven AI measurement, (2) an empirical case study showing feasibility, (3) lessons and opportunities for implementing this approach more broadly.
Data & Methods
- Participants: 19 school social workers from a local organization (organization employs ~30 social workers across preschool–elementary schools).
- Timeline: Eight iterative workshops over eight weeks, plus pre-workshop asynchronous prompts and formative interviews with a subset of workers.
- Three-phase collaborative process:
- Identify candidate LLM use cases (pre-workshop submissions; workshop discussion) and select a focused use case. Workers produced 39 current and 18 desired LLM uses (from nine workers) during intake.
- Systematize measurement goals by jointly inspecting model responses and annotating desirable/undesirable properties; surfaced the “people challenger” goal.
- Operationalize into a benchmark: workers designed 16 realistic test cases and iteratively developed an LLM-as-a-judge rubric to rate how well a model’s outputs would help them reflect and expand perspectives.
- Model response generation: For rubric development, each test case was paired with multiple model responses generated from different base models and system instructions (drawn from OpenAI, Anthropic, Google) to expose variation.
- Validation: Compared the LLM-as-a-judge ratings with worker judgments and evaluated six state-of-the-art LLMs on the worker-created benchmark. Reported findings: strong alignment between worker and LLM-judge ratings; benchmark can distinguish relative model performance. (Paper reports “strong agreement” and discrimination across six models; exact numeric reliability/statistics not provided in the excerpt.)
- Tools: Collaborative workshops (Figma boards for synthesis), model customization via OpenAI ChatGPT customization service for the organization’s internal model; facilitators provided light guidance.
Implications for AI Economics
- Better alignment of evaluations with worker value: Benchmarks designed by workers capture dimensions of value (e.g., stimulating reflection, surfacing biases) that standard benchmarks miss. This reduces measurement error in estimating the economic benefits of AI augmentation (productivity, quality of service), improving decision-making about investment and deployment.
- Impacts on assessments of complementarity vs. substitution: Worker-driven metrics emphasize augmentation qualities (e.g., challenging thinking, enabling learning) that signal complementarity with human labor rather than autonomous substitution. Incorporating such metrics into economic estimates can change labor impact projections (e.g., slower displacement, increased skill premiums).
- Procurement and model choice: Organizations (and purchasers) can use worker-derived benchmarks to select or customize models that better deliver on mission-critical, non-quantitative outcomes. This can channel spending toward models that improve organizational outcomes valued by workers, altering demand dynamics in model markets.
- ROI and adoption economics: Worker-created measures can reveal benefits not captured by standard productivity metrics (e.g., improved teacher practice, better child outcomes over time). Capturing these benefits can change cost–benefit calculations for adopting LLMs, potentially justifying higher upfront costs for customized or higher-performing models.
- Incentives for model developers: If buyers (organizations) increasingly require worker-aligned benchmarks, model developers may invest more in capabilities that support human reflection, contextual nuance, and dialogic augmentation—shifting development priorities and market competition.
- Organizational governance and labor bargaining: Worker-driven measurement can empower labor and professional bodies to define acceptable automation boundaries and evaluation standards, which could be used in procurement clauses, compliance checks, or bargaining over work redesign—affecting wages, task allocation, and monitoring regimes.
- Policy and regulation: Regulators and funders evaluating social-impact deployments may adopt worker-derived metrics as complementary evidence of utility, safety, or fairness. This could alter subsidy, procurement, or regulatory criteria, influencing which AI products are deployed in public services.
- Externalities & measurement challenges: Worker-driven benchmarks are context-specific and may be costly to scale. Economists and policymakers should account for heterogeneity in measurement across organizations and the transaction costs of co-design when modeling adoption trajectories and aggregate impacts.
- Research agenda for AI economics: Incorporate worker-defined outcome measures into empirical studies of LLMs’ labor-market effects (e.g., field experiments, difference-in-differences) to better capture welfare effects beyond narrow productivity proxies. Evaluate how worker-driven evaluations shift observable adoption patterns, wage dynamics, and human capital investments.
Limitations and cautions relevant to economic interpretation: - Single-organization, small sample case study; results may not generalize across occupations, firm sizes, or institutional contexts. - Worker-driven metrics are value-laden and may reflect local norms; aggregation across firms/sectors requires careful harmonization to avoid biased macro estimates. - The labor and time cost of co-design (workshops, rubric development, validation) should be included in economic analyses of adopting such evaluation regimes.
Suggested applications for economists and policymakers: - Use worker-driven benchmarks as complementary outcome measures in empirical evaluations of AI deployments. - Factor the value of augmentation-related outcomes (learning, decision quality, trust) into cost–benefit models for AI procurement in public services. - Study how procurement standards that require worker-aligned evaluation affect market structure and innovation incentives in the AI industry.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study proposes worker-driven AI measurement, a bottom-up approach in which workers collaboratively decide which tasks AI should augment, what successful augmentation means, and how it should be measured. Governance And Regulation | positive | Worker participation in defining AI augmentation goals and evaluation criteria |
Reading fidelity
high
Study strength
speculative
|
n=19
|
| Across eight workshops, 19 school social workers collaboratively designed an evaluation benchmark for LLM augmentation in their work. Governance And Regulation | positive | Completion of a worker-designed AI evaluation benchmark |
Reading fidelity
high
Study strength
medium
|
n=19
8 workshops
|
| The workers selected generating reflective questions to challenge their thinking and assumptions about classroom observations as a high-priority LLM use case. Task Allocation | positive | Worker-perceived priority and desirability of an LLM use case |
Reading fidelity
high
Study strength
medium
|
n=8
|
| Workers collectively designed 16 realistic test cases intended to challenge LLM capabilities for generating useful reflective questions. Training Effectiveness | positive | Number and realism of worker-designed benchmark test cases |
Reading fidelity
high
Study strength
medium
|
n=16
16 test cases
|
| The workers identified measuring whether LLMs can act as “people challengers”—prompting reflection on users’ assumptions and biases and expanding their perspectives—as an important evaluation goal. Decision Quality | positive | LLM ability to challenge workers' assumptions, biases, and perspectives |
Reading fidelity
high
Study strength
medium
|
n=13
|
| The workers collaboratively developed an LLM-as-a-judge rubric that operationalized their context-specific notion of meaningful AI augmentation. Governance And Regulation | positive | Contextual validity and operationalization of an AI evaluation rubric |
Reading fidelity
high
Study strength
medium
|
n=19
|
| Validation of the worker-created benchmark found strong agreement between worker ratings and ratings produced by the LLM judge. Decision Quality | positive | Agreement between worker evaluations and LLM-judge evaluations |
Reading fidelity
high
Study strength
low
|
strong agreement
|
| The worker-created benchmark differentiated performance across six state-of-the-art LLMs. Output Quality | positive | Differences in LLM performance on the worker-defined use case |
Reading fidelity
high
Study strength
low
|
n=6
six state-of-the-art LLMs
|
| Before the workshops, nine workers reported 39 current and 18 desired uses of LLMs for assisting their work. Adoption Rate | positive | Breadth of current and desired LLM use cases |
Reading fidelity
high
Study strength
medium
|
n=9
39 current uses and 18 desired uses
|