The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Social workers co‑design a benchmark to judge whether LLMs can productively challenge practitioners' assumptions; the worker-created rubric aligns closely with practitioners' ratings and distinguishes performance across six leading LLMs.

"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work
Anna Kawakami, Chloe Qianhui Zhao, Renee Shelby, Fernando Diaz, Haiyi Zhu, Kenneth Holstein · August 23, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Anna Kawakami unresolved corpus identity
  2. Chloe Qianhui Zhao unresolved corpus identity
  3. Renee Shelby unresolved corpus identity
  4. Fernando Diaz unresolved corpus identity
  5. Haiyi Zhu unresolved corpus identity
  6. Kenneth Holstein unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Anna Kawakami unresolved corpus identity
  2. Chloe Qianhui Zhao unresolved corpus identity
  3. Renee Shelby unresolved corpus identity
  4. Fernando Diaz unresolved corpus identity
  5. Haiyi Zhu unresolved corpus identity
  6. Kenneth Holstein unresolved corpus identity
Through eight participatory workshops with 19 school social workers the authors co-designed a worker-driven benchmark (an LLM-as-a-judge rubric and 16 test cases) to measure whether LLMs can productively 'challenge' practitioners, and validated that the worker-created rubric aligns with worker judgments and differentiates among six LLMs.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what "successful" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can "challenge" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their professional and lived expertise. In validations of the worker-created benchmark, we find that there is strong agreement between worker and LLM judge ratings and that the resulting benchmark can differentiate performance across six state-of-the-art LLMs. Based on our case study, we discuss opportunities for future work to support worker-driven AI measurement as a complementary approach to existing top-down AI evaluation approaches.

Summary

Main Finding

Worker-driven AI measurement—where frontline workers collaboratively define what useful AI augmentation looks like and design the evaluation instruments—can produce valid, discriminating benchmarks for LLM augmentation in real work. In a case study with 19 school social workers who co-designed an LLM-as-a-judge rubric and 16 realistic test cases focused on generating reflective questions for teachers, the worker-created benchmark showed strong agreement with worker judgments and could differentiate performance across six state-of-the-art LLMs.

Key Points

  • Problem framed: Typical AI evaluations are top-down and often miss what frontline workers value; this can lead to misaligned deployments in occupations like social work and education.
  • Proposed approach: Worker-driven AI measurement — a bottom-up, collaborative lifecycle for evaluation that has workers (a) pick use cases, (b) articulate what “good” looks like, and (c) operationalize those concepts into measurement instruments (benchmarks/rubrics).
  • Case study context: A local school social work organization exploring LLM support (19 workers across ~30 staff organization).
  • Use case chosen by workers: LLM-generated reflective questions to “challenge” social workers’ assumptions about classroom observations and to broaden teachers’ perspectives — characterized as measuring the LLM’s ability to be a “people challenger” rather than a people-pleaser.
  • Design principles for the workshops: support expressive participation, enable informed iterative design (inspect instrument behavior on realistic data), and foster meaningful collaboration/co-learning among workers.
  • Outcome: Workers jointly created 16 realistic test cases and an LLM-as-a-judge rubric derived from professional and lived expertise. Validation showed strong worker–LLM-judge agreement and the benchmark discriminated among six SOTA LLMs.
  • Contributions: (1) articulation of worker-driven AI measurement, (2) an empirical case study showing feasibility, (3) lessons and opportunities for implementing this approach more broadly.

Data & Methods

  • Participants: 19 school social workers from a local organization (organization employs ~30 social workers across preschool–elementary schools).
  • Timeline: Eight iterative workshops over eight weeks, plus pre-workshop asynchronous prompts and formative interviews with a subset of workers.
  • Three-phase collaborative process:
  • Identify candidate LLM use cases (pre-workshop submissions; workshop discussion) and select a focused use case. Workers produced 39 current and 18 desired LLM uses (from nine workers) during intake.
  • Systematize measurement goals by jointly inspecting model responses and annotating desirable/undesirable properties; surfaced the “people challenger” goal.
  • Operationalize into a benchmark: workers designed 16 realistic test cases and iteratively developed an LLM-as-a-judge rubric to rate how well a model’s outputs would help them reflect and expand perspectives.
  • Model response generation: For rubric development, each test case was paired with multiple model responses generated from different base models and system instructions (drawn from OpenAI, Anthropic, Google) to expose variation.
  • Validation: Compared the LLM-as-a-judge ratings with worker judgments and evaluated six state-of-the-art LLMs on the worker-created benchmark. Reported findings: strong alignment between worker and LLM-judge ratings; benchmark can distinguish relative model performance. (Paper reports “strong agreement” and discrimination across six models; exact numeric reliability/statistics not provided in the excerpt.)
  • Tools: Collaborative workshops (Figma boards for synthesis), model customization via OpenAI ChatGPT customization service for the organization’s internal model; facilitators provided light guidance.

Implications for AI Economics

  • Better alignment of evaluations with worker value: Benchmarks designed by workers capture dimensions of value (e.g., stimulating reflection, surfacing biases) that standard benchmarks miss. This reduces measurement error in estimating the economic benefits of AI augmentation (productivity, quality of service), improving decision-making about investment and deployment.
  • Impacts on assessments of complementarity vs. substitution: Worker-driven metrics emphasize augmentation qualities (e.g., challenging thinking, enabling learning) that signal complementarity with human labor rather than autonomous substitution. Incorporating such metrics into economic estimates can change labor impact projections (e.g., slower displacement, increased skill premiums).
  • Procurement and model choice: Organizations (and purchasers) can use worker-derived benchmarks to select or customize models that better deliver on mission-critical, non-quantitative outcomes. This can channel spending toward models that improve organizational outcomes valued by workers, altering demand dynamics in model markets.
  • ROI and adoption economics: Worker-created measures can reveal benefits not captured by standard productivity metrics (e.g., improved teacher practice, better child outcomes over time). Capturing these benefits can change cost–benefit calculations for adopting LLMs, potentially justifying higher upfront costs for customized or higher-performing models.
  • Incentives for model developers: If buyers (organizations) increasingly require worker-aligned benchmarks, model developers may invest more in capabilities that support human reflection, contextual nuance, and dialogic augmentation—shifting development priorities and market competition.
  • Organizational governance and labor bargaining: Worker-driven measurement can empower labor and professional bodies to define acceptable automation boundaries and evaluation standards, which could be used in procurement clauses, compliance checks, or bargaining over work redesign—affecting wages, task allocation, and monitoring regimes.
  • Policy and regulation: Regulators and funders evaluating social-impact deployments may adopt worker-derived metrics as complementary evidence of utility, safety, or fairness. This could alter subsidy, procurement, or regulatory criteria, influencing which AI products are deployed in public services.
  • Externalities & measurement challenges: Worker-driven benchmarks are context-specific and may be costly to scale. Economists and policymakers should account for heterogeneity in measurement across organizations and the transaction costs of co-design when modeling adoption trajectories and aggregate impacts.
  • Research agenda for AI economics: Incorporate worker-defined outcome measures into empirical studies of LLMs’ labor-market effects (e.g., field experiments, difference-in-differences) to better capture welfare effects beyond narrow productivity proxies. Evaluate how worker-driven evaluations shift observable adoption patterns, wage dynamics, and human capital investments.

Limitations and cautions relevant to economic interpretation: - Single-organization, small sample case study; results may not generalize across occupations, firm sizes, or institutional contexts. - Worker-driven metrics are value-laden and may reflect local norms; aggregation across firms/sectors requires careful harmonization to avoid biased macro estimates. - The labor and time cost of co-design (workshops, rubric development, validation) should be included in economic analyses of adopting such evaluation regimes.

Suggested applications for economists and policymakers: - Use worker-driven benchmarks as complementary outcome measures in empirical evaluations of AI deployments. - Factor the value of augmentation-related outcomes (learning, decision quality, trust) into cost–benefit models for AI procurement in public services. - Study how procurement standards that require worker-aligned evaluation affect market structure and innovation incentives in the AI industry.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic qualitative evidence from an iterative, participatory case study (19 social workers, eight workshops) plus a validation showing strong agreement between worker and LLM-judge ratings and the ability to differentiate among six LLMs; however, evidence is limited to a single organization, small sample, and a narrowly scoped benchmark (16 test cases), limiting claims about broader effectiveness or impact. Methods Rigormedium — Methods are appropriate for exploratory HCI research: structured, iterative workshops, use-case selection, co-created test cases, and a quantitative validation step using multiple LLMs and worker judgments; but the study has limited sample size and scope, potential selection and facilitator biases, and limited external validation or longitudinal assessment. SampleCollaborative case study with 19 frontline school social workers at one Pennsylvania-based school social work organization (organization employs ~30 social workers serving preschool, kindergarten, and elementary schools in low-SES communities). Data come from eight iterative workshops over eight weeks, pre-workshop asynchronous reports (39 current and 18 desired LLM uses from nine workers), 5 formative interviews, co-designed 16 test cases for a targeted use case (generating reflective questions), and a validation comparing worker ratings to an LLM-as-a-judge across responses generated from multiple base models/providers (OpenAI, Anthropic, Google) and six state-of-the-art LLMs. Themeshuman_ai_collab adoption org_design skills_training GeneralizabilitySingle organization and geographic location (Pennsylvania, USA), Small sample of workers (n=19) not representative of broader social work or other occupations, Use case specific to teacher-reflection questions in K-5 school contexts; may not transfer to other tasks or professions, Benchmark built from 16 test cases—limited coverage of real-world variability, Possible facilitator influence and participant selection bias (workers who opted in), Customization of the organization's LLM may limit transferability to off-the-shelf models

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study proposes worker-driven AI measurement, a bottom-up approach in which workers collaboratively decide which tasks AI should augment, what successful augmentation means, and how it should be measured. Governance And Regulation positive Worker participation in defining AI augmentation goals and evaluation criteria
Reading fidelity high
Study strength speculative
n=19
0.03
Across eight workshops, 19 school social workers collaboratively designed an evaluation benchmark for LLM augmentation in their work. Governance And Regulation positive Completion of a worker-designed AI evaluation benchmark
Reading fidelity high
Study strength medium
n=19
8 workshops
0.18
The workers selected generating reflective questions to challenge their thinking and assumptions about classroom observations as a high-priority LLM use case. Task Allocation positive Worker-perceived priority and desirability of an LLM use case
Reading fidelity high
Study strength medium
n=8
0.18
Workers collectively designed 16 realistic test cases intended to challenge LLM capabilities for generating useful reflective questions. Training Effectiveness positive Number and realism of worker-designed benchmark test cases
Reading fidelity high
Study strength medium
n=16
16 test cases
0.18
The workers identified measuring whether LLMs can act as “people challengers”—prompting reflection on users’ assumptions and biases and expanding their perspectives—as an important evaluation goal. Decision Quality positive LLM ability to challenge workers' assumptions, biases, and perspectives
Reading fidelity high
Study strength medium
n=13
0.18
The workers collaboratively developed an LLM-as-a-judge rubric that operationalized their context-specific notion of meaningful AI augmentation. Governance And Regulation positive Contextual validity and operationalization of an AI evaluation rubric
Reading fidelity high
Study strength medium
n=19
0.18
Validation of the worker-created benchmark found strong agreement between worker ratings and ratings produced by the LLM judge. Decision Quality positive Agreement between worker evaluations and LLM-judge evaluations
Reading fidelity high
Study strength low
strong agreement
0.09
The worker-created benchmark differentiated performance across six state-of-the-art LLMs. Output Quality positive Differences in LLM performance on the worker-defined use case
Reading fidelity high
Study strength low
n=6
six state-of-the-art LLMs
0.09
Before the workshops, nine workers reported 39 current and 18 desired uses of LLMs for assisting their work. Adoption Rate positive Breadth of current and desired LLM use cases
Reading fidelity high
Study strength medium
n=9
39 current uses and 18 desired uses
0.18

Notes