The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI 'executives' steer group chats to surface more evidence of teamwork and project skills, and AI autoraters score those conversations on par with human experts; the approach promises scalable assessment of creativity, critical thinking and collaboration but remains validated only in a limited sample and in-protocol outcomes.

Towards Scalable Measurement of Durable Skills
Amir Globerson, Amy Keeling, Anisha Choudhury, Anna Iurchenko, Aviad Segal, Avinatan Hassidim, Ayça Çakmakli, Ben Gomes, Benn Witt, Cathy Cheunga, Cristine Legare, Diana Akrong, Eliad Carmi, Elisabeth Bauer, Gal Elidan, Hadas Gelbart, Hairong Mu, Katherine Chou, Lev Borovoi, Nir Kerem, Niv Efron, Noa Kerrem Gilo, Preeti Singh, Rajvi Kapadia, Rena Levitt, Roni Rabin, Ronit Levavi Morad, Rotem Yulzary, Shashank Agarwal, Sophie Allweis, Tracey Lee-Joe, Tzvika Stein, Yael Bar Moshe, Yael Haramaty, Yaniv Carmel, Yishay Mor, Yoav Bar Sinai, Yoav Bergner, Yossi Matias, Yuri Lev · September 14, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Amir Globerson unresolved corpus identity
  2. Amy Keeling unresolved corpus identity
  3. Anisha Choudhury unresolved corpus identity
  4. Anna Iurchenko unresolved corpus identity
  5. Aviad Segal unresolved corpus identity
  6. Avinatan Hassidim unresolved corpus identity
  7. Ayça Çakmakli unresolved corpus identity
  8. Ben Gomes unresolved corpus identity
  9. Benn Witt unresolved corpus identity
  10. Cathy Cheunga unresolved corpus identity
  11. Cristine Legare unresolved corpus identity
  12. Diana Akrong unresolved corpus identity
  13. Eliad Carmi unresolved corpus identity
  14. Elisabeth Bauer unresolved corpus identity
  15. Gal Elidan unresolved corpus identity
  16. Hadas Gelbart unresolved corpus identity
  17. Hairong Mu unresolved corpus identity
  18. Katherine Chou unresolved corpus identity
  19. Lev Borovoi unresolved corpus identity
  20. Nir Kerem unresolved corpus identity
  21. Niv Efron unresolved corpus identity
  22. Noa Kerrem Gilo unresolved corpus identity
  23. Preeti Singh unresolved corpus identity
  24. Rajvi Kapadia unresolved corpus identity
  25. Rena Levitt unresolved corpus identity
  26. Roni Rabin unresolved corpus identity
  27. Ronit Levavi Morad unresolved corpus identity
  28. Rotem Yulzary unresolved corpus identity
  29. Shashank Agarwal unresolved corpus identity
  30. Sophie Allweis unresolved corpus identity
  31. Tracey Lee-Joe unresolved corpus identity
  32. Tzvika Stein unresolved corpus identity
  33. Yael Bar Moshe unresolved corpus identity
  34. Yael Haramaty unresolved corpus identity
  35. Yaniv Carmel unresolved corpus identity
  36. Yishay Mor unresolved corpus identity
  37. Yoav Bar Sinai unresolved corpus identity
  38. Yoav Bergner unresolved corpus identity
  39. Yossi Matias unresolved corpus identity
  40. Yuri Lev unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Amir Globerson provider ID
  2. A. Keeling provider ID
  3. Anisha Choudhury provider ID
  4. A. Iurchenko provider ID
  5. Aviad Segal provider ID
  6. Avinatan Hassidim provider ID
  7. Ayça Çakmakli provider ID
  8. Ben Gomes provider ID
  9. B. Witt provider ID
  10. Cathy Cheung provider ID
  11. C. Legare provider ID
  12. Diana Akrong provider ID
  13. Eliad Carmi provider ID
  14. Elisabeth Bauer provider ID
  15. G. Elidan provider ID
  16. Hadas Gelbart provider ID
  17. Hairong Mu provider ID
  18. Katherine Chou provider ID
  19. Lev Borovoi provider ID
  20. Nir Kerem provider ID
  21. Niv Efron provider ID
  22. Noa Kerrem Gilo provider ID
  23. Preeti Singh provider ID
  24. Rajvi Kapadia provider ID
  25. Rena Levitt provider ID
  26. Ron Rabin provider ID
  27. Ronit Levavi Morad provider ID
  28. Rotem Yulzary provider ID
  29. Shashank Agarwal provider ID
  30. Sophie Allweis provider ID
  31. T. Lee-Joe provider ID
  32. T. Stein provider ID
  33. Y. Bar provider ID
  34. Y. Haramaty provider ID
  35. Yaniv Carmel provider ID
  36. Yishay Mor provider ID
  37. Yoav Bar Sinai provider ID
  38. Yoav Bergner provider ID
  39. Y. Matias provider ID
  40. Yuri Lev provider ID
Orchestrated LLM teammates (an Executive LLM) steer group conversations to elicit significantly more evidence of durable skills, and an LLM-based evaluator rates transcripts at similar agreement levels to expert human raters in a randomized experiment with 188 participants.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.

Summary

Main Finding

LLMs can be orchestrated to create scalable, controllable, and ecologically valid assessments of durable skills (collaboration, creativity, critical thinking). An “Executive LLM” that jointly generates AI-teammate behavior and steers conversation toward targeted behaviors increases the amount of observable evidence for a given skill; an LLM-based autorater attains inter-rater agreement with human expert raters at similar levels to human-human agreement. The paper presents Vantage, a full protocol and implementation demonstrating these claims on human and simulated conversations.

Key Points

  • Motivation: Durable (future-ready) skills are important but hard to measure; measurement gaps reduce instruction and credentialing. The paper targets the trade-off between ecological validity (naturalistic interaction) and psychometric control (reliability, reproducibility).
  • Executive LLM concept: A single LLM generates utterances for multiple AI teammates and is prompted with pedagogical rubrics to steer the interaction toward eliciting informative evidence about specific rubric dimensions (e.g., conflict resolution, project management).
  • AI Evaluator (autorater): An LLM grades turn-level evidence and aggregates to conversation-level ratings. Turn-level ratings are produced 20 times per turn; NA-handling and mode aggregation are used, followed by a regression model to map turn labels to conversation scores.
  • Empirical findings:
    • Executive LLMs focused on a skill elicit significantly more turn- and conversation-level evidence for that skill than uncoordinated “Independent Agents.”
    • Steering has a cross-effect: focusing on one skill increases evidence for that skill and can reduce evidence for others (desirable discriminant behavior).
    • LLM autorating achieves agreement with human experts comparable to inter-expert agreement (human-human Cohen’s Kappa ~0.45–0.64; LLM-human in similar range).
    • Simulations (LLM-as-subject) can be used to test recovery of set skill levels and to develop the protocol before costly human data collection.
    • For a real-world creativity task (high-school students), a Gemini-based autorater matched expert human ratings.
  • Practical UI & feedback: Vantage produces quantitative maps and qualitative conversation excerpts to support user feedback and formative use.
  • Limitations noted: coding of complex interaction dimensions remained challenging for human raters (moderate inter-rater agreement), so rubric design and rater calibration are critical.

Data & Methods

  • Participants & data:
    • Human study: 188 U.S.-based native English speakers (ages 18–25) recruited on Prolific.
    • Each participant produced two conversations; total N = 373 conversations (30-minute chat interactions).
  • Experimental variations:
    • Three AI teammate strategies: Executive LLM focused on Conflict Resolution, Executive LLM focused on Project Management, and Independent Agents (no steering).
    • Randomized instruction to participants: no focus vs. instruction to pay attention to one skill.
  • Models:
    • Executive LLM (content generation): Gemini 2.5 Pro used for main collaboration experiments; Gemini 3 used in some creativity/critical thinking experiments.
    • AI Evaluator (turn-level autorater): Gemini 3.0 used to generate turn-level ratings.
  • Autorater procedure:
    • For each turn, the evaluator generated 20 independent ratings.
    • If any of the 20 was NA, final turn label = NA; otherwise final turn score = mode of the 20.
    • Conversation-level aggregation: linear regression for numeric scores and logistic regression for NA presence; performance assessed with leave-one-out cross-validation.
  • Human annotation:
    • Two pedagogical raters annotated turns and holistic conversation scores using 1–4 scales plus NA for rubric dimensions.
    • Used as gold/comparison for inter-rater consistencies and for validating autorater outputs.
  • Simulations:
    • Gemini prompted to act as a student at specified rubric levels; used for recovery tests (50-turn conversations, 100 repetitions per level) and for protocol development.

Implications for AI Economics

  • Lower-cost, scalable measurement of durable skills
    • Measurement-as-a-service: Vantage-style autoraters can dramatically reduce the marginal cost of scoring complex social-cognitive tasks, enabling larger samples and longitudinal data.
    • Economists can obtain richer, standardized measures of non-cognitive skills at scale for labor-market, education, and training studies.
  • Improved identification of returns to durable skills
    • With scalable, repeatable measures, research designs (RCTs, instrumental variables, difference-in-differences) can better estimate causal returns of collaboration/creativity/critical thinking to wages, productivity, job matches, and career trajectories.
    • Enables linking of granular skill profiles to firm-level performance, occupation task content, and complementarity with technical skills or AI tools.
  • Market and policy applications
    • Credentialing & signaling: automated, rubric-driven assessments could form components of micro-credentials or hiring signals; this may alter labor-market sorting and wages for tasks requiring durable skills.
    • Targeting upskilling investments: public and private training programs can be evaluated and adapted at scale using such measures, improving allocative efficiency of training funds.
  • Potential for new data infrastructures
    • Large panel datasets combining Vantage-like assessments with administrative outcomes (earnings, employment records) enable richer models of human capital accumulation and depreciation, measurement error correction, and heterogeneity analysis.
  • Risks, caveats, and research priorities
    • Measurement validity & predictive validity: economists should validate these LLM-based measures against long-run outcomes (earnings, promotions, productivity) before using them as dependent variables or policy levers.
    • Fairness & measurement invariance: assessments must be tested for differential functioning across demographic groups, languages, and cultural contexts to avoid biased estimates and harmful redistribution.
    • Strategic behavior & gaming: widespread use in high-stakes settings (hiring, admissions) creates incentives to game interactions; researchers should study robustness, adversarial manipulation, and design anti-gaming protocols (e.g., randomized tasks, multiple modalities).
    • Model drift & transparency: reliance on opaque LLMs raises reproducibility/regulatory issues. Economists should document model versions, prompts, and calibration data; implement monitoring for drift over time.
    • Privacy and consent: conversational data is sensitive; economic uses (linking to admin data) must respect privacy, informed consent, and legal constraints.
  • Concrete suggestions for economists wanting to use this approach
    • Pilot linking: run pilots linking Vantage-style measures to short-term productivity proxies (grades, supervisor ratings) and to administrative outcomes to establish predictive validity.
    • Use as outcome in RCTs: incorporate Vantage measures as intermediate outcomes in training or pedagogy RCTs to improve statistical power and diagnostic value.
    • Measurement error modeling: incorporate autorater uncertainty (e.g., turn-level NA probabilities, repeat-rating variance) in structural models or IV designs.
    • Equity audits: pre-register and conduct subgroup analyses and measurement invariance tests before policy deployment.
    • Cost-benefit analyses: compare costs of Vantage-style assessments with alternative measurement strategies to evaluate scalability and welfare impacts.

Overall, the paper provides a practical, empirically validated blueprint for using orchestrated LLMs to measure complex durable skills at scale. For AI economists, it opens opportunities to study human capital dynamics, policy interventions, and labor-market matching with richer behavioral measures — provided careful validation, fairness checks, and governance mechanisms are in place.

Assessment

Paper Typerct Evidence Strengthmedium — The paper uses randomized assignment to agent protocols and benchmarks automated scoring against human expert ratings, providing credible internal identification for the main claims (Executive LLM increases elicited evidence; LLM evaluator approximates human ratings). However, outcomes are proxied (fraction of turns/conversations coded as evidence and inter-rater agreement) rather than downstream, real-world measures of skill transfer or labor-market outcomes. Sample is moderate and restricted (Prolific, ages 18–25, US English), and some validation relies on the same model family (Gemini) for generation and evaluation, which limits external validity. Methods Rigormedium — Strengths: randomized experimental design, pre-specified rubrics, human rater calibration, comparison of LLM autorater to human experts, use of leave-one-out cross-validation and simulation recovery tests. Limitations: moderate sample and narrow participant population, potential circularity because LLMs are used both to steer conversations and to evaluate them (though different variants/models used in parts), limited reporting of predictive/construct validity beyond agreement statistics, and no pre-registration or long-term outcome validation reported in the supplied text. Sample188 human participants recruited on Prolific, ages 18–25, native English speakers based in the United States; each produced two conversations (total 373 conversations after filtering). Conversations were 30 minutes in a chat-like interface (text and optional voice). Executive LLM experiments used Gemini 2.5 Pro; additional creativity/critical thinking analyses used Gemini 3; human expert ratings were provided by two pedagogical raters from NYU. Simulated-subject experiments used Gemini to produce 50-turn conversations with 100 repetitions per rubric level. An external autorater evaluation included high-school student creativity tasks (details/N not fully specified in supplied text). Themesskills_training human_ai_collab IdentificationBetween-subjects random assignment of participants (N=188) to agent protocols: Executive LLM focused on Conflict Resolution, Executive LLM focused on Project Management, or Independent Agents; additional randomization of explicit pre-task focus instruction (none vs skill-specific). Causal inference about the effect of the Executive LLM on elicited evidence is thus identified by randomization; validity of automated scoring is established by comparing LLM autorater outputs to human expert ratings (Cohen's Kappa) and using leave-one-out cross-validation for the turn-to-conversation aggregation models. Simulated-agent recovery tests (LLM-simulated participants with known skill levels) are used for additional validation. GeneralizabilityRestricted age range (18–25) — results may not generalize to younger students or older adults., US-based, native English Prolific sample — limited cross-cultural and language generalizability., Laboratory-style, short (30-minute) chat tasks — may not reflect longer-term classroom/group dynamics., Use of Gemini-family models (proprietary) — results may depend on specific model capabilities and prompting and may not transfer to other LLMs or future model versions., Evaluation focuses on elicited evidence and inter-rater agreement rather than predictive validity for real-world outcomes (e.g., classroom performance, employment), limiting external validity., Potential circularity: some validation uses models related to those used for generation (risk of overfitting to model behavior).

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Agreement between the LLM AI Evaluator and human experts was similar to the agreement between two human expert raters when scoring conflict-resolution and project-management conversations. Decision Quality positive Agreement between automated and human ratings of collaboration-skill evidence and scores
Reading fidelity high
Study strength medium
n=188
Human inter-rater Cohen's Kappa 0.45-0.64; LLM–expert agreement described as similar
0.6
An Executive LLM focused on a target collaboration skill elicited significantly more evidence of that skill than the Independent Agents protocol. Decision Quality positive Fraction of participant turns and conversations containing evidence of conflict resolution or project management
Reading fidelity high
Study strength medium
n=188
Statistically significant difference, p≤0.05
0.6
Steering the conversation toward conflict resolution increased conflict-resolution evidence but reduced project-management evidence, while steering toward project management produced the opposite pattern. Task Allocation mixed Skill-specific evidence elicited in participant turns and conversations
Reading fidelity high
Study strength medium
n=188
Significant difference in three of four comparisons
0.6
Human raters found coding conversations for conflict resolution and project management challenging, even after multiple rounds of calibration. Decision Quality negative Inter-rater agreement and consistency in coding collaboration-skill evidence
Reading fidelity high
Study strength medium
n=188
Cohen's Kappa 0.45-0.64
0.6
The proposed Vantage protocol provides a scalable and controllable approach for assessing complex durable skills through human interactions with AI teammates and automated transcript evaluation. Organizational Efficiency positive Scalability and controllability of durable-skill assessment
Reading fidelity high
Study strength medium
n=188
373 conversations generated; three conversations filtered due to technical issues
0.6
For complex creativity tasks completed by high-school students, a Gemini-based AI Evaluator was reported to perform effectively as a creativity assessor and to be on par with human expert raters. Creativity positive Agreement or effectiveness of automated creativity assessment relative to expert ratings
Reading fidelity high
Study strength low
On par with human expert raters
0.3

Notes