0 cumulative citations
View corpus contextAI 'executives' steer group chats to surface more evidence of teamwork and project skills, and AI autoraters score those conversations on par with human experts; the approach promises scalable assessment of creativity, critical thinking and collaboration but remains validated only in a limited sample and in-protocol outcomes.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.
Summary
Main Finding
LLMs can be orchestrated to create scalable, controllable, and ecologically valid assessments of durable skills (collaboration, creativity, critical thinking). An “Executive LLM” that jointly generates AI-teammate behavior and steers conversation toward targeted behaviors increases the amount of observable evidence for a given skill; an LLM-based autorater attains inter-rater agreement with human expert raters at similar levels to human-human agreement. The paper presents Vantage, a full protocol and implementation demonstrating these claims on human and simulated conversations.
Key Points
- Motivation: Durable (future-ready) skills are important but hard to measure; measurement gaps reduce instruction and credentialing. The paper targets the trade-off between ecological validity (naturalistic interaction) and psychometric control (reliability, reproducibility).
- Executive LLM concept: A single LLM generates utterances for multiple AI teammates and is prompted with pedagogical rubrics to steer the interaction toward eliciting informative evidence about specific rubric dimensions (e.g., conflict resolution, project management).
- AI Evaluator (autorater): An LLM grades turn-level evidence and aggregates to conversation-level ratings. Turn-level ratings are produced 20 times per turn; NA-handling and mode aggregation are used, followed by a regression model to map turn labels to conversation scores.
- Empirical findings:
- Executive LLMs focused on a skill elicit significantly more turn- and conversation-level evidence for that skill than uncoordinated “Independent Agents.”
- Steering has a cross-effect: focusing on one skill increases evidence for that skill and can reduce evidence for others (desirable discriminant behavior).
- LLM autorating achieves agreement with human experts comparable to inter-expert agreement (human-human Cohen’s Kappa ~0.45–0.64; LLM-human in similar range).
- Simulations (LLM-as-subject) can be used to test recovery of set skill levels and to develop the protocol before costly human data collection.
- For a real-world creativity task (high-school students), a Gemini-based autorater matched expert human ratings.
- Practical UI & feedback: Vantage produces quantitative maps and qualitative conversation excerpts to support user feedback and formative use.
- Limitations noted: coding of complex interaction dimensions remained challenging for human raters (moderate inter-rater agreement), so rubric design and rater calibration are critical.
Data & Methods
- Participants & data:
- Human study: 188 U.S.-based native English speakers (ages 18–25) recruited on Prolific.
- Each participant produced two conversations; total N = 373 conversations (30-minute chat interactions).
- Experimental variations:
- Three AI teammate strategies: Executive LLM focused on Conflict Resolution, Executive LLM focused on Project Management, and Independent Agents (no steering).
- Randomized instruction to participants: no focus vs. instruction to pay attention to one skill.
- Models:
- Executive LLM (content generation): Gemini 2.5 Pro used for main collaboration experiments; Gemini 3 used in some creativity/critical thinking experiments.
- AI Evaluator (turn-level autorater): Gemini 3.0 used to generate turn-level ratings.
- Autorater procedure:
- For each turn, the evaluator generated 20 independent ratings.
- If any of the 20 was NA, final turn label = NA; otherwise final turn score = mode of the 20.
- Conversation-level aggregation: linear regression for numeric scores and logistic regression for NA presence; performance assessed with leave-one-out cross-validation.
- Human annotation:
- Two pedagogical raters annotated turns and holistic conversation scores using 1–4 scales plus NA for rubric dimensions.
- Used as gold/comparison for inter-rater consistencies and for validating autorater outputs.
- Simulations:
- Gemini prompted to act as a student at specified rubric levels; used for recovery tests (50-turn conversations, 100 repetitions per level) and for protocol development.
Implications for AI Economics
- Lower-cost, scalable measurement of durable skills
- Measurement-as-a-service: Vantage-style autoraters can dramatically reduce the marginal cost of scoring complex social-cognitive tasks, enabling larger samples and longitudinal data.
- Economists can obtain richer, standardized measures of non-cognitive skills at scale for labor-market, education, and training studies.
- Improved identification of returns to durable skills
- With scalable, repeatable measures, research designs (RCTs, instrumental variables, difference-in-differences) can better estimate causal returns of collaboration/creativity/critical thinking to wages, productivity, job matches, and career trajectories.
- Enables linking of granular skill profiles to firm-level performance, occupation task content, and complementarity with technical skills or AI tools.
- Market and policy applications
- Credentialing & signaling: automated, rubric-driven assessments could form components of micro-credentials or hiring signals; this may alter labor-market sorting and wages for tasks requiring durable skills.
- Targeting upskilling investments: public and private training programs can be evaluated and adapted at scale using such measures, improving allocative efficiency of training funds.
- Potential for new data infrastructures
- Large panel datasets combining Vantage-like assessments with administrative outcomes (earnings, employment records) enable richer models of human capital accumulation and depreciation, measurement error correction, and heterogeneity analysis.
- Risks, caveats, and research priorities
- Measurement validity & predictive validity: economists should validate these LLM-based measures against long-run outcomes (earnings, promotions, productivity) before using them as dependent variables or policy levers.
- Fairness & measurement invariance: assessments must be tested for differential functioning across demographic groups, languages, and cultural contexts to avoid biased estimates and harmful redistribution.
- Strategic behavior & gaming: widespread use in high-stakes settings (hiring, admissions) creates incentives to game interactions; researchers should study robustness, adversarial manipulation, and design anti-gaming protocols (e.g., randomized tasks, multiple modalities).
- Model drift & transparency: reliance on opaque LLMs raises reproducibility/regulatory issues. Economists should document model versions, prompts, and calibration data; implement monitoring for drift over time.
- Privacy and consent: conversational data is sensitive; economic uses (linking to admin data) must respect privacy, informed consent, and legal constraints.
- Concrete suggestions for economists wanting to use this approach
- Pilot linking: run pilots linking Vantage-style measures to short-term productivity proxies (grades, supervisor ratings) and to administrative outcomes to establish predictive validity.
- Use as outcome in RCTs: incorporate Vantage measures as intermediate outcomes in training or pedagogy RCTs to improve statistical power and diagnostic value.
- Measurement error modeling: incorporate autorater uncertainty (e.g., turn-level NA probabilities, repeat-rating variance) in structural models or IV designs.
- Equity audits: pre-register and conduct subgroup analyses and measurement invariance tests before policy deployment.
- Cost-benefit analyses: compare costs of Vantage-style assessments with alternative measurement strategies to evaluate scalability and welfare impacts.
Overall, the paper provides a practical, empirically validated blueprint for using orchestrated LLMs to measure complex durable skills at scale. For AI economists, it opens opportunities to study human capital dynamics, policy interventions, and labor-market matching with richer behavioral measures — provided careful validation, fairness checks, and governance mechanisms are in place.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Agreement between the LLM AI Evaluator and human experts was similar to the agreement between two human expert raters when scoring conflict-resolution and project-management conversations. Decision Quality | positive | Agreement between automated and human ratings of collaboration-skill evidence and scores |
Reading fidelity
high
Study strength
medium
|
n=188
Human inter-rater Cohen's Kappa 0.45-0.64; LLM–expert agreement described as similar
|
| An Executive LLM focused on a target collaboration skill elicited significantly more evidence of that skill than the Independent Agents protocol. Decision Quality | positive | Fraction of participant turns and conversations containing evidence of conflict resolution or project management |
Reading fidelity
high
Study strength
medium
|
n=188
Statistically significant difference, p≤0.05
|
| Steering the conversation toward conflict resolution increased conflict-resolution evidence but reduced project-management evidence, while steering toward project management produced the opposite pattern. Task Allocation | mixed | Skill-specific evidence elicited in participant turns and conversations |
Reading fidelity
high
Study strength
medium
|
n=188
Significant difference in three of four comparisons
|
| Human raters found coding conversations for conflict resolution and project management challenging, even after multiple rounds of calibration. Decision Quality | negative | Inter-rater agreement and consistency in coding collaboration-skill evidence |
Reading fidelity
high
Study strength
medium
|
n=188
Cohen's Kappa 0.45-0.64
|
| The proposed Vantage protocol provides a scalable and controllable approach for assessing complex durable skills through human interactions with AI teammates and automated transcript evaluation. Organizational Efficiency | positive | Scalability and controllability of durable-skill assessment |
Reading fidelity
high
Study strength
medium
|
n=188
373 conversations generated; three conversations filtered due to technical issues
|
| For complex creativity tasks completed by high-school students, a Gemini-based AI Evaluator was reported to perform effectively as a creativity assessor and to be on par with human expert raters. Creativity | positive | Agreement or effectiveness of automated creativity assessment relative to expert ratings |
Reading fidelity
high
Study strength
low
|
On par with human expert raters
|