The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-assisted scoring of short Spanish essays matches human raters closely in a national exam, and a targeted human-in-the-loop workflow can cut grading effort without materially changing certification outcomes.

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
María Eugenia Curi, Germán Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adrián Silveira, Andrés Peri · September 04, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. María Eugenia Curi unresolved corpus identity
  2. Germán Capdehourat unresolved corpus identity
  3. Isabel Amigo unresolved corpus identity
  4. Magdalena Romano unresolved corpus identity
  5. Rosana Serra unresolved corpus identity
  6. Adrián Silveira unresolved corpus identity
  7. Andrés Peri unresolved corpus identity
LLM-based rubric scoring achieved moderate-to-high agreement with expert human raters on large-scale Spanish national writing tests (2024–2025), and a human-in-the-loop workflow can identify critical cases for review and substantially reduce manual grading workload while preserving pass/fail decisions.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.

Summary

Main Finding

LLM-based, prompt-engineered scoring of short Spanish argumentative texts (≈150–200 words) can achieve moderate-to-high agreement with trained human raters in a national large-scale exam. When integrated into a carefully designed Human-in-the-Loop (HITL) workflow that flags cases likely to affect pass/fail outcomes for human review, AI assistance can meaningfully reduce manual grading effort while preserving decision quality and fairness — but only with sustained human oversight, monitoring for model drift, and safeguards against bias.

Key Points

  • Context: Applied to the Writing section of Uruguay’s national accreditation exam (Acredita EB), a high-stakes test used for certification of lower secondary education.
  • Data: Operational datasets from two consecutive editions (2024 and 2025), each with roughly 5,000–6,000 student responses, enabling cross-year validation.
  • Performance: Prompt-based LLM scoring shows moderate-to-high alignment with expert raters on analytic rubric dimensions (agreement metrics reported at the rubric-item level).
  • Decision focus: Evaluation emphasized decision-oriented impact (e.g., whether AI scores would change pass/fail outcomes) rather than only model-centric metrics.
  • HITL design: Proposed workflow uses AI to pre-score all responses and selectively routes
    • potentially critical or uncertain cases (those likely to change final proficiency/pass–fail) and
    • cases with flagged uncertainty to human experts for review; no final high-stakes decision is made without human oversight.
  • Uncertainty & triage: The study reviews and employs multiple uncertainty measures (semantic and categorical entropy, Max-Agree-Rate, and IRT-based surprises) and evaluates their effectiveness (AUC, C-index) to prioritize human review.
  • Operational implications: Simulation of the HITL pipeline demonstrates realistic workload reductions under deployment assumptions while identifying types of errors that require human attention.
  • Limits and risks: Authors stress need for continuous monitoring (longitudinal model–human alignment), analysis of potential cognitive biases introduced by AI-assisted review, and caution about transferring English-focused findings to Spanish-language contexts.

Data & Methods

  • Data sources:
    • Two real-world national test editions (2024 and 2025), ~5k–6k written responses per edition.
    • Ground truth: human scores produced under the exam’s standard rubric and multi-phase evaluation procedure (calibration, double-scoring, then single-scoring with supervision).
  • Modeling approach:
    • Prompt-engineering of LLMs (no large supervised fine-tuning reported) to produce rubric-aligned analytic scores per response.
    • Prompts included rubric definitions and examples to guide scoring in Spanish.
  • Validation strategy:
    • Rubric-item level comparisons between AI and human raters using agreement measures (accuracy, Cohen’s Kappa / Quadratic Weighted Kappa, inter-rater metrics).
    • Cross-year tests to assess generalization and robustness across test forms (2024 → model development; 2025 → validation).
    • Use of uncertainty/consistency metrics: semantic entropy, categorical entropy, Max-Agree-Rate (MAR), and IRT-based “surprise” between expected and assigned grades.
    • Evaluation of uncertainty metrics via AUC and C-index to quantify how well they discriminate AI errors that need human review.
  • HITL workflow simulation:
    • Scoring pipeline simulated end-to-end to estimate how many responses would be routed to human reviewers based on different triage thresholds and to measure effects on pass/fail outcomes.
    • Emphasis on minimizing the set of reviews needed to preserve final decision integrity rather than maximizing raw AI-human agreement.

Implications for AI Economics

  • Cost and productivity
    • Potential to reduce grading labor costs and shorten turnaround time for high-stakes assessments by automating routine cases and reallocating human graders to review edge/critical cases.
    • Net savings depend on triage thresholds, human oversight intensity, and costs of monitoring/auditing infrastructure; HITL reduces but does not eliminate human labor.
  • Labor market effects
    • Partial substitution of routine grading tasks could shift demand toward fewer but more expert supervisory roles (training, calibration, auditing), altering skill premiums in assessment teams.
    • Reallocation may increase productivity of assessment systems, enabling scaling or re-investment in pedagogy/feedback services.
  • Quality, fairness, and welfare risks
    • Errors that systematically under-grade or over-grade certain groups can have distributional welfare consequences (e.g., unequal access to certification, labor market entry). Decision-oriented triage reduces but does not remove these risks.
    • Cognitive biases introduced by AI-supported review workflows (e.g., reviewers deferring to AI or being anchored by model outputs) may erode human adjudication quality; these effects require explicit study and mitigation (blinding, randomized audits).
  • Adoption incentives and regulation
    • Governments and exam boards face trade-offs between faster/cheaper certification and legal/ethical obligations to ensure fairness and appealability. The decision-oriented HITL approach aligns better with regulatory constraints because it guarantees human sign-off on high-stakes outcomes.
    • Demonstrated cross-year robustness in a non-English setting (Spanish) lowers one barrier to adoption in other low-resource-language contexts but does not obviate the need for local validation.
  • Market-level effects
    • Faster certification can change labor supply timing (earlier entry), potentially affecting wages and credential signaling; broader adoption might increase credential supply, with implications for qualification inflation.
    • Cost savings for public assessment systems could free funds for other educational investments, but savings estimates must incorporate ongoing monitoring, audit, retraining, and model update costs.
  • Research and policy priorities for economists
    • Quantify net economic benefits: rigorous cost–benefit analyses that include human oversight, monitoring infrastructure, error costs (false passes/false fails), and distributional impacts.
    • Study behavioral effects on human graders (anchoring/automation bias) and design incentives/controls to preserve adjudication quality.
    • Model long-run labor reallocation and credential market responses to scaled AI-assisted assessment.
    • Develop metrics and regulation standards (auditability, appeals, transparency) tailored to high-stakes educational markets.

Summary conclusion: The paper provides pragmatic, decision-focused evidence that LLMs can materially assist large-scale, Spanish-language writing assessment when embedded in a conservative HITL framework. From an AI economics perspective, this enables meaningful productivity gains and reconfiguration of grading labor, but net economic gains hinge on careful design, ongoing monitoring, and policies that address fairness, accountability, and labor transitions.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Uses large-scale operational data (two national-test cohorts of ~5k–6k responses) and cross-year validation to show alignment between LLM outputs and human raters and to simulate an operational HITL workflow; however, there is no experimental/randomized assignment, limited detail (in the provided text) on model selection/training, potential selection and rubric drift concerns, and no independent causal test of downstream outcomes. Methods Rigormedium — Design leverages real-world administrative data, rubric-based human scores, double-scoring calibration samples, and cross-year validation—strengths for operational validity—but lacks randomized comparison, limited transparency in modeling/prompting details and error analysis in the excerpt, and potential unaddressed risks (bias, adversarial prompts, long-term model drift). SampleOperational data from Uruguay's national Acredita EB test (writing section) for two recent editions (2024 and 2025), each with approximately 5,000–6,000 short argumentative student responses (~150–200 words) written in Spanish; responses were evaluated with an expert-designed analytic rubric, including calibration and double-scoring subsamples used for inter-rater reliability and model validation. Themesproductivity human_ai_collab GeneralizabilityFindings are specific to short (150–200 word) Spanish argumentative essays and may not generalize to longer essays, other genres, or other languages., Results reflect the particular rubric, rater training, and operational procedures of Uruguay's Acredita EB exam and may not transfer to different scoring rubrics or assessment cultures., Performance depends on the specific LLM(s), prompt designs, and model versions used; results may change as models are updated., High-stakes testing contexts with different stakes or appeals processes may reveal different error costs (e.g., equity concerns) not fully captured here., Simulated workload reductions may differ from effects in fully deployed systems (operational integration, adversarial inputs, or student appeals could alter outcomes).

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study reports moderate to high agreement between AI-generated scores and human evaluations across most rubric dimensions. Decision Quality positive Agreement between AI-generated rubric scores and expert human scores
Reading fidelity high
Study strength medium
n=5000
0.18
The proposed human-in-the-loop correction workflow identifies cases in which human review is most valuable, enabling more efficient allocation of expert effort. Organizational Efficiency positive Allocation of expert human review effort
Reading fidelity high
Study strength medium
n=5000
0.18
The paper concludes that AI-assisted scoring can be integrated into large-scale assessment processes only when combined with carefully designed human oversight. Governance And Regulation positive Feasibility and governance of AI-assisted scoring in high-stakes assessment
Reading fidelity high
Study strength medium
n=5000
0.18
The AI-assisted scoring framework is intended to reduce human scoring workload while preserving the quality of assessment decisions. Organizational Efficiency positive Human scoring workload and assessment decision quality
Reading fidelity high
Study strength low
n=5000
0.09
The study uses operational data from two recent editions of Uruguay's national assessment, with approximately 5,000–6,000 participants in each edition. Other null_result Scope of the assessment dataset
Reading fidelity high
Study strength high
n=5000
approximately 5,000–6,000 participants per edition
0.3
The 2024 assessment data were used to develop the AI-based scoring model, while the 2025 data were used for validation. Decision Quality positive Out-of-year validation of the AI scoring model
Reading fidelity high
Study strength high
n=5000
0.3
The Writing section is the most time- and resource-intensive component of the assessment because written responses are evaluated manually using a detailed rubric. Task Completion Time negative Human time and resources required for writing-section scoring
Reading fidelity high
Study strength high
not reported
0.3
Applicants typically receive their final assessment results no sooner than three months after the examination. Task Completion Time negative Time from examination to final result reporting
Reading fidelity high
Study strength medium
no sooner than three months
0.18
The proposed operational principle is that no final decision affecting test outcomes should be made without appropriate human oversight. Governance And Regulation positive Human oversight of consequential pass/fail assessment decisions
Reading fidelity high
Study strength low
not reported
0.09

Notes