0 cumulative citations
View corpus contextAI-assisted scoring of short Spanish essays matches human raters closely in a national exam, and a targeted human-in-the-loop workflow can cut grading effort without materially changing certification outcomes.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.
Summary
Main Finding
LLM-based, prompt-engineered scoring of short Spanish argumentative texts (≈150–200 words) can achieve moderate-to-high agreement with trained human raters in a national large-scale exam. When integrated into a carefully designed Human-in-the-Loop (HITL) workflow that flags cases likely to affect pass/fail outcomes for human review, AI assistance can meaningfully reduce manual grading effort while preserving decision quality and fairness — but only with sustained human oversight, monitoring for model drift, and safeguards against bias.
Key Points
- Context: Applied to the Writing section of Uruguay’s national accreditation exam (Acredita EB), a high-stakes test used for certification of lower secondary education.
- Data: Operational datasets from two consecutive editions (2024 and 2025), each with roughly 5,000–6,000 student responses, enabling cross-year validation.
- Performance: Prompt-based LLM scoring shows moderate-to-high alignment with expert raters on analytic rubric dimensions (agreement metrics reported at the rubric-item level).
- Decision focus: Evaluation emphasized decision-oriented impact (e.g., whether AI scores would change pass/fail outcomes) rather than only model-centric metrics.
- HITL design: Proposed workflow uses AI to pre-score all responses and selectively routes
- potentially critical or uncertain cases (those likely to change final proficiency/pass–fail) and
- cases with flagged uncertainty to human experts for review; no final high-stakes decision is made without human oversight.
- Uncertainty & triage: The study reviews and employs multiple uncertainty measures (semantic and categorical entropy, Max-Agree-Rate, and IRT-based surprises) and evaluates their effectiveness (AUC, C-index) to prioritize human review.
- Operational implications: Simulation of the HITL pipeline demonstrates realistic workload reductions under deployment assumptions while identifying types of errors that require human attention.
- Limits and risks: Authors stress need for continuous monitoring (longitudinal model–human alignment), analysis of potential cognitive biases introduced by AI-assisted review, and caution about transferring English-focused findings to Spanish-language contexts.
Data & Methods
- Data sources:
- Two real-world national test editions (2024 and 2025), ~5k–6k written responses per edition.
- Ground truth: human scores produced under the exam’s standard rubric and multi-phase evaluation procedure (calibration, double-scoring, then single-scoring with supervision).
- Modeling approach:
- Prompt-engineering of LLMs (no large supervised fine-tuning reported) to produce rubric-aligned analytic scores per response.
- Prompts included rubric definitions and examples to guide scoring in Spanish.
- Validation strategy:
- Rubric-item level comparisons between AI and human raters using agreement measures (accuracy, Cohen’s Kappa / Quadratic Weighted Kappa, inter-rater metrics).
- Cross-year tests to assess generalization and robustness across test forms (2024 → model development; 2025 → validation).
- Use of uncertainty/consistency metrics: semantic entropy, categorical entropy, Max-Agree-Rate (MAR), and IRT-based “surprise” between expected and assigned grades.
- Evaluation of uncertainty metrics via AUC and C-index to quantify how well they discriminate AI errors that need human review.
- HITL workflow simulation:
- Scoring pipeline simulated end-to-end to estimate how many responses would be routed to human reviewers based on different triage thresholds and to measure effects on pass/fail outcomes.
- Emphasis on minimizing the set of reviews needed to preserve final decision integrity rather than maximizing raw AI-human agreement.
Implications for AI Economics
- Cost and productivity
- Potential to reduce grading labor costs and shorten turnaround time for high-stakes assessments by automating routine cases and reallocating human graders to review edge/critical cases.
- Net savings depend on triage thresholds, human oversight intensity, and costs of monitoring/auditing infrastructure; HITL reduces but does not eliminate human labor.
- Labor market effects
- Partial substitution of routine grading tasks could shift demand toward fewer but more expert supervisory roles (training, calibration, auditing), altering skill premiums in assessment teams.
- Reallocation may increase productivity of assessment systems, enabling scaling or re-investment in pedagogy/feedback services.
- Quality, fairness, and welfare risks
- Errors that systematically under-grade or over-grade certain groups can have distributional welfare consequences (e.g., unequal access to certification, labor market entry). Decision-oriented triage reduces but does not remove these risks.
- Cognitive biases introduced by AI-supported review workflows (e.g., reviewers deferring to AI or being anchored by model outputs) may erode human adjudication quality; these effects require explicit study and mitigation (blinding, randomized audits).
- Adoption incentives and regulation
- Governments and exam boards face trade-offs between faster/cheaper certification and legal/ethical obligations to ensure fairness and appealability. The decision-oriented HITL approach aligns better with regulatory constraints because it guarantees human sign-off on high-stakes outcomes.
- Demonstrated cross-year robustness in a non-English setting (Spanish) lowers one barrier to adoption in other low-resource-language contexts but does not obviate the need for local validation.
- Market-level effects
- Faster certification can change labor supply timing (earlier entry), potentially affecting wages and credential signaling; broader adoption might increase credential supply, with implications for qualification inflation.
- Cost savings for public assessment systems could free funds for other educational investments, but savings estimates must incorporate ongoing monitoring, audit, retraining, and model update costs.
- Research and policy priorities for economists
- Quantify net economic benefits: rigorous cost–benefit analyses that include human oversight, monitoring infrastructure, error costs (false passes/false fails), and distributional impacts.
- Study behavioral effects on human graders (anchoring/automation bias) and design incentives/controls to preserve adjudication quality.
- Model long-run labor reallocation and credential market responses to scaled AI-assisted assessment.
- Develop metrics and regulation standards (auditability, appeals, transparency) tailored to high-stakes educational markets.
Summary conclusion: The paper provides pragmatic, decision-focused evidence that LLMs can materially assist large-scale, Spanish-language writing assessment when embedded in a conservative HITL framework. From an AI economics perspective, this enables meaningful productivity gains and reconfiguration of grading labor, but net economic gains hinge on careful design, ongoing monitoring, and policies that address fairness, accountability, and labor transitions.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study reports moderate to high agreement between AI-generated scores and human evaluations across most rubric dimensions. Decision Quality | positive | Agreement between AI-generated rubric scores and expert human scores |
Reading fidelity
high
Study strength
medium
|
n=5000
|
| The proposed human-in-the-loop correction workflow identifies cases in which human review is most valuable, enabling more efficient allocation of expert effort. Organizational Efficiency | positive | Allocation of expert human review effort |
Reading fidelity
high
Study strength
medium
|
n=5000
|
| The paper concludes that AI-assisted scoring can be integrated into large-scale assessment processes only when combined with carefully designed human oversight. Governance And Regulation | positive | Feasibility and governance of AI-assisted scoring in high-stakes assessment |
Reading fidelity
high
Study strength
medium
|
n=5000
|
| The AI-assisted scoring framework is intended to reduce human scoring workload while preserving the quality of assessment decisions. Organizational Efficiency | positive | Human scoring workload and assessment decision quality |
Reading fidelity
high
Study strength
low
|
n=5000
|
| The study uses operational data from two recent editions of Uruguay's national assessment, with approximately 5,000–6,000 participants in each edition. Other | null_result | Scope of the assessment dataset |
Reading fidelity
high
Study strength
high
|
n=5000
approximately 5,000–6,000 participants per edition
|
| The 2024 assessment data were used to develop the AI-based scoring model, while the 2025 data were used for validation. Decision Quality | positive | Out-of-year validation of the AI scoring model |
Reading fidelity
high
Study strength
high
|
n=5000
|
| The Writing section is the most time- and resource-intensive component of the assessment because written responses are evaluated manually using a detailed rubric. Task Completion Time | negative | Human time and resources required for writing-section scoring |
Reading fidelity
high
Study strength
high
|
not reported
|
| Applicants typically receive their final assessment results no sooner than three months after the examination. Task Completion Time | negative | Time from examination to final result reporting |
Reading fidelity
high
Study strength
medium
|
no sooner than three months
|
| The proposed operational principle is that no final decision affecting test outcomes should be made without appropriate human oversight. Governance And Regulation | positive | Human oversight of consequential pass/fail assessment decisions |
Reading fidelity
high
Study strength
low
|
not reported
|