3 cumulative citations
View corpus contextAI-augmented TAs improve student writing: students whose teaching assistants received LLM suggestions revised essays to higher quality, and the boost grew with TA uptake of those suggestions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Despite growing interest in using LLMs to generate feedback on students' writing, little is known about how students respond to AI-mediated versus human-provided feedback. We address this gap through a randomized controlled trial in a large introductory economics course (N=354), where we introduce and deploy FeedbackWriter - a system that generates AI suggestions to teaching assistants (TAs) while they provide feedback on students' knowledge-intensive essays. TAs have the full capacity to adopt, edit, or dismiss the suggestions. Students were randomly assigned to receive either handwritten feedback from TAs (baseline) or AI-mediated feedback where TAs received suggestions from FeedbackWriter. Students revise their drafts based on the feedback, which is further graded. In total, 1,366 essays were graded using the system. We found that students receiving AI-mediated feedback produced significantly higher-quality revisions, with gains increasing as TAs adopted more AI suggestions. TAs found the AI suggestions useful for spotting gaps and clarifying rubrics.
Summary
Main Finding
AI-mediated feedback—where teaching assistants (TAs) receive LLM-generated, rubric-aligned suggestions via the FeedbackWriter system and can adopt/edit/dismiss them—led to substantially better student revisions in a large undergraduate economics course. Students receiving AI-mediated feedback produced higher-quality revised drafts (Cohen’s d = 0.50; roughly moving a student from the 50th to the 70th percentile). Effects grew with greater TA adoption of AI suggestions. AI-mediated feedback also scored higher on actionability and promotion of independent learning; TAs reported the system helped them spot missed issues and align with rubrics.
Key Points
- Intervention: FeedbackWriter supplies one AI-generated feedback message per rubric item, anchored to relevant sentences, with transparency controls (sentence highlighting, rubric judgments, option to use historical instructor comments) and requires TA confirmation before sending.
- Experimental design: Randomized controlled trial in an intro economics course (N = 354 students, 11 TAs). Students were randomized at semester start: assignment 1 half received human-only feedback / half AI-mediated; conditions reversed for assignment 2 (crossover). Total of 1,366 essays graded through the system.
- Primary outcome: Quality of revised drafts — AI-mediated feedback produced significantly higher-quality revisions (Cohen’s d = 0.50).
- Dose response: Larger improvement when TAs adopted more AI-suggested constructive feedback.
- Secondary outcomes: Revised-draft quality correlated with higher learning on a post-test, but there was no significant difference in post-test scores between the two experimental conditions.
- Feedback quality analysis: Automated evaluation showed AI-mediated feedback outperformed human-only feedback on actionability and supporting independent learning.
- TA experience: TAs found suggestions generally accurate, useful for spotting overlooked gaps and aligning with rubrics; appreciated ease of detecting and correcting AI mistakes. TAs relied on rubrics and benefited from AI reducing time/effort to find relevant sentences and phrase constructive questions.
- Limitations noted by authors: dependence on well-specified rubrics, course context (knowledge-intensive econ essays), LLM errors still present but mitigated by TA oversight, generalizability to other domains and longer-term learning not fully established.
Data & Methods
- Setting: Introduction to Economics (ECON101) at a large public R1 university; two writing assignments with detailed, knowledge-focused rubrics; TA graders trained and normed.
- Participants: 354 students, 11 TAs. Essays graded: 1,366 across drafts and assignments.
- Treatment conditions: Human-only feedback (baseline) vs. AI-mediated feedback (TAs received AI suggestions via FeedbackWriter and could adopt/edit/dismiss).
- Randomization/crossover: Students randomized into conditions for assignment 1; conditions reversed for assignment 2 (within-course crossover design).
- Outcomes measured:
- Blind-graded revision quality (primary outcome).
- Post-test for learning (secondary outcome).
- Automated/annotated analyses of feedback properties (actionability, specificity, etc.).
- TA qualitative responses and formative study observations on workflow.
- Analysis: Effect size reported (Cohen’s d = 0.50) for revision quality; heterogeneity analysis showing stronger gains with higher TA adoption of AI suggestions; automated metrics showed improvements in actionable and independence-supporting feedback. No significant difference found on post-test between conditions.
Implications for AI Economics
- Productivity and labor augmentation
- Human-AI pairing can materially raise the productivity and effectiveness of semi-skilled graders (TAs) in knowledge-intensive assessment tasks. This suggests AI is more likely to augment rather than fully replace expert graders in contexts that require judgment and domain knowledge.
- The design—AI suggestions + human oversight—mitigates risks from hallucinations while speeding the identification of missed points and phrasing of actionable guidance.
- Cost-effectiveness and scaling personalized education
- Better-quality feedback at scale (same TA workforce) implies potential cost-effective scaling of higher-quality formative assessment in large classes. Economic evaluations should quantify TA time savings, incremental learning gains, and system deployment/training costs.
- Human capital and skill complementarity
- The results show complementarities: AI-generated scaffolding (spotting evidence, phrasing probe questions) enhances a human grader’s ability to deliver pedagogically valuable feedback. Demand may shift toward graders skilled in oversight, calibration, and pedagogical judgment rather than raw grading throughput.
- Implementation and adoption considerations
- Value depends on availability of clear rubrics and some initial knowledge engineering; institutions need to invest in rubric design and grader training to realize benefits.
- Interfaces that preserve human agency (adopt/edit/dismiss) appear critical for trust and error correction; policy or procurement should prioritize such human-in-the-loop designs.
- Distributional and long-term effects
- Short-term gains in revision quality do not automatically translate into detected improvements on immediate post-tests; longitudinal studies are needed to assess persistence of learning, credential value, and labor market impacts.
- Equitable effects: research should test whether benefits are uniform across student ability levels or whether AI-mediated feedback widens/narrows performance gaps.
- Research and policy priorities for economists
- Conduct cost–benefit analyses comparing instructor time saved, quality-adjusted learning gains, and deployment costs across course sizes and disciplines.
- Study labor-market impacts: will demand shift from graders to higher-skilled pedagogical overseers? What retraining is required?
- Examine generalizability across domains with varying knowledge-intensity, and test pure substitution vs. augmentation scenarios.
- Consider regulatory and ethical issues: data privacy, transparency about AI use, quality assurance standards for LLM suggestions in educational settings.
- Practical recommendation for institutions
- Pilot human-in-the-loop AI tooling anchored to explicit rubrics, with mandatory human review and logging of AI adoption rates. Monitor revision-quality impacts, TA time-to-grade, and student learning outcomes before scaling.
If you want, I can produce a one-page brief for university administrators estimating potential TA time savings and a simple cost–benefit framework to evaluate adopting FeedbackWriter-like systems.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We conducted a randomized controlled trial in a large introductory economics course (N=354) comparing AI-mediated feedback to handwritten TA feedback. Other | null_result | assignment to feedback condition (AI-mediated vs handwritten) |
Reading fidelity
high
Study strength
high
|
n=354
|
| Students receiving AI-mediated feedback produced significantly higher-quality revisions than students receiving handwritten feedback. Output Quality | positive | quality of revised essays (revision quality) |
Reading fidelity
high
Study strength
high
|
n=354
|
| Gains in revision quality increased as TAs adopted more AI suggestions. Output Quality | positive | change in revision quality as a function of TA adoption rate of AI suggestions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| TAs found the AI suggestions useful for spotting gaps and clarifying rubrics. Worker Satisfaction | positive | TA-perceived usefulness of AI suggestions |
Reading fidelity
medium
Study strength
low
|
not reported
|
| TAs had the full capacity to adopt, edit, or dismiss the AI-generated suggestions while providing feedback. Other | null_result | TA interaction options with AI suggestions (adopt/edit/dismiss) |
Reading fidelity
high
Study strength
high
|
not reported
|
| In total, 1,366 essays were graded using the system during the study. Other | null_result | number of essays graded |
Reading fidelity
high
Study strength
high
|
n=1366
|
| Students revised their drafts based on the feedback, and those revisions were subsequently graded. Other | null_result | submission and grading of revised drafts |
Reading fidelity
high
Study strength
high
|
not reported
|