The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-augmented TAs improve student writing: students whose teaching assistants received LLM suggestions revised essays to higher quality, and the boost grew with TA uptake of those suggestions.

AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course
Xinyi Lu, Kexin Phyllis Ju, Mitchell Dudley, Larissa Sano, Xu Wang · February 18, 2026
arxiv rct high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Xinyi Lu unresolved corpus identity
  2. Kexin Phyllis Ju unresolved corpus identity
  3. Mitchell Dudley unresolved corpus identity
  4. Larissa Sano unresolved corpus identity
  5. Xu Wang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Xinyi Lu provider ID
  2. Kexin Phyllis Ju provider ID
  3. Mitchell Dudley provider ID
  4. Larissa Sano provider ID
  5. Xu Wang provider ID
In an RCT of 354 students, providing TAs with LLM-generated suggestions (via FeedbackWriter) led to significantly higher-quality student essay revisions, with larger gains when TAs adopted more AI suggestions.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Despite growing interest in using LLMs to generate feedback on students' writing, little is known about how students respond to AI-mediated versus human-provided feedback. We address this gap through a randomized controlled trial in a large introductory economics course (N=354), where we introduce and deploy FeedbackWriter - a system that generates AI suggestions to teaching assistants (TAs) while they provide feedback on students' knowledge-intensive essays. TAs have the full capacity to adopt, edit, or dismiss the suggestions. Students were randomly assigned to receive either handwritten feedback from TAs (baseline) or AI-mediated feedback where TAs received suggestions from FeedbackWriter. Students revise their drafts based on the feedback, which is further graded. In total, 1,366 essays were graded using the system. We found that students receiving AI-mediated feedback produced significantly higher-quality revisions, with gains increasing as TAs adopted more AI suggestions. TAs found the AI suggestions useful for spotting gaps and clarifying rubrics.

Summary

Main Finding

AI-mediated feedback—where teaching assistants (TAs) receive LLM-generated, rubric-aligned suggestions via the FeedbackWriter system and can adopt/edit/dismiss them—led to substantially better student revisions in a large undergraduate economics course. Students receiving AI-mediated feedback produced higher-quality revised drafts (Cohen’s d = 0.50; roughly moving a student from the 50th to the 70th percentile). Effects grew with greater TA adoption of AI suggestions. AI-mediated feedback also scored higher on actionability and promotion of independent learning; TAs reported the system helped them spot missed issues and align with rubrics.

Key Points

  • Intervention: FeedbackWriter supplies one AI-generated feedback message per rubric item, anchored to relevant sentences, with transparency controls (sentence highlighting, rubric judgments, option to use historical instructor comments) and requires TA confirmation before sending.
  • Experimental design: Randomized controlled trial in an intro economics course (N = 354 students, 11 TAs). Students were randomized at semester start: assignment 1 half received human-only feedback / half AI-mediated; conditions reversed for assignment 2 (crossover). Total of 1,366 essays graded through the system.
  • Primary outcome: Quality of revised drafts — AI-mediated feedback produced significantly higher-quality revisions (Cohen’s d = 0.50).
  • Dose response: Larger improvement when TAs adopted more AI-suggested constructive feedback.
  • Secondary outcomes: Revised-draft quality correlated with higher learning on a post-test, but there was no significant difference in post-test scores between the two experimental conditions.
  • Feedback quality analysis: Automated evaluation showed AI-mediated feedback outperformed human-only feedback on actionability and supporting independent learning.
  • TA experience: TAs found suggestions generally accurate, useful for spotting overlooked gaps and aligning with rubrics; appreciated ease of detecting and correcting AI mistakes. TAs relied on rubrics and benefited from AI reducing time/effort to find relevant sentences and phrase constructive questions.
  • Limitations noted by authors: dependence on well-specified rubrics, course context (knowledge-intensive econ essays), LLM errors still present but mitigated by TA oversight, generalizability to other domains and longer-term learning not fully established.

Data & Methods

  • Setting: Introduction to Economics (ECON101) at a large public R1 university; two writing assignments with detailed, knowledge-focused rubrics; TA graders trained and normed.
  • Participants: 354 students, 11 TAs. Essays graded: 1,366 across drafts and assignments.
  • Treatment conditions: Human-only feedback (baseline) vs. AI-mediated feedback (TAs received AI suggestions via FeedbackWriter and could adopt/edit/dismiss).
  • Randomization/crossover: Students randomized into conditions for assignment 1; conditions reversed for assignment 2 (within-course crossover design).
  • Outcomes measured:
    • Blind-graded revision quality (primary outcome).
    • Post-test for learning (secondary outcome).
    • Automated/annotated analyses of feedback properties (actionability, specificity, etc.).
    • TA qualitative responses and formative study observations on workflow.
  • Analysis: Effect size reported (Cohen’s d = 0.50) for revision quality; heterogeneity analysis showing stronger gains with higher TA adoption of AI suggestions; automated metrics showed improvements in actionable and independence-supporting feedback. No significant difference found on post-test between conditions.

Implications for AI Economics

  • Productivity and labor augmentation
    • Human-AI pairing can materially raise the productivity and effectiveness of semi-skilled graders (TAs) in knowledge-intensive assessment tasks. This suggests AI is more likely to augment rather than fully replace expert graders in contexts that require judgment and domain knowledge.
    • The design—AI suggestions + human oversight—mitigates risks from hallucinations while speeding the identification of missed points and phrasing of actionable guidance.
  • Cost-effectiveness and scaling personalized education
    • Better-quality feedback at scale (same TA workforce) implies potential cost-effective scaling of higher-quality formative assessment in large classes. Economic evaluations should quantify TA time savings, incremental learning gains, and system deployment/training costs.
  • Human capital and skill complementarity
    • The results show complementarities: AI-generated scaffolding (spotting evidence, phrasing probe questions) enhances a human grader’s ability to deliver pedagogically valuable feedback. Demand may shift toward graders skilled in oversight, calibration, and pedagogical judgment rather than raw grading throughput.
  • Implementation and adoption considerations
    • Value depends on availability of clear rubrics and some initial knowledge engineering; institutions need to invest in rubric design and grader training to realize benefits.
    • Interfaces that preserve human agency (adopt/edit/dismiss) appear critical for trust and error correction; policy or procurement should prioritize such human-in-the-loop designs.
  • Distributional and long-term effects
    • Short-term gains in revision quality do not automatically translate into detected improvements on immediate post-tests; longitudinal studies are needed to assess persistence of learning, credential value, and labor market impacts.
    • Equitable effects: research should test whether benefits are uniform across student ability levels or whether AI-mediated feedback widens/narrows performance gaps.
  • Research and policy priorities for economists
    • Conduct cost–benefit analyses comparing instructor time saved, quality-adjusted learning gains, and deployment costs across course sizes and disciplines.
    • Study labor-market impacts: will demand shift from graders to higher-skilled pedagogical overseers? What retraining is required?
    • Examine generalizability across domains with varying knowledge-intensity, and test pure substitution vs. augmentation scenarios.
    • Consider regulatory and ethical issues: data privacy, transparency about AI use, quality assurance standards for LLM suggestions in educational settings.
  • Practical recommendation for institutions
    • Pilot human-in-the-loop AI tooling anchored to explicit rubrics, with mandatory human review and logging of AI adoption rates. Monitor revision-quality impacts, TA time-to-grade, and student learning outcomes before scaling.

If you want, I can produce a one-page brief for university administrators estimating potential TA time savings and a simple cost–benefit framework to evaluate adopting FeedbackWriter-like systems.

Assessment

Paper Typerct Evidence Strengthhigh — A randomized controlled trial with a reasonably large sample (N=354 students, 1,366 graded essays) gives strong internal validity for the main comparison between AI-mediated and human-only feedback; results are robust to observed variation in TA adoption of suggestions. Methods Rigorhigh — Uses randomized allocation, multiple essay observations per student, and direct measurement of essay quality after revision; however, there are plausible concerns (noted below) about potential grader/TA non-blinding, clustering at TA level, and interpretation of adoption effects which slightly temper the assessment but do not undermine the randomized comparison. SampleUndergraduate students enrolled in a large introductory economics course (N=354 students) who produced multiple drafts of knowledge-intensive essays; 1,366 essays were graded in total. Teaching assistants graded and provided feedback; in the treatment arm TAs received LLM-generated suggestions via FeedbackWriter which they could adopt, edit, or dismiss. Themeshuman_ai_collab skills_training IdentificationRandomized assignment of students to two feedback conditions (handwritten TA feedback vs AI-mediated feedback where TAs received LLM-generated suggestions via FeedbackWriter). This randomization provides causal identification of the effect of AI-mediated feedback on subsequent essay revisions; variation in TA adoption of suggestions is used to assess dose-response (heterogeneous treatment effects), though adoption itself is not randomized. GeneralizabilitySingle course (introductory economics) and likely one university context — may not generalize to other subjects or institutional settings, Undergraduate student population; effects may differ for K-12, graduate students, or working professionals, Intervention involves TAs mediating AI suggestions (human-in-the-loop) — findings may not apply to fully automated feedback systems, Short-term outcome: revision quality on subsequent draft(s); long-term learning gains or lasting skill acquisition were not measured, Results may depend on the specific LLM, prompts, and FeedbackWriter interface used, and on TA experience and incentives

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We conducted a randomized controlled trial in a large introductory economics course (N=354) comparing AI-mediated feedback to handwritten TA feedback. Other null_result assignment to feedback condition (AI-mediated vs handwritten)
Reading fidelity high
Study strength high
n=354
1.0
Students receiving AI-mediated feedback produced significantly higher-quality revisions than students receiving handwritten feedback. Output Quality positive quality of revised essays (revision quality)
Reading fidelity high
Study strength high
n=354
1.0
Gains in revision quality increased as TAs adopted more AI suggestions. Output Quality positive change in revision quality as a function of TA adoption rate of AI suggestions
Reading fidelity high
Study strength medium
not reported
0.6
TAs found the AI suggestions useful for spotting gaps and clarifying rubrics. Worker Satisfaction positive TA-perceived usefulness of AI suggestions
Reading fidelity medium
Study strength low
not reported
0.18
TAs had the full capacity to adopt, edit, or dismiss the AI-generated suggestions while providing feedback. Other null_result TA interaction options with AI suggestions (adopt/edit/dismiss)
Reading fidelity high
Study strength high
not reported
1.0
In total, 1,366 essays were graded using the system during the study. Other null_result number of essays graded
Reading fidelity high
Study strength high
n=1366
1.0
Students revised their drafts based on the feedback, and those revisions were subsequently graded. Other null_result submission and grading of revised drafts
Reading fidelity high
Study strength high
not reported
1.0

Notes