The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A four-week practitioner-led micro-RCT finds an AI tutoring platform (Medly) produced a measurable short-term gain in GCSE science revision—about one-third of a standard deviation—versus ordinary self-directed revision, but high attrition, non-standardised outcomes and limited process data make the result a provisional signal rather than a definitive effectiveness estimate.

Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science
Wayne Harrison, Rahil Khowaja, Emma Dobson, Germaine Uwimpuhwe, Steve Higgins · September 13, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Wayne Harrison unresolved corpus identity
  2. Rahil Khowaja unresolved corpus identity
  3. Emma Dobson unresolved corpus identity
  4. Germaine Uwimpuhwe unresolved corpus identity
  5. Steve Higgins unresolved corpus identity

Semantic Scholar

Latest observation:

  1. W. Harrison provider ID
  2. Rahil Khowaja provider ID
  3. E. Dobson provider ID
  4. Germaine Uwimpuhwe provider ID
  5. S. Higgins provider ID
In a multisite micro-RCT with 644 post-test completers, allocation to the Medly AI tutoring platform raised short-term GCSE-science revision scores by about Hedges' g = 0.33 versus business-as-usual self-directed revision.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial intelligence (AI) systems in education are developing on timescales that sit uneasily with conventional evaluation. By the time a large-scale trial has been designed, delivered, analysed and published, the technology under study may have changed materially. This creates a temporal problem for evidence-informed education: the need for timely evidence can encourage reliance on weak observational or usage data, while conventional rigorous evaluation may produce evidence too slowly to guide rapidly evolving practice. We examine teacher-led micro-randomised controlled trials (micro-RCTs) as one response to this problem. The empirical case is a four-week multisite individually randomised evaluation of Medly, an AI-powered tutoring platform, in GCSE Biology, Chemistry and Physics in English secondary schools. Of 929 students completing baseline assessment, 644 completed post-testing. In the primary ITT analysis, students allocated to Medly achieved higher post-test attainment than students undertaking business-as-usual self-directed revision (Hedges' g = 0.33, 95% CI 0.18 to 0.48). Positive estimates were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52), with no evidence of differential impact by disadvantage status. Greater platform engagement was associated with higher attainment, but these post-randomisation analyses are treated as exploratory rather than causal. Attrition was substantial (30.7%), outcome measures were curriculum-aligned rather than standardised, and process evaluation response was limited. We therefore interpret the findings as preliminary. We argue that the value of micro-RCTs for educational AI lies not in replacing definitive evaluation with small studies, but in enabling a rapid, cumulative evaluation architecture in which randomised estimates can be generated, replicated and updated as technologies and their implementation evolve.

Summary

Main Finding

A practitioner-led, four-week multisite micro-randomised trial (Npre = 929, Npost = 644) found that allocation to Medly — an AI-powered, curriculum-aligned tutoring platform for GCSE Science — produced a positive short-term impact on post-test attainment versus business-as-usual self-directed revision. Intention-to-treat: adjusted difference = +2.56 marks (95% CI 1.39 to 3.73); Hedges’ g = 0.33 (95% CI 0.18 to 0.48). Positive effects were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52). No clear differential effect by Pupil Premium (disadvantage proxy) was detected. Engagement (questions answered) was positively associated with attainment, but these analyses are exploratory and non-causal.

Key Points

  • Intervention: Medly, a generative-AI tutoring platform combining LLM layers with an exam-specific knowledge base; students received one ~30-minute activity per week for 4 weeks.
  • Comparator: 30 minutes/week of business-as-usual self-directed revision (heterogeneous, frequently technology-supported).
  • Primary outcome: curriculum-aligned GCSE-style assessment (5 questions, max 35); AI-marked then teacher-verified.
  • Sample & attrition: 929 baseline; 644 post-test (30.7% attrition). Attrition higher in control (33.2%) than intervention (27.7%).
  • Effect sizes: overall g = 0.33; subject-specific g = 0.31 (Physics), 0.32 (Chemistry), 0.52 (Biology). Subject-by-treatment interaction not significant.
  • Engagement: each additional question answered ~ +0.18 post-test marks (associational).
  • Implementation variability: delivery modes varied (homework vs in-class), device/login friction reported; only 6 of 39 teacher trials returned process surveys.
  • Limitations stressed by authors: short follow-up, curriculum-aligned (not standardised) outcomes, substantial attrition, limited process data, evolving product (moving intervention). Findings treated as preliminary.

Data & Methods

  • Design: Practitioner-led parallel micro-RCTs (three subject trials) conducted in mainstream English secondary schools; individual randomisation via WhatWorked Teachers platform; 4-week intervention window.
  • Participants: Year 9–10 pupils; each pupil enrolled in one subject trial only.
  • Randomisation & masking: Individual randomisation handled by platform; teachers could not give control students access to Medly content during trial.
  • Outcome measurement:
    • Baseline: 30-minute pre-test.
    • Post: equivalent curriculum-aligned test (different items).
    • Scoring: AI auto-mark then teacher manual check/amendment.
  • Engagement measure: platform logs — number of Medly questions answered during intervention.
  • Covariates: baseline score; Pupil Premium status (free-school-meal eligibility proxy).
  • Analysis:
    • Primary: intention-to-treat ANCOVA-style mixed-effects models adjusting for baseline and clustering at school level; reported as adjusted mean differences and Hedges’ g with 95% CIs.
    • Subgroup: treatment-by-Pupil Premium interaction.
    • Exploratory associations: engagement vs outcome adjusted for baseline (not claimed causal).
  • Ethical/data handling: anonymised records; processed under UK GDPR public-interest basis and headteacher consent; limited teacher survey response.

Implications for AI Economics

  • Reducing information asymmetry and procurement risk:
    • Micro-RCTs embedded in routine practice can generate timely causal signals that reduce buyer/deployer uncertainty about incremental educational value, improving procurement decisions for schools and districts.
    • Rapid, repeatable randomised tests can create a stream of version-specific evidence; this helps purchasers evaluate current product iterations rather than outdated trial results.
  • Product-market fit, iteration and valuation:
    • A repeatable micro-RCT architecture supports faster product iteration cycles: developers can A/B or version-test pedagogical or interface changes with causal readouts, informing R&D priorities and feature investments.
    • Investors and acquirers should value firms that operationalise continuous randomized evaluation (capacity to demonstrate incremental impact across versions/cohorts), since this reduces execution and adoption risk.
  • Pricing, contracts and performance-based procurement:
    • Evidence of positive short-run effects (g ≈ 0.3) provides a basis for modeling willingness-to-pay and cost-effectiveness, but firms and buyers need cost-per-unit-effect metrics (e.g., cost per 0.1 SD improvement) and long-run outcome linkage to exam performance to price sustainably.
    • Micro-RCTs enable alternative contracting: pay-for-performance or milestone contracts could be conditioned on repeated, pre-specified randomized estimates, though contract design must account for versioning and nonstationarity.
  • Scalability, heterogeneity and risk:
    • Heterogeneous implementation (in-class vs homework, device constraints) and attrition (30.7%) imply realized ROI will vary across contexts. Economic models should incorporate heterogeneity in uptake, engagement, and access costs (devices, connectivity, teacher time).
    • The absence of a clear differential effect by disadvantage (Pupil Premium) is cautiously encouraging from an equity perspective, but sample sizes and attrition limit conclusions; economists modelling distributional impacts should demand larger/longer trials that test equity outcomes explicitly.
  • Engagement, monetization and incentives:
    • Positive association between engagement and outcomes suggests business models that increase effective engagement (e.g., nudges, gamification, teacher integration) could raise realized value — but causality is unproven. Monetization (subscription, freemium, licensing to schools) must consider whether paid features actually causally increase engagement and learning.
  • Evaluation economics & market for evidence services:
    • Platforms like WhatWorked Teachers that automate randomisation, data collection and reporting lower the fixed and transaction costs of repeated trials, creating a market for continuous-evidence-as-a-service. This capability is a potential competitive advantage and could become part of product bundles.
  • Policy and regulatory implications:
    • Fast-moving AI products highlight the limits of one-off large trials. Regulators, funders and procurement bodies should consider requiring rolling evaluation, version disclosure, and pre-registered core outcomes to maintain an up-to-date evidence base.
  • Research and investment priorities for robust economic assessment:
    • Longer-term and standardised outcomes are needed to link short-term gains to high-stakes metrics (GCSE grades, progression) and to compute cost-effectiveness.
    • Designs to identify causal effects of engagement (encouragement designs, instrumental variables, complier-based approaches) are necessary to value marginal investments that increase usage.
    • Meta-analytic and hierarchical models over repeated micro-RCTs can quantify mean effects, heterogeneity, and the value of learning over time — useful for portfolio valuation and for forecasting adoption returns under uncertainty.
  • Practical caution for economic modelling:
    • Versioning (“moving intervention”) means that historic effect estimates depreciate; any economic valuation should incorporate decay/obsolescence of evidence and a plan for evidence refresh, which affects expected returns and pricing strategies.

Summary takeaway: The trial provides a preliminary positive signal that a curriculum-aligned AI tutor can raise short-term attainment (g ~ 0.3). For AI economics, the most valuable contribution of this work is methodological: demonstrating a feasible, low-latency randomized evaluation architecture that can reduce uncertainty, inform product iteration and support evidence-based procurement — but full economic valuation requires repeated trials, standardized longer-run outcomes, causal identification of engagement channels, and careful treatment of version risk and implementation heterogeneity.

Assessment

Paper Typerct Evidence Strengthmedium — Randomised allocation gives credible causal identification for the short-term ITT contrast, and estimates are reported with CIs; however, substantial and differential attrition (~31% overall, higher in control), non-standard curriculum-aligned outcome measures that were AI-marked then teacher-checked, short follow-up (4 weeks), heterogeneous active control, and limited process data weaken internal validity and external applicability, so the causal signal is provisional rather than definitive. Methods Rigormedium — Strong elements: individual randomisation, prespecified ITT, baseline adjustment, mixed-effects for clustering, realistic active control; Weaknesses: high and differential attrition without detailed attrition sensitivity analyses reported, outcome measurement not independently standardized and subject to teacher adjustment, short duration, limited process evaluation, and post-randomisation engagement analyses treated as non-causal. SampleMainstream English secondary-school students in Year 9 or Year 10; 929 completed pre-test across three subject trials (Physics n=505, Chemistry n=275, Biology n=149); 644 completed post-test (Physics 312, Chemistry 217, Biology 115); randomised individually to Medly (intervention) or business-as-usual self-directed revision (control); engagement measured via platform logs; Pupil Premium (free-school-meal eligibility) recorded as binary covariate. Themesskills_training human_ai_collab IdentificationIndividual-level randomisation to Medly versus business-as-usual self-directed revision, allocation concealed through the WhatWorked Teachers platform; primary analysis follows intention-to-treat using ANCOVA (post-test ~ treatment + pre-test) in a mixed-effects model to account for school clustering; engagement analyses are observational and treated as exploratory. GeneralizabilityShort-term (4-week) revision context — unknown persistence of effects over longer periods or real exam performance, Curriculum- and exam-board-specific GCSE content in England — limits transferability to other countries, age groups, or subjects, Outcome is a curriculum-aligned, non-standardised assessment with AI initial marking and teacher amendments — potential measurement bias relative to independent standardized tests, Volunteer schools/teachers and low teacher process-evaluation response (6/39) — implementation variation and selection by adopters may limit representativeness, Active, heterogeneous control (some students used other digital platforms) — effect is incremental over variable existing practice and not versus no-digital condition, Platform-specific (Medly) — results may not generalize to other AI tutoring systems or configurations

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Students allocated to Medly achieved higher post-test attainment than students allocated to business-as-usual self-directed revision. Output Quality positive Score on a subject-specific GCSE-aligned post-test with a maximum of 35 marks
Reading fidelity high
Study strength medium
n=644
Hedges’ g = 0.33 (95% CI 0.18 to 0.48); adjusted difference = 2.56 marks (95% CI 1.39 to 3.73)
0.6
Medly showed positive attainment effects in Physics, Chemistry, and Biology. Output Quality positive Subject-specific GCSE-aligned post-test attainment
Reading fidelity high
Study strength medium
n=644
Physics Hedges’ g = 0.31 (95% CI 0.11 to 0.51); Chemistry Hedges’ g = 0.32 (95% CI 0.06 to 0.58); Biology Hedges’ g = 0.52 (95% CI 0.19 to 0.84)
0.6
There was no clear evidence that Medly’s effect differed by Pupil Premium status. Inequality null_result Differential effect of Medly on GCSE-aligned post-test attainment by disadvantage status
Reading fidelity high
Study strength medium
n=644
Treatment-by-status interaction = 0.57 marks (95% CI -2.25 to 3.39)
0.6
Within students assigned to Medly, greater platform engagement was associated with higher post-test attainment. Output Quality positive GCSE-aligned post-test attainment
Reading fidelity high
Study strength low
n=332
Approximately 0.18 additional post-test marks per additional question answered (95% CI 0.13 to 0.22)
0.3
The positive association between Medly engagement and attainment cannot establish that increasing the number of questions answered causes higher attainment. Decision Quality mixed Causal relationship between platform engagement and post-test attainment
Reading fidelity high
Study strength high
n=332
1.0
Attrition was substantial, with 30.7% of students lost between baseline and post-testing, and attrition was higher in the control group than in the intervention group. Other negative Post-test completion and study retention
Reading fidelity high
Study strength high
n=929
30.7% overall attrition; 33.2% control versus 27.7% intervention
1.0
Implementation of Medly varied across schools and teachers, with technical friction—especially mobile-device access and login—identified as a practical barrier. Organizational Efficiency mixed Implementation consistency and practical usability of the AI tutoring platform
Reading fidelity high
Study strength low
n=6
0.3

Notes