Does AI actually teach better?
Four new meta-analyses of randomised trials reach different verdicts on whether AI-based teaching beats the traditional kind.
This Week in MedEd
3 new funding calls, 2 new events, 21 jobs (including 6 new and 3 close this week)
⭐ Closes Today: University of Buckingham is hiring a Senior Lecturer, MLA Lead (Clinical), to lead its Medical Licensing Assessment programme.
The full database, with deadlines, links and eligibility, free for 30 days:
Every medical school that has waved piloted AI in learning and teaching has, sooner or later, had to answer the same question: does your pilot actually work? Since the arrival of genAI the honest answer has been “we don’t know.” Thankfully, given the increasing volume of scholarship and research, that is starting to change.
Four meta-analyses, published within weeks of each other, have each pooled dozens of randomised controlled trials to ask whether AI-based teaching genuinely improves what students know and can do, compared with the traditional teaching (delivered by humans!). However the question isn’t settled; I’m covering all four together this week precisely because they don’t agree cleanly with each other. The gaps between them, in scope, population and how strictly they weigh the underlying trials, tell you more about the current state of the evidence than any single pooled effect size does.
Key points:
Four new meta-analyses of RCTs each ask whether AI-based teaching beats traditional teaching, and they reach visibly different conclusions.
The two smaller, more optimistic reviews found large, consistent benefits for knowledge, skills and satisfaction.
The two larger, more methodologically cautious reviews found the picture is patchy: strong for some outcomes and populations, flat or unclear for others.
Every review flags the same structural weakness: nobody has found a way to blind a trial when one arm gets a chatbot and the other doesn’t.
The most consistent finding across all four is that AI helps more with skills practice than with rote knowledge, and more in some specialties and student groups than others.
Pillai et al. offer the most positive result1. Their meta-analysis pooled 14 RCTs and 1,116 students comparing AI-based teaching with traditional teaching across a mix of specialties, from paediatrics to hepatobiliary surgery. AI-based teaching came out ahead on every outcome measured: knowledge, skills and teaching satisfaction. Knowledge show the smallest gains. Practical courses benefited more than theoretical ones, and the effect held whether the course ran for a single session or several weeks.
There were some useful critical findings:
None of the 14 trials could blind students to which arm they were in
Funnel plots suggested small negative studies were under-represented
There was high heterogeneity between studies.
The authors’ own recommendation is that future trials needs to move into postgraduate and subspecialty training, where the evidence is thinnest.
Wang et al. cover similar ground with a larger and more recent evidence base. They reviewed 20 RCTs with 1,413 participants, published in 2024 or 20252. GenAI-assisted teaching was associated with large gains in post-intervention knowledge and skills, and the knowledge gains were still detectable at follow-up in the four studies that measured retention.
This review also flags some issues with the existing research:
None of the included trials was rated low risk of bias overall
The prediction interval for the knowledge result actually crossed zero, meaning a future trial could plausibly find no benefit at all
16 of the 20 studies came from a single country.
Interesting, two crossover trials compared AI feedback directly against expert human feedback and found that students did slightly worse with the AI..
Lai et al. have undertook the most comprehensive review3. They covered a whopping 66 RCTs with 4,911 participants, screened down from nearly 40,000 records. The RCTs were split into 13 categories of AI tool and seven outcome domains.
Large language model-based personalised learning aids, the biggest single category, showed positive effects on satisfaction, confidence and knowledge, but every one of those results was rated very low certainty, most of the included trials had a high risk of bias, and not a single study measured whether any of this translated into better clinical performance or patient outcomes further down the training pipeline.
Their conclusion is noncommittal: treat AI applications as something to trial in your own context, not as an intervention with settled evidence behind it.
Xie et al. reviewed 31 studies with 2,615 undergraduate participants4. They found no significant improvement in either knowledge or skills for medicine students using GenAI teaching compared with conventional methods. Nursing students, by contrast, saw a genuinely large improvement in knowledge.
The skills result for clinical medicine students became significant once a single outlying study on fine motor procedural skills was removed in a sensitivity analysis, which suggests the type of skill matters as much as the presence of AI.
The authors argue for what they call a “differentiated integration strategy”. AI’s benefit is not uniform across specialties, student populations or skill types, so it shouldn’t be adopted uniformly either.
So, there you have it. Despite literally thousands of studies there is still no clear evidence that genAI benefits learning.
The positive story seems to revolve mostly around skills practice, in nursing education specifically. Knowledge gains for clinical medicine students are present, but less grand than the biggest headline numbers suggest.
These four reviews all flag the same structural problem: you cannot blind a student to whether they are talking to a chatbot, so some share of the measured benefit may be enthusiasm for something new rather than durable learning.
None of this means AI-assisted teaching doesn’t work. It means that for now it is sensible is to treat genAI as a supplement worth piloting with your own outcome measures in your own context, rather than an intervention that has already proven itself.
Beyond the scope of this review, and perhaps something we’ll cover in another future edition, is the long-term impact of genAI on how we all learn and think, and how this begins to filter into the clinical workplace. Stay tuned for more!
Pratyush Pillai, Shayan Saadat, Imogen Ward, Shahab Hajibandeh, Shahin Hajibandeh, Meta-analysis of randomized controlled trials comparing outcomes of artificial intelligence-based teaching versus traditional-based teaching in medical education, Postgraduate Medical Journal, 2026;, qgag077, https://doi.org/10.1093/postmj/qgag077
Wang, C., Sun, N., Pan, X. et al. The impact of integrating generative artificial intelligence into medical education on short-term learning outcomes: a systematic review and meta-analysis of randomized controlled trials. BMC Med Educ 26, 983 (2026). https://doi.org/10.1186/s12909-026-09320-6
Lai NM, Lim YS, Win MT, Bhargava P,Thomas P, Ong QC. The Effectiveness of Artificial Intelligence in Undergraduate Health Professions Education: Systematic Review and Meta-Analysis of Randomized Controlled Trials. JMIR Med Educ 2026;12:e88933 doi: 10.2196/88933
Xie, G., Liao, J., Tang, X. et al. The impact of generative artificial intelligence on clinical skills and knowledge acquisition in medical undergraduates: a systematic review and meta-analysis. BMC Med Educ 26, 1144 (2026). https://doi.org/10.1186/s12909-026-09458-3



