This Week in MedEd
9 open funding calls, 40 upcoming events, 47 live jobs (22 new, 15 closing this week), including:
💰 Medical Student Education Program (up to $1.625M from the US Health Resources and Services Administration, deadline 1 Sep 2026)
🎓 AI in the Global Classroom: Between Innovation and Inequity (ASME MEGAbite, online, 9 Sep 2026)
💼 Faculty Open Rank (Tenure), Data Science and AI in Health Professions Education (University of Illinois Chicago)
The full list, with deadlines, links and eligibility": Try everything free for 30 days.
The simulated patient encounter is one of the most expensive things a medical school does. Actors have to be recruited, trained, and paid. They need car parking and cups of tea. OSCE circuits need rooms, examiners, and a whole day of everyone’s diary. Simulation centres need managers, technicians and other support staff. In the past, we had assumed that these activities would be the last to be automated, because simulation is relational, embodied, and stubbornly human.
However, that assumption is now out of date. Not only has AI become more capable and freely available, but the challenges associated with scaling traditional simulation are becoming insurmountable. I initially wrote about this capacity challenge, and the potential for AI to be part of the solution in Simulation in Healthcare last year1.
I have selected four recent papers that map how AI is being used in simulation, and between them they cover the whole cycle: writing the case, playing the patient, marking the station, and running the debrief afterwards. You can read more about my own AI patient simulator, SimPatient, here:
Key points
Of 39 implemented AI innovations in undergraduate clinical skills curricula since 2022, 19 were LLM-powered virtual patients.
162 medical students at the Charité gained a mean of 0.94 points of self-rated communication competence (Cohen d=0.58) after one 22-minute conversation with a role-prompted GPT-4o patient.
The one case where the chatbot failed to produce a gain was motivational interviewing, which requires the simulated patient to be ambivalent and resistant rather than agreeable. This is a clear practical consequence of acquiescence bias, or sycophancy.
Across 22 OSCE studies, AI scoring agreed best with humans on visually observable tasks and worst on communication-dependent elements, with transcript-based agreement ranging from 26% to 83%.
When LLMs stood in as simulated patients, 70% of faculty rated their performance good or excellent, but 85% flagged the absence of non-verbal communication and the inability to be examined.
Thind et al. undertook a scoping review, which screened 1,130 records down to 392. The review describes how AI was implemented in undergraduate clinical skills teaching between January 2022 and January 2026. They identified eight categories:
LLM-based virtual patient and clinical simulation systems account for 19 of the 39 papers
AI-augmented OSCE and simulation assessment (6)
Embodied and robotic simulations (4)
Procedural skills training (3)
Documentation and multimodal analytics (2 each)
Educator-facing case authoring (2)
Clinical reasoning tutoring (1).
The pace of development is increasing; 25 of the 39 were published in 2025 alone. Feasibility and acceptability were consistently reported while longitudinal retention and transfer to real patients were almost never measured, which we have been echoed in other recent editions of this newsletter.
Müller et al. built four GPT-4o patient personas mapped to core objectives in the German national curriculum3:
Shared decision-making
Motivational interviewing
Sexual history taking
Breaking bad news
A total of 162 students each had a single conversation, which lasted a median 22 minutes and a mean of 13 prompt-response pairs, then received structured AI feedback across eight criteria. Self-rated communication competence rose from 5.89 to 6.83 on a 0 to 10 scale (mean difference 0.94, Cohen d=0.58), with significant gains in three of the four cases and the largest gains among students who started lowest.
Interestingly, motivational interviewing produced no significant change. The authors suggest this is because the AI is not very good at portraying a patient who expresses ambivalence and resistance, which "a fundamentally supportive AI chatbot may struggle to represent adequately". I think this might be the best material example of sycophancy or acquiescence bias in medical education.
Students rated the AI feedback useful (7.92 out of 10), but feedback adequacy did not correlate with competence gain at all (r=0.08, p=.32), and the model awarded 85% to 92 of the 162 students. More sycophancy?! It should be noted that there was no control group and the student outcomes are self-reported, so read this as evidence about confidence rather than competence.
León-Ariza et al. completed a scoping review to consider how AI can be used to mark OSCEs4. They screened 421 records to 22 included studies and used the FACETS and SAMR frameworks. Automated video scoring produced higher and more uniform marks than human examiners (intramuscular injection 28.23 versus 25.25, knot tying 16.07 versus 10.44), with agreement strongest for visually observable tasks and weakest for communication-dependent elements.
ChatGPT-4.0 scored clinical records at a level statistically indistinguishable from residents while 3.5 did not. When generative models played the simulated patient, 35% of faculty rated the performance excellent and another 35% good, yet 85% identified limitations in non-verbal communication and the obvious inability to be physically examined.
Their conclusion is that AI currently augments rather than transforms OSCE assessment. Their paper contains the now ubiquitous recommendations / frameworks for use, including local validation before adoption in a new curriculum or language, bias audit across learner subgroups at fixed intervals, a documented appeal pathway when AI-supported scores affect progression, and a flat recommendation against fully autonomous AI scoring in high-stakes decisions.
Khan et al. close the loop by examining how AI can be used to support simulation debriefs5.. Reviewing a decade of work on AI in extended reality debriefing, they identify four functions:
Converting XR telemetry into performance analytics
Using NLP to transcribe and code the discourse
Deploying conversational agents to scaffold reflective dialogue
Fusing action, speech, gaze and physiological signals into adaptive feedback.
The evidence base is feasibility studies, pilots and prototypes, with almost no controlled comparisons and no psychometric validation of the AI-derived metrics. They land on a hybrid model in which AI prepares transparent artefacts (transcripts, event highlights, suggested advocacy-inquiry prompts, each with timestamps and quotations) while the human facilitator retains psychological safety and high-stakes interpretation. They also flag that NLP systems mis-transcribe non-native accents, which in most medical schools is not a marginal concern.
So, based on these papers, it seems that AI may be useful in some parts of the simulated encounter. Particularly ones that are observable, repeatable, and countable: generating the case, holding a conversation to a brief, checking whether a step happened, flagging the ninety seconds worth returning to.
It seems AI is poor at the parts of simulation that are relational and situated: ambivalence, non-verbal cues, professionalism, and the collective sense-making that turns an event into learning.
So, presumably if we hand the observable parts to the machine and keep the relational parts for ourselves, we should be honest that we are keeping the half we have the least evidence about. So my suggestion: use these tools to multiply the number of low-stakes repetitions a student can get before they meet a real actor, and keep the actor, the examiner, and the debrief for the encounter that actually counts.
O’Malley, A. (2025). Reflections on Confronting a Capacity Challenge With an AI-Powered Patient Simulator (SimPatient). Simulation in Healthcare: The Journal of the Society for Simulation in Healthcare, 21(3), 209–210. https://doi.org/10.1097/SIH.0000000000000890
Thind BS, Javidi D, Schwartz LM. Artificial intelligence in undergraduate medical education clinical skills curricula: a scoping review of implementations since 2022. Frontiers in Digital Health. 2026;8:1830254. https://doi.org/10.3389/fdgth.2026.1830254
Müller Y, Lemberg H, Gellert P, Zoellick JC. Training Empathetic Communication Skills in Medical Students With a Role-Prompted GPT-4o Chatbot: Quasi-Experimental Intervention Study. JMIR Medical Education. 2026;12:e88092. https://doi.org/10.2196/88092
León-Ariza SA, Orobio-Pinzón MC, Ibáñez-Gutiérrez HM, García-Serna J, Vivas-Giraldo DA. Mapping artificial intelligence integration in objective structured clinical examinations: A scoping review. Medical Teacher. 2026. https://doi.org/10.1080/0142159X.2026.2685794
Khan AS, Hasan S, Ismail FW. Integration of Artificial Intelligence Into Extended Reality Debriefing in Healthcare Simulation: A Narrative Review. Advances in Medical Education and Practice. 2026;17. https://doi.org/10.2147/AMEP.S578857




