This Week in MedEd
You may have noticed that there was no newsletter last week. I completely forgot to post anything because I was enjoying AMEE 2026 in Vienna! However, I have continued tracking opportunities, including 9 open funding calls, 35 upcoming events, 34 live jobs (11 new, 9 closing this week), including:
💰 Connecting Champions for Diversity, Equity, Inclusion, and Belonging in the Health Professions (up to $2M from the Robert Wood Johnson Foundation, deadline 24 Sep 2026)
🎓 AIiH 2026, the International Conference on AI in Healthcare (London, 26-28 Aug 2026)
💼 Associate Dean, Artificial Intelligence, College of Medicine (University of Arizona, Tucson)
The full list, with deadlines, links and eligibility, is one tap away. Try everything free for 30 days.
This week the newsletter returns with a more worrying question than usual: what happens to a clinician’s or student doctor’s own reasoning when a confident machine is always available to step in? It’s difficult to measure, it takes years to detect, and in the meantime everything appears fine and there is little impetus to change our practice as educators.
I’ve selected four recent papers that all summarise the same problem, which is that that primary risk may not be is not accuracy or robustness (i.e. how models get things wrong), but that using the models could change how we think and learn.
Key points
A systematic review of 29 studies found automation bias reported in 10 of them and clinician deskilling in 9, though its author is candid that direct longitudinal evidence of sustained deskilling is still missing.
The nearest thing to a hard number: endoscopists showed reductions of up to 20% in unassisted colonoscopy diagnostic accuracy following AI adoption.
The proposed mechanism is gradual rather than sudden, a slide from cognitive augmentation through “epistemic vulnerability” to outright epistemic dependency, driven by repeated small acts of delegation.
What erodes is disposition rather than knowledge: tolerance of ambiguity, reflective scepticism, the habit of monitoring your own reasoning. None of which our assessments measure.
All four papers land on the same cheap intervention: make learners commit to a judgment before they are allowed to consult the machine.
Lets start with a systematic review1. Al-Anezi et al screened 11,269 studies down to 29 published 2022 - 2025. Three themes recur:
Automation bias
Clinician deskilling
Effects on diagnostic reasoning.
The concrete findings are borrowed from the included studies rather than pooled, and they cut both ways. Two of the studies report efficiency improvements (up to 45%) on diagnostic and prognostic tasks, and one found a reduction in medication-dose errors (up to 88%). However others found AI use led to automation bias, including a reduction of up to 20% in unassisted colonoscopy accuracy following AI adoption.
One hypothesis that emerged in the review was of a "self-referential learning loop", in which reliance on AI both degrades clinician skill and contaminates the data used to retrain the model.
There is no direct longitudinal evidence demonstrating clinician deskilling, and this review was undertaken by a single author with no prospective protocol.
Botello Jaimes et al. supply the theory the review is missing2. They argue the relationship between genAI and learning is best understood as a continuum running from cognitive augmentation, through what they call “epistemic vulnerability”, to “epistemic dependency”. Here’s how they think this slide happens:
"Epistemic vulnerability rarely emerges through abrupt replacement of human cognition. The process is more subtle and usually develops through repeated patterns of cognitive delegation."
Two properties of language models drive it:
They produce outputs that "frequently display the linguistic attributes of expertise, even under conditions of incomplete, uncertain, or erroneous knowledge"
They deliver immediate interpretive closure precisely in the situations that used to demand sustained analytical effort.
The result they worry about is "inferential fragility", where visible success coexists with a diminished capacity for prolonged reasoning. Their table of seven educational domains is the most immediately usable thing I read this week.

Ahmady et al. point out that the thing we are integrating is changing underneath us3. Their perspective distinguishes conventional language models, which "function primarily as reactive text generators based on pattern recognition", from AI agents, which add memory, planning, goal-directed action and integration with external tools such as clinical databases, guidelines and simulation environments.
I think this is the first time I've read someone distinguish between pedagogical consequence of the two:
A chatbot's characteristic risk is hallucination, which at least gives learners something to catch
An agent's characteristic risk is "reduction in independent reasoning and diagnostic engagement if core cognitive tasks are fully delegated".
I’m not sure it’s quite as simple as this (for example, I’m sure chatbots can diminish the act of learning and agents can hallucinate so it is presented as a bit of a false dichotomy), but the principle is sound.
They frame what is at stake as phronesis, practical wisdom, and propose a four-part response: mandated faculty development alongside student AI literacy, deliberately failure-based simulation4, agents positioned as cognitive partners that learners must argue with, and assessment that treats the ability to interpret and validate AI output as a competency in its own right.
Its worth noting that this is philosophical argument, with no empirical evidence cited.
Teng et al. consider a much more clinically-applied context5. Their central observation concerns AI harm in clinical environment already saturated with it, where AI predicts outcomes, summarises charts, drafts letters etc. They point out that "harms may appear less as wrong numbers and more as distortions of context: omissions, overconfident summaries, or misframed trends".
I think this echoes what the distinction between simple technical hallucinations and more problematic agentic misalignment we mentioned above. The authors say "the most dangerous failures may be silent: an AI-generated sign-out that sounds coherent yet misses a rising vasopressor requirement".
Their answer is to build adaptive expertise rather than fixed competencies, and to prioritise verification over technical fluency: "Clinicians do not need to code, but must understand common failure modes: bias, hallucinations, omissions, and calibration drift." The provide two practical examples:
The AI sign-out drill, in which a team reviews an AI-generated ICU summary and must link every key claim to source data, identify what is missing, articulate uncertainty, and decide what would trigger a stop-the-line escalation.
A score audit, in which learners take a familiar tool such as SOFA or a sepsis trigger and trace its development, assumptions and performance variation across sites.
I think it’s important to note that three of these four papers are argument rather than evidence. We discussed the glut of opinions and frameworks a few weeks ago:
Anyway, full credit to Al-Anezi for proposing longitudinal studies to measure time-to-decision without AI access and retention of differential diagnosis accuracy. And the new thinking around the pedagogical consequence of chatbot use vs agent use is also interesting and I'm sure will lead to further discussion here in future. The other recommendations are maybe useful, and maybe not. I’ll leave it to you to decide.
Al-Anezi FM. Generative Artificial Intelligence in Healthcare: Automation Bias, Deskilling, and Cognitive Implications - A Systematic Review. Journal of Healthcare Leadership. 2026;18:1-13. https://doi.org/10.2147/JHL.S590498
Botello Jaimes JJ, Turriago Castañeda AK, Alzate Mejía OA, Lombana Jimenez AE. Generative artificial intelligence and the epistemology of medical learning: from cognitive augmentation to epistemic dependency. Frontiers in Education. 2026;11:1894187. https://doi.org/10.3389/feduc.2026.1894187
Ahmady S, Kohan N, Monajemi A. AI Agents and the Future of Clinical Judgment in Medical Education: Opportunities, Challenges, and the Need for Human-Centered Integration. Health Science Reports. 2026;9(8):e73042. https://doi.org/10.1002/hsr2.73042
I love the idea of simulations that are specifically designed to fail, like the Kobayashi Maru in Star Trek, which inspired this paper published in print earlier this week: Wong, K. S. S. (2025). A ‘Star Trek’-inspired, doomed-to-fail simulation for teaching coping with physician grief. Medical Education, 60(8), 930–931. https://doi.org/10.1111/medu.70161
Teng MM, Tsuei SH, Koblanski M, Celi LA. Cultivating adaptive expertise in critical care: reimagining intensive care unit education in the era of Artificial Intelligence. Critical Care Science. 2026;38:e20260040. https://doi.org/10.62675/2965-2774.20260040




