Can AI write the exam?
Four 2026 papers show AI-generated questions now match human ones on paper, but validity, fairness and the human sign-off are where the real work still sits.
Most of the debate about AI in medical education has been about learning: how students revise, how we teach, whether a chatbot helps or hinders. Assessment has had less attention, which is odd, because assessment is where the stakes are highest. A wrong answer in a revision session is a teaching moment. A flawed item in a licensing exam is a fairness problem with a person’s career attached to it. Assessment is also where AI is arriving fastest, drafting questions, marking free text, and feeding back on performance in seconds.
This week I have picked four 2026 papers that map how far AI has moved into assessment, from writing the exam items to running the feedback loop itself. The reassuring news is that AI-generated questions now hold up psychometrically. The less comfortable news is that the evidence for validity, fairness and governance has not kept pace, and the human examiner is still doing the load-bearing work.
Key points:
AI-generated multiple-choice questions now reach difficulty and discrimination levels broadly comparable to human-written items.
The catch is validity: content accuracy, contextual alignment and, above all, evidence about how students actually process the items remain thin.
Almost no study has examined whether AI-generated items are fair across different learner groups.
The most interesting use of AI may not be writing summative exams at all, but supporting formative, assessment-for-learning feedback.
Across all four papers the recommendation is the same: keep a human in the loop, with auditable checkpoints.
Yavuz Selim Kıyak et al. offer the widest view, a systematic review of 71 empirical studies of large language model-generated MCQs across undergraduate, postgraduate and continuing education, most of them testing GPT-4, GPT-4o or GPT-3.51. Their headline is encouraging: AI-generated items frequently reached acceptable difficulty (around 0.62 to 0.78) and discrimination (roughly 0.20 to 0.30).
The qualifier is more sobering. Problems with content accuracy and contextual alignment were reported frequently, and response-process evidence, the validity domain that asks how learners actually interpret an item, appeared in fewer than half the studies (35 of 71) and was almost always indirect. Well-structured prompts consistently improved output quality. The practical message is that AI can draft plausible items at scale, but the validity case is far from complete, and prompt design plus expert review are not optional extras.
Alani et al. narrow the question to a head-to-head comparison, pooling 12 studies (two randomised trials, ten non-randomised) in the first meta-analysis of AI versus human-written questions2. They found no significant difference in either difficulty (SMD 0.05) or discrimination (SMD −0.10): on the numbers, AI items behave like human ones.
Two caveats matter. Human-written questions were harder in medical licensing examinations, and no study reported any equity-focused psychometric analysis, a striking blind spot given how easily language models reproduce the biases in their training data. The authors also quantify the temptation: generation costs can fall below one dollar for 100 to 150 items, with one study estimating a 31-fold time saving. Their recommendation is a staged, auditable pipeline with safety checkpoints (provenance, continuous psychometric surveillance and human gatekeeping), positioning AI to supplement expert item development in lower-stakes settings for now.
Zhao et al. pull back to survey the whole field with a bibliometric mapping study and thematic evidence map of AI for assessment and feedback3. They chart a steep rise in output over the past decade and cluster the literature by theme, and it is telling which concerns organise the map: validity, fairness and governance recur as the field’s central preoccupations.
The value here is orientation. The volume of research is expanding quickly but unevenly, and the community is circling the same unresolved questions that the two reviews above expose in detail. It is a useful guide to where the gaps, and the next studies, should be.
Nguyen et al. reframe the entire discussion. The point of AI in assessment, they argue, may not be the summative exam at all. They define AI-assisted formative assessment as the intentional use of AI to generate, organise and support the interpretation of performance information for learning rather than grading.
Students already use dashboards, large language models and emerging agentic tools to rehearse clinical reasoning, communication and procedural skills and receive individualised feedback within seconds, and the authors work through concrete examples across preclinical case-based learning, OSCE and communication practice, procedural skills laboratories, clerkships and programmatic assessment portfolios.
They are clear about the risks: hallucination, automation bias, epistemic overtrust, hidden curricular effects, and wider concerns about professional identity, data privacy and unequal access. The implication is that the richest opportunity is low-stakes and formative, where fast feedback compounds, provided we design deliberately against ‘over-trust'.
A pattern runs through all four papers. AI has closed the gap on the mechanical part of assessment, generating items that behave like human ones and returning feedback faster than any tutor could manage. What it has not closed is the gap on the parts that make assessment trustworthy: content accuracy, response-process validity, fairness across groups, and clear governance.
For practising educators the useful question is no longer whether we can use AI to assess, but where on the stakes ladder it is safe to. Right now the answer points downward, toward formative, feedback-rich, low-stakes uses where a hallucinated distractor is a learning point rather than a career event, and upward only with the kind of auditable human checkpoints Alani and colleagues describe.
Assessment was always the hardest thing to get right in medical education. AI has not changed that. It has raised the speed at which we can get it wrong.
Are you looking to progress?
I am now tracking over £11m of medical education grants and 18 jobs over on the Career Intelligence pages, which are availably for AI × MedEd Members.
Kıyak YS, Kaya AB, Emekli E. Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal, 2026. https://academic.oup.com/pmj/advance-article/doi/10.1093/postmj/qgag057/8688271
Alani NHS, Amaranathan J, Dandoush A, Rye S, King PT, Chin MH. Can large language models generate exam questions comparable to humans? A systematic review and meta-analysis study in medical education. Medical Teacher, 2026. https://www.tandfonline.com/doi/full/10.1080/0142159X.2026.2691072
Zhao Z, Liu Z, Guo L, et al. Artificial Intelligence for Assessment and Feedback in Medical Education: Bibliometric Mapping Study and Thematic Evidence Map. JMIR Medical Education, 2026. https://mededu.jmir.org/2026/1/e98949




