How I use AI in AI × MedEd
The answer may surprise you!
Last week Substack took the controversial step of adding AI detection into its user interface. I can see why they did this. They are seeking to deter AI slop from filling their platform and reward writers that produce high-quality original content. But things are never that simple…
False Positives
We have all seen instances of false positive results, when detection software identifies a human-authored piece of work as AI-generated. While it is impossible to test exactly why this happens (AI detection models are as much a “black box” as the models themselves), it seems that detection software simply looks through a piece of writing to assess how perfect it is in spelling and formatting, how regular its structure is, and for the appearance of the fabled ‘—’, ‘:’, “quietly” and “delve”.
With this in mind, AI text detectors can work as a rough signal in controlled settings, but the evidence shows that they are not reliable enough to treat as proof in real-world disputes. A recent systematic review found 80–99% accuracy on in-domain benchmarks, but only 60–75% against adversarial or cross-domain text1, and a stress test showed paraphrasing could cut DetectGPT from 70.3% to 4.6% at a fixed 1% false-positive rate2. A separate large evaluation found detector performance varied wildly, with false-positive rates from 0.05% to 68.6% and false-negative rates from 0.3% to 99.6%3.
Authors Alliance illustrated the absurdity of false positive AI detections in their article last week, when ZeroGPT (a popular detection tool) judged that the book of Genesis was “88.2% AI GPT”, whatever that means. Now, I know many people think AI is overhyped but the word of God it is not.
The practical implication of a false positive is a false accusation. In one study of authentic student essays, both detectors split the same ten human essays into five “AI-generated” and five “Human” labels4, and another audit found ESL writing was disproportionately flagged as AI5.
False Negatives
The opposite is also a problem: when AI detectors classify something as human-authored when it has in fact been created entirely (or partially) by an LLM. It is very difficult to quality the scale of this particular problem, since we can really never know ‘unknown unknowns’. However, I can attest that the phenomenon real.
I published one of my weekly roundups last week, on four colossal meta-analyses of learning effectiveness studies. I’ll share more below about exactly how I use AI in AI × MedEd below, but in this case I used Claude to polish my writing and entirely generate the jobs/grants/events summary at the beginning of the post. Despite this the post scored 100% human.
For an online blog a false negative isn’t necessarily an issue. As educators, however, false negatives are a serious problem. False negatives conceal breaches of academic integrity. But they don’t do it uniformly; they disadvantage students to need more support with their writing (e.g. international students and students with conditions like dyslexia) and advantage students who have well-developed AI skills, or who can afford the more complex models to evade detection. Dr Sam Illingworth of Slow AI focussed on this issue last week.
So AI detection software is useful for triage, especially when text is obviously machine-made or the detector was tuned to a narrow benchmark; it is not dependable as sole evidence for academic integrity, moderation, or legal decisions.
Does anyone even care?
This is disputed and deeply contentious. Some don’t mind if AI generates text wholesale. Many believe that AI can function as a great leveller, especially for people who struggle with English or writing generally. Others believe that as long as the original thought comes from the human author and the text represents the their option or intent, what does it matter if AI generates the actual words? Substack summed up these opinions nicely last week:
AI × MedEd
So how do I use AI to produce AI × MedEd? Probably not as you’d expect! I use AI to build things for this newsletter and to act as a copy editor, but never to write original articles. Decisions about what makes it into the newsletter are always mine, and the style and tone of the posts come from me.
Discovery
One of the most valuable ways AI helps with AI × MedEd is in discovery of new content. An AI agent gets to work every week by scooping up everything published in the medical education (and adjacent) journals that touches on AI. These papers are then screened by me and the agent, and the ones that make the cut are included in a database. This database contains every paper that has featured in this newsletter (currently 260) and is resolved against PubMed. This forms the basis of the AI × MedEd Paper Library.
Similar agents trawl the internet, jobs boards, funder websites and various other places for other jobs and funding opportunities in Medical Education (not just related to AI), and record these in a similar way.
The AI × MedEd Paper Library is now live for paid subscribers, alongside the curated jobs, funding and conference lists. Subscription fees help me pay for the costs associated with running these agents.





Generation
I also use AI to build amazing tools that Medical Educators might find useful in their work, including SBA Studio (which is available for paid subscribers) and SimPatient. These tools are made using research-led design to overcome many of the issues that come with AI, including representation and accuracy. SBA Studio overcomes one of the main issues with LLM-generated questions, which is that the questions are generally too easy!



In terms of my writing, an LLM generates a block of text that summarises new funding calls, job adverts and conferences for the weekly roundup. I usually use the same LLM to proof-read my weekly roundup and suggest improvements, which I either implement or ignore. Sometimes I don’t use an LLM at all.
Academic writing is devoid of personality, and this is especially true for AI-edited academic writing. Hopefully my intros and outros make this newsletter a bit more personal!
Alethary & Aliwy, 2026. A Systematic Review of AI-Generated Text Detection: Approaches, Tools, and Datasets. Al-Salam Journal for Engineering and Technology.
Krishna et al., 2023. Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. Neural Information Processing Systems.
Layton et al., 2026. AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text Detectors. IEEE Symposium on Security and Privacy.
Bendo, 2026. False Positives in AI Writing Detection: A Small-Scale Empirical Study Using Authentic Filipino Student Essays. ASEAN Journal of Open and Distance Learning.
Lege, 2025. Auditing the Fairness of AI-Detection Tools: A Comparative Study of ESL, Published, and AI-Generated Texts and Their Misclassification Risks. International Journal of Teaching, Learning and Education.







