Today I’m writing about three stories that are worth putting together. They didn’t appear in the same week, but they make the same argument three different ways. Epic’s revised sepsis algorithm turns out to predict sepsis hours before clinicians do, but still misfires on the majority of patients it flags. A chatbot posed as a licensed psychiatrist, invalid license number included, until Pennsylvania sued to stop it. And a growing body of research has started documenting residents who never build a skill that AI already performs for them. If you consider them separately, they’re three unrelated complaints about AI in medicine. But, if you consider them together, there is a similar argument in each one: everything about clinical AI today works because a clinician’s judgment is doing real work somewhere in the system. Whether that’s validating a model’s numbers before trusting them, standing between a patient and a bot with no license, or teaching a resident before AI does the work for them, these three stories are what happens when clinical judgement doesn't show up when it's supposed to.
A model that’s better, and a problem that’s worse
Epic released the original Sepsis Model in 2018, and outside validation was not pretty. A 2021 external validation from Michigan found poor discrimination and calibration, partly attributed to using antibiotics as a predictor and comparing against outcome labels that differed across health systems. Epic revised the model in 2022. This year, a multicenter prospective validation in JAMA Network Open tested the update across 227,091 inpatient encounters at four health systems: Michigan, Oregon Health & Science University, Emory, and MetroHealth.
The new model is a real improvement. Researchers set both models to catch the same share of sepsis cases, 60 percent, before comparing everything else, so neither one gets an advantage by being tuned differently. The original model's ability to tell which patients would actually go on to develop sepsis, scored on a 0.5-to-1.0 scale where 0.5 is a coin flip and 1.0 is perfect, ranged from 0.65 to 0.84 depending on the hospital. In practice, most decent clinical prediction models land somewhere between 0.7 and 0.9. The revised model scored 0.82 to 0.92. It beat the original everywhere. It also predicted sepsis ahead of clinicians by a median of 1.4 to 7.1 hours before the first antibiotic order, lactate order, or blood culture that marked clinical recognition.
What didn’t improve is the false alarm rate. At the sensitivity threshold Epic recommends, positive predictive value ran from 0.13 to 0.26 across sites, meaning at the low end, roughly seven of every eight alerts fire on a patient who isn’t developing sepsis. At a 4-hour prediction window, the number needed to evaluate to catch one true case ran from 24 patients at the best-performing site to 69 at the worst. And performance wasn't even consistent from one hospital to the next. On the same scale above, the model scored 0.82 at Michigan and 0.92 at MetroHealth. So a hospital using it has no real way of knowing which end of that range it's going to land on until it tries the model on its own patients.
The authors didn’t conclude with a binary trust or don’t trust decision about the model. They concluded that institutions need to run their own local validation before deployment, build clinical workflows that can absorb the false-positive volume, and use alert-silencing strategies to keep the noise from burying the signal. That focuses on the important question of “does your hospital know what it’s actually getting before it turns the alerts on for every patient in the building.” Most hospitals running this model haven’t done that work, because Epic’s own reported numbers make it look like a plug-in that doesn’t need it.
Somebody still has to sign off on the sepsis model before it reaches a patient: a hospital validating it locally, deciding how to handle its false positives. The next story is about a product nobody signed off on at all.
Where there’s no one to validate anything
Character.AI lets users build custom chatbot “characters.” According to the Commonwealth’s complaint, filed May 1 in Commonwealth Court, a state investigator opened a free account and talked to a character named “Emilie,” described on the platform as a “Doctor of psychiatry.” When the investigator described symptoms of depression, Emilie produced a Pennsylvania medical license number, PS306189, that state authorities confirmed doesn’t exist. The same character separately claimed licensure with the UK’s General Medical Council and prior practice experience in Philadelphia. The Commonwealth is suing under Pennsylvania’s Medical Practice Act (63 P.S. § 422.38) for unauthorized practice of medicine, and asking the court to order Character Technologies to cease and desist. The platform has more than 20 million monthly active users.
The American Psychiatric Association has since issued a formal advisory telling patients and clinicians not to rely on general-purpose chatbots for mental health support, and states have been drafting chatbot safety legislation since.
Compare that to a sepsis model where a hospital has to manage a false-positive rate with workflows and validation. This chatbot has no equivalent safety net at all. There’s no chart, no clinician, no local validation study standing between a person describing depression symptoms and a chatbot that invented a license number to sound credible. The sepsis model needs a clinician’s judgment to interpret its output correctly. Character.AI fabricated a clinician, for a person in one of the most vulnerbale sates.
The sepsis model shows judgment that showed up, just late and at the institutional level. The chatbot shows judgment that didn't show up until a lawsuit forced it. The last story is about whether it shows up at all in the residents training right now.
The skill nobody’s required to have anymore
Medical education researchers have started distinguishing deskilling, which is losing a capability you once had, from never-skilling, which is not acquiring it in the first place. A paper in Nature Medicine this year lays out the distinction directly, along with a third variant, mis-skilling, where a trainee absorbs an AI tool’s errors as fact rather than catching them.
The authors are careful about their own evidentiary limits. They write plainly that direct evidence from medical training is still absent, and ground the concern instead in learning theory and early signals from outside medicine, including data from high school math instruction showing that generative AI use without guardrails can hurt learning outcomes rather than help it. Their proposed fix is a three-phase framework: establish AI-independent baseline competency first, build calibration through structured teaching second, and only then integrate AI use under supervision.
That order is important. A resident who’s reviewed a hundred AI-generated differentials before ever building one independently doesn’t have the same diagnostic reasoning as one who built a hundred from scratch, got a meaningful share of them wrong, and learned from the mistakes.
The sepsis model needs a clinician who can tell a real alert from noise. The chatbot story shows what happens with no clinician there at all. And this story is about what happens to the supply of clinicians whose judgement we need, if training doesn’t change the order AI gets introduced.
What’s the solution?
None of these are arguments against using AI tools.
If your health system runs Epic’s sepsis prediction model, ask your informatics team whether it’s been locally validated against your patient population.
If your health system offers any AI-based mental health resource, know exactly where the human handoff triggers are located before a patient in crisis needs them.
If you supervise residents, pay attention to which of their skills are AI-adjacent and which have become AI-dependent, because the difference won’t show up on a rotation evaluation. By the time it’s visible, the training window that would have fixed it is closed.
Your license says that judgment is supposed to show up. These three stories are what happens when it doesn't.



The order in the third story is the part I would underline. A resident who builds a differential, gets it wrong and sees why is calibrating something no review of AI output can supply: a sense of where their own reasoning tends to fail. It connects to your sepsis section as well. Deciding which alerts are noise is itself a judgment that has to be learned before it can be delegated.