Who Gets to Decide When AI Is Ready to Perform Medical Work?
Utah’s AI Pilot Exposed a Governance Question That Medicine Has Never Had to Answer
I’ve been thinking about the Utah AI prescription pilot since it began in April. Medicine spent a century building a process for deciding when a physician is ready to practice. Nobody has built the equivalent process for software. The Utah experiment didn’t answer that question. It just made clear how badly we need one. Let’s talk about it in detail.
As always, if you enjoy reading, subscribe and tell a friend.
Sam
Who gets to decide that software is ready to perform work that has historically required a medical license?
Medicine has spent more than a century building institutions that determine when physicians are competent to care for patients. Medical schools educate them. Residency programs train them. Specialty boards certify them. State medical boards license them. Hospitals credential them. Peer review evaluates them. Courts review their actions when patients are harmed.
Those institutions are not perfect, but together they establish something fundamental. Society has created a process for determining when a physician is competent enough to practice medicine.
Artificial intelligence introduces a new challenge because software is beginning to perform tasks that have traditionally depended on physician judgment. I am not talking about whether AI will participate in medicine. I’m adking who has jurisdiction to determine when the evidence is sufficient for software to assume greater clinical responsibility.
Earlier this year, Utah became one of the first states to confront that question directly. The debate that followed was widely portrayed as one about artificial intelligence and prescription renewals. I believe it was fundamentally a debate about jurisdiction over medical competence.
The Utah Pilot Structure
In December 2025, Utah’s Office of Artificial Intelligence Policy authorized a health technology company called Doctronic to begin renewing prescriptions for Utah patients using an AI system, under a 12-month regulatory sandbox agreement announced January 6, 2026.
The Office of Artificial Intelligence Policy was created in 2024 under Utah’s AI Policy Act, SB 149, which gives the office authority to temporarily waive regulatory requirements so companies can test AI systems under state supervision. Utah has used that same authority for a mental health chatbot aimed at teenagers and for an AI tool that reads dental radiographs. Doctronic’s agreement lets its AI process 30-, 60-, or 90-day renewals for medications that a licensed physician has already prescribed, and screen for drug interactions along the way.
The pilot is structured in three phases. In Phase One, every AI-generated renewal is reviewed and approved by a licensed physician before it reaches a pharmacy. Phase Two would move that review to shortly after a prescription is issued. Phase Three would allow the system to renew prescriptions with only periodic sampling of AI outputs for oversight. State officials say the pilot will not advance to the next phase until safety has been demonstrated in the one before it.
Before the pilot launched, Doctronic shared data with state regulators comparing its AI’s recommendations to physicians’ decisions across 500 urgent care cases. The company reported that its recommendations matched physicians’ 99.2 percent of the time.
Medication renewals consume substantial physician time, and many involve stable chronic conditions with predictable decision pathways. Exploring whether AI can safely assist physicians with those tasks is a reasonable objective.
The Utah Pilot Significance
On April 20, 2026, eleven of the fourteen members of the Utah Medical Licensing Board sent a letter to the Office of Artificial Intelligence Policy. They wrote that the board “was made aware of this agreement only after its implementation, once the system was already live and available for use.” The letter argued that prescription refills require reassessment of dose, side effects, contraindications, and drug interactions, work the board said belongs to a licensed physician, and it warned that “patients who continue refilling medications without assessment may remain on outdated or suboptimal therapy for months or years.” It closed by recommending that the pilot “be immediately suspended pending further discussion.”
The next day, the directors of the Office of Artificial Intelligence Policy and the Division of Professional Licensing responded in writing. They said the pilot had been “rigorously reviewed by several medical professionals prior to launch,” a process that produced “a large number of suggested substantive adjustments and guardrails,” and that the state would not suspend the pilot because it remained in Phase One, where a physician reviews every renewal before it is filled. They committed to involving the board more closely going forward.
Reasonable people can disagree about whether that was the correct decision. The more enduring question concerns the process itself. Which institution should determine that the available evidence justifies introducing software into a clinical role that has traditionally required physician judgment?
Evidence and Jurisdiction Cannot Be Separated
Much of the public debate over the Utah pilot focused on the AI itself: was it accurate, was physician oversight sufficient, should refill decisions be delegated to software at all.
Those questions naturally followed once the pilot was underway. They also assume that someone had already concluded the available evidence justified launching it. Let’s take a closer look.
The Doctronic pilot shows how the same evidence can support different conclusions. A 99.2 percent concordance rate across 500 urgent care cases was sufficient for the state and the medical professionals who reviewed the pilot before launch. It was not sufficient for the Medical Licensing Board, whose objection was not about the AI’s accuracy in aggregate. It was about what a refill requires in every individual case: a reassessment that a benchmark conducted before launch cannot fully substitute for once the system is operating on real patients. Both readings of the evidence are defensible. They come from institutions with different responsibilities.
The institution responsible for authorizing clinical AI also determines what counts as convincing evidence. A software engineer may prioritize benchmark performance and operational reliability. A practicing physician may focus on uncommon but consequential clinical presentations. A statistician may emphasize study design and external validity. A regulator may concentrate on statutory authority, while a hospital executive may view liability and implementation as equally important.
Each perspective is legitimate. Each emphasizes different forms of evidence. The authority that evaluates the evidence decides the standard by which readiness is judged.
Medicine Already Has a Model
Imagine a hospital announced that it had created a pathway allowing physicians to practice emergency medicine without residency training. The first question would not concern examination scores or clinical outcomes. It would concern authority. Who approved this process?
Medicine has developed a distributed system for answering that question. Medical education is accredited. Residency programs are supervised. Licensing examinations are standardized. Board certification is independently administered. Hospitals grant privileges only after reviewing credentials and competence.
No single organization controls that process. Government, professional organizations, accrediting bodies, and hospitals all participate. Each contributes a different perspective before a physician is entrusted with patient care.
Clinical AI is developing without an equally mature framework for determining when software is ready to assume comparable responsibilities.
The Same Evidence, Different Conclusions
One reason these discussions become contentious is that stakeholders evaluate evidence through different lenses.
Software developers ask whether the model performs accurately. Researchers ask whether a study demonstrates benefit. Medical boards ask whether patient safety has been adequately protected. Attorneys ask who assumes liability when errors occur. Patients ask whether they can trust the recommendations being made.
These are not competing questions, but complementary ones.
Evidence never speaks for itself. People decide what the evidence is sufficient to support.
A Different Way to Begin
Imagine Utah had started somewhere else. Before enrolling the first patient, the state convenes a multidisciplinary commission composed of primary care physicians, pharmacists, medical licensing officials, software engineers, biostatisticians, health services researchers, patient representatives, ethicists, and regulators. Their first task is not evaluating the AI, but instead defining the rules.
What evidence will be required? How should safety be measured? What outcomes matter? What level of physician oversight is necessary? Who has authority to suspend the pilot if safety concerns emerge?
Only after those questions are answered does the commission evaluate the software.
That sequence changes the discussion. The debate isn’t about defending or criticizing artificial intelligence. It is about establishing a legitimate process before patients become part of the evaluation.
Why Utah Matters
Prescription renewals are unlikely to be the last area where this issue appears.
Utah has already used the same regulatory sandbox authority outside prescribing: for a mental health chatbot aimed at teenagers and for an AI tool that reads dental radiographs. The same questions will emerge in radiology, pathology, dermatology, emergency medicine, imaging interpretation, treatment recommendations, autonomous documentation, and other clinical applications as AI capabilities continue to expand.
Every new application will generate studies, benchmark results, and performance claims. Those discussions remain essential. They do not eliminate the need for an accepted process that determines when the available evidence justifies broader clinical use.
The American Medical Association has already staked out a position on where that leaves physicians. Responding to the Utah pilot, AMA chief executive Dr. John Whyte said that “while AI has limitless opportunity to transform medicine for the better, without physician input it also poses serious risks to patients and physicians alike.”
Without an agreed framework, different organizations will continue reaching different conclusions from similar evidence because they are applying different standards and serving different responsibilities.
The Question Utah Leaves Behind
The public conversation often frames medical AI as a choice between accelerating innovation and slowing it down. Utah suggests the more important issue comes earlier.
Before software performs work that has historically required a medical license, society should determine who has jurisdiction to evaluate the evidence, what standards that evidence must satisfy, and how disagreements between institutions will be resolved.
The Utah pilot may ultimately prove to be an important success. It may demonstrate that AI can safely improve efficiency in routine clinical care. Regardless of its eventual outcome, it has already exposed a governance question that will become increasingly difficult to ignore.
Medicine has long established how society decides that physicians are ready to care for patients. Artificial intelligence now requires an equally thoughtful process for deciding when software is ready to assume comparable responsibilities. Utah did not answer that question. It demonstrated that we need one.


