I’ve written a lot about AI in Medicine and how we must be careful that its output is validated before applying it in any clinical setting. There are so many cautionary tales that it may seem like we should just throw it away and make medicine an AI-free zone. But today I’d like to highlight a potentially positive use for AI in medical research. Let’s get into the details.
As always, if you enjoy reading this newsletter, subscribe and tell a friend.
Sam
A lot of my writing and discussion in medicine in general has focused on the danger of artificial intelligence. I’ve discussed AI hallucinations, incorrect conclusions, concern over physician overreliance, data privacy and consent, and more.
But today, I want to discuss one of the positive aspects of combining AI and medical research.
Medicine as a research source
As I wrote about in my last article, medical data storehouses are increasing in number.
Hospital electronic health records, insurance company claims databases, government datasets, and clinical trial data are just a few examples. Recently, consumer-facing products have begun building their own databases of patient information. Some notable examples: Apple Research, Google Health Connect, Evidation. Users are able to upload their health records, lab records, imaging, and even connect their wearables. That data is being voluntarily given to large companies, which can build their own databases.
Historically, researchers have accessed these databases to ask new questions and find correlations between therapies and specific populations. This type of research fits into a category called “secondary data analysis”. The data already exists, and researchers are using it to answer questions.
The diagram below shows just how many different kinds of research studies can be derived from secondary data analyses.
Secondary data analyses aren’t just limited to observational studies. Researchers regularly use data from previously published randomized controlled trials to look for new outcomes, perform subgroup analysis, evaluate prognostic relationships, and answer questions that were not addressed in the original publication.
Oftentimes people think of secondary data analysis as a lesser form of research. But that’s not the case at all.
Influential medical research uses existing data
You don’t have to take my word for it. Look at any major medical journal, and you can find numerous examples.
The National Health and Nutrition Examination Survey (NHANES) produced important estimates of obesity and chronic kidney disease in the United States.
A 2009 study in the New England Journal of Medicine looked at data from almost 12 million Medicare beneficiaries, with a focus on hospital readmissions and their costs. That publication led to a national discussion surrounding hospital readmission.
Medicaid data was used to investigate whether atypical antipsychotic medications were associated with sudden cardiac death.
Medicare databases have also been used to examine the effectiveness and safety of multiple anticoagulants in patients with atrial fibrillation.
SEER-Medicare data were used to investigate differences in breast cancer outcomes between Black and White women.
Secondary analysis of data from major randomized controlled studies like the SPRINT and FOURIER trials has answered additional questions.
All across the spectrum of medical specialties, secondary analyses have contributed significantly to our medical knowledge.
Researchers don’t need to start over every time
A single medical database can support many different questions
With access to the medical histories of millions of patients, the number and types of questions are seemingly unlimited.
You could ask a descriptive question like “How common is chronic kidney disease among patients with diabetes?”
You could ask a comparative effectiveness question like “Among patients with atrial fibrillation, how do outcomes differ between patients receiving two commonly used anticoagulants?”
You could investigate safety by asking “Are patients who are receiving a particular medication more likely to develop acute kidney injury?”
You could study prognosis by asking “Which characteristics predict readmission following hospitalization for heart failure?”
You could examine disparities by asking “What percentage of patients in different demographic or geographic populations receive guideline-recommended treatment?”
You could study changes over time by asking “Does an FDA-issued safety warning change prescribing behavior?”
Each of these questions can be asked of the same database and result in truly significant findings.
The barriers
Physicians frequently come across unanswered questions while treating patients.
A cardiologist may notice a certain group of patients seems to respond unusually well to a treatment.
An emergency physician may wonder why the same diagnosis causes some patients to return to the emergency department while others don’t.
A primary care physician may notice an unreported adverse effect of a medication.
An oncologist may wonder if a treatment seems to perform differently in a particular patient population.
These questions and observations are the foundation of medical research. But access to what’s needed to perform that research is not the same across the spectrum of physicians.
Turning that question into a research project requires numerous steps like:
Defining a study population, inclusion and exclusion criteria, and exposures and outcomes.
Understanding the structure of a database, how to write queries and extract data
Choosing a statistical approach, identifying confounders, and performing sensitivity analyses
Until now, those steps required a team of epidemiologists, biostatisticians, database specialists, and programmers. Although those people provide very valuable expertise, they are all scarce resources. A physician working at a large academic medical center may have access to that team. A community physician with an interesting clinical observation is more likely to just let it slip away.
AI may change all of that.
The distance between question and analysis
Imagine you are a hospitalist and notice that patients you discharge on drug A don’t bounce back as often as patients on drug B. You check your favorite clinical reference, but there is no published data comparing the two drugs. AI has the potential to assist in the research needed to answer the question: “Among patients with condition X who initiated Drug A versus Drug B, was there a difference in hospitalization within 12 months?”
Based on your question, AI might help identify a study design like a retrospective cohort study.
It could help define the cohort by asking questions like: Which patients should be included or excluded? How should initiation of the medication be defined? How much past medical history is required?
It could help define a comparator group and outcome.
It could identify possible confounders.
It could help identify an existing medical database ideally suited for your needs.
It could help generate a database query.
It could assist with statistical analysis and sensitivity testing.
It could help interpret the results and identify weaknesses in the analysis.
The process is familiar because it is the same one medical researchers perform routinely. But the ability for a practicing community physician to complete that process has now changed.
Clinical observation → research question → study design → cohort definition → database query → analysis → evidence
AI could assist with each step.
The catch
Making an analysis easier to perform doesn’t automatically make the analysis trustworthy. Medical databases can be messy.
Patients are not randomly assigned to treatment groups in routine clinical practice.
Diagnoses can be coded incorrectly.
Outcomes may be missing.
Patients can leave a healthcare system and disappear from the database, leaving no follow-up data.
Laboratory testing occurs more frequently in sicker patients.
Medication records show what was prescribed, not what was actually taken.
Researchers can inadvertently define populations that bias their results.
And if enough variables are included, significant associations can emerge just by chance. AI doesn’t make any of those problems disappear. In fact, AI could compound the problem by allowing a poorly designed analysis to be performed with extraordinary speed.
That means physicians using these tools will still need to understand concepts like confounding, selection bias, misclassification, missing data, multiple comparisons, causal inference, statistical validity, and reproducibility.
Correct research methodology, biostatisticians, epidemiologists, and expert review will all remain essential. The opportunity comes from allowing those resources to function differently. Instead of every research question requiring extensive technical work before it can even be explored, AI may allow physicians to perform preliminary analyses, test feasibility, refine hypotheses, and identify questions that deserve deeper methodological involvement.
AI and the research cycle
One last area AI can impact.
Secondary data analysis often occurs before prospective research.
Suppose an analysis of five million patient records identifies an unexpected association between a medication and a clinical outcome. That doesn’t prove the medication caused the outcome. But it does give researchers a valuable signal that may justify an observational study, which in turn may lead to a prospective study. And eventually, a randomized controlled trial.
The same process can work in reverse. Researchers can use completed randomized trials and ask additional questions of datasets that cost millions of dollars and years of effort to create.
Here, AI has the potential to accelerate multiple stages of the research cycle:
Hypothesis generation.
Feasibility assessment.
Safety-signal detection.
Comparative effectiveness research.
Identification of high-risk populations.
Subgroup exploration.
Study design.
And identification of questions worthy of prospective trials. The acceleration allows more questions to be investigated.
Should we broaden the conversation around AI?
We have to establish boundaries around the use of artificial intelligence in medical research. We need standards for privacy, validation, reproducibility, transparency, methodology, and human accountability just as we have non-AI-related standards for all these areas now.
These safeguards are important precisely because these tools may become extraordinarily capable.
Medicine has enormous datasets describing diseases, treatments, complications, and outcomes across millions of patients. The scientific opportunity contained in that data is enormous. AI may give us a new way to interrogate it.



