The Two Pipelines
One hides your name. The other doesn't bother.
In this installment of Ashoo Review: AI in Medicine, I’m taking a closer look at how our data is being sold to software companies AND how those same companies are now asking for us to provide it voluntarily. You may think your consent is required; think again.
As always, if you enjoy reading, consider subscribing and tell a friend.
Sam
In July 2026, OpenAI rolled out ChatGPT Health to every U.S. adult user, inviting them to connect their Apple Health data, medical records from Epic and Oracle Health systems, and wellness apps directly to the AI. OpenAI announced the nationwide launch on July 23, promising “secure” connections and personalized health insights. The move was widely covered as a consumer convenience, a smarter way to track medications, interpret lab results, and prepare for doctor visits.
What most coverage missed is that this is only one half of a much larger story. While consumers debate whether to trust OpenAI with their blood pressure readings, a separate, invisible pipeline has already moved hundreds of millions of patient records into the hands of AI companies. The difference is structural, legal, and deeply asymmetrical, the back-door pipeline strips your name before it sells your data. The front-door pipeline does not have to.
The Front Door, HIPAA Ends Here
Here is the first thing to understand about ChatGPT Health, it is not covered by HIPAA.
HIPAA, the Health Insurance Portability and Accountability Act, governs a narrow class of “covered entities”, hospitals, health insurers, and healthcare clearinghouses, plus their “business associates.” OpenAI is none of these. OpenAI states explicitly that “Health in ChatGPT is not intended for clinical or covered-entity use and does not offer a Business Associate Agreement.”
What this means in practice is stark. When a hospital holds your medical record, federal law requires them to protect it, limit its use, and obtain authorization before sharing it broadly. When you upload that same record to ChatGPT, those protections evaporate. As health platform b.well noted, “medical records uploaded into ChatGPT Health are no longer covered by HIPAA, because the service isn’t a covered entity under that law.”
OpenAI does not need to de-identify your data. It does not need to strip your name, your birth date, your medical record number, or your diagnostic codes. Its only constraints are its own privacy policies. OpenAI’s Health Privacy Notice says that by default, health data received via Health Features is not used to improve foundational models. But the policy also notes that “a limited number of authorized OpenAI personnel and trusted service providers might access data... to improve model safety, unless you have opted out.” That is a significant carve-out, governed by a terms-of-service agreement most users will not read.
The consumer is being asked to make a privacy decision without understanding the legal terrain. They assume HIPAA protects their health data everywhere. It does not.
The Back Door, The Pipeline You Never Signed Up For
While OpenAI courts consumers on the front end, a separate industry has built a lucrative infrastructure for selling hospital data to AI companies on the back end. The key players are data middlemen, companies that contract with hospitals, strip patient records of direct identifiers, and license the resulting datasets to AI developers and researchers.
Truveta is among the largest. Founded by a consortium of major U.S. health systems, Truveta now includes data from more than 130 million de-identified patients across over 900 hospitals and 20,000 clinics. Its members include Providence, Advocate Health, Trinity Health, Tenet Healthcare, Northwell Health, AdventHealth, and Novant Health. In September 2021, Microsoft announced a strategic partnership and investment in Truveta, and Truveta uses Azure-based AI to process its data. Truveta has also launched its own Truveta Language Model trained on this clinical corpus.
Protege is the emerging contender. In January 2026, Protege raised a $30 million Series A led by Andreessen Horowitz, bringing its total funding to $65 million. The company works with what it describes as “close to 20 data partners” in healthcare and more than 100 across all sectors. Protege licenses what it calls “private, real-world data” including clinical notes, medical images, video, audio, pathology slides, genomic data, lab reports, wearables data, and social determinants of health. In a partnership with Syndesis Health, Protege gained access to 70 million de-identified patient lives from more than 500 facilities across 15 countries.
Protege states on its website that it works directly with “leading foundation model labs” to define licensing standards. It does not name them. What is clear is that the pipeline is vast, well-funded, and growing rapidly.
The Asymmetry
Consider what this means for a single patient.
If you are treated at a Truveta-member hospital, your de-identified record may be harmonized, aggregated, and sold to AI developers. The hospital has removed your name, address, medical record number, and the other 15 identifiers required by HIPAA’s Safe Harbor method. The AI company receives a rich longitudinal record of your diagnoses, medications, procedures, and lab values, but without a name attached.
If you then upload your own medical records to ChatGPT Health to ask about your medication list, OpenAI receives the same clinical information, but with your name on it, your dates intact, and your full identifying context, because HIPAA does not apply to them.
The hospital is legally required to de-identify your data before sharing it. The consumer AI company is not required to do anything of that kind. The regulated system forces anonymization on the back door while the consumer app collects richer, more identifiable data through the front door with no federal health privacy guardrails at all.
Here is the part that should give consumers pause. The software companies building these health AI products may already have more data on you than you would want them to know, de-identified records from your hospital, imaging from your health system, lab values from your clinic. If you then provide them with a small sampling of identifiable data, your own uploaded records, your Apple Health metrics, a few lab results, you may inadvertently give them the key to re-identify what they already possess as belonging to you. Careful what you share, because even a narrow slice of identifiable health data can be used to frame a much more detailed picture of you than you may think.
The De-identification Myth
There is a second problem with the back-door pipeline, and it is mathematical.
For decades, the legal and ethical framework for health data sharing has relied on a simple premise, remove 18 identifiers, and the data is no longer “personally identifiable.” Once de-identified, it falls outside HIPAA’s scope and can be “freely used, shared, and sold” without patient consent.
Modern AI has broken that premise. In 2019, researchers Luc Rocher, Julien Hendrickx, and Yves-Alexandre de Montjoye published a study in Nature Communications titled “Estimating the success of re-identifications in incomplete datasets using generative models.” Their finding was startling, “Using our model, we find that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes.”
The authors were explicit about the policy implications, “Our results suggest that even heavily sampled anonymized datasets are unlikely to satisfy the modern standards for anonymization set forth by GDPR and seriously challenge the technical and legal adequacy of the de-identification release-and-forget model.”
Other studies have reinforced the point. Researchers have demonstrated that algorithms can accurately match physical activity data and demographic information to 95% of adults in a dataset. As Scripps News reported, “Health data shared with AI systems can be stripped of names and sold legally, then re-identified using AI tools.”
The data middlemen are selling records that are legally de-identified but practically re-identifiable. The “without your name” framing is a legal fiction that AI has rendered increasingly untrue.
The Consent Theater
If the pipeline is invisible, how do patients find out about it? In most cases, they don’t.
When you check into a hospital, you sign a group of forms and receive a Notice of Privacy Practices, a long document that mentions treatment, payment, and “health care operations.” De-identification and data sharing may be mentioned in a single clause. The notice almost never says, “Your records may be sold to an AI training data broker.”
Some hospitals are becoming more explicit. Archbold Medical Center in Georgia publishes a Notice of Privacy Practices that states, in plain language, “De-identified data is no longer subject to privacy or security laws.” It also notes that “AI use within applications uses/discloses your medical information.” This is unusually candid. Most notices are not.
The legal framework creates what might be called consent theater, patients are given a document they do not read, which authorizes broad data uses they do not understand, for purposes, including AI model training, that were not contemplated when HIPAA was written in 1996.
The Regulatory Vacuum
One might expect that the wave of state consumer privacy laws would close this gap. They largely do not.
Washington’s My Health My Data Act (MHMDA), enacted in 2023, broadly defines “consumer health data” and requires opt-in consent for collection and sale. But the Act explicitly excludes “deidentified data” from its definition of personal information. A hospital selling de-identified records to Protege falls outside the law.
Maryland’s Online Data Privacy Act (MODPA), effective October 2025, has some restrictions on using consumer health data for AI model development. But properly de-identified or aggregated data is largely exempt.
Nevada’s SB 370 and Connecticut’s CDPA amendments follow the same pattern, they regulate identifiable consumer health data but leave the de-identified pipeline untouched.
At the federal level, the FTC has been active on data broker enforcement, but not on this specific subject. In February 2026, the FTC sent warning letters to 13 data brokers about compliance with the Protecting Americans’ Data from Foreign Adversaries Act (PADFAA), focusing on foreign data transfers rather than domestic AI training. There is no federal enforcement action specifically targeting the sale of de-identified health records for commercial AI model training.
The Convergence
The two pipelines are separate in law but convergent in effect. OpenAI is building the infrastructure to ingest medical data on both sides.
On the front end, OpenAI acquired Torch, a health records unification startup, for roughly $60-100 million in January 2026. Torch’s technology consolidates fragmented patient data from hospitals, labs, and visit recordings. The stated goal, per HLTH reporting, is a “unified medical memory” for ChatGPT Health.
On the back end, OpenAI’s health data ambitions are clear even if its purchases from middlemen are not publicly documented. The company is positioning itself to be a central platform for health information, whether that data arrives through consumer upload or hospital pipeline.
The patient, meanwhile, is left with a false sense of security. They are told HIPAA protects their medical records. They are told de-identification makes their data anonymous. They are told they can “securely” connect their health records to an AI assistant. None of these statements are entirely false. But none of them capture the full picture of where their data is going, how it is being used, or how thin the legal protections actually are.
The health data economy has bifurcated into two tracks, one regulated and de-identified, the other unregulated and fully identifiable. Both are feeding the same models. And the patient is the last to know.


