AI Has Already Crawled Your Data. Here's Why Deleting It May Not Be Enough

Introduction
For the last few years, "delete my data" has meant something specific. You find the data broker, you submit a request, and eventually, sometimes grudgingly, your record disappears from that site. It's not a perfect system, but it's a system. Laws made it one.
California's Delete Act is the clearest example. As of 2026, a single request through the state's new deletion platform reaches every registered data broker at once. Brokers have to check it every 45 days. If your information is in their database, they have to remove it, inferences included, or face a fine of $200 per request, per day, until they comply. Europe has its own version through GDPR's "right to erasure," which for two decades has given people a real, enforceable claim to have their personal data removed from a company's systems.
It took years of lawsuits, regulatory pressure, and bad press to build that system. And it only applies to one kind of company: one that stores your data in a database it can query, search, and delete from.
AI doesn't store your data that way. That's the part almost nobody has fully priced in yet.
Is your cellphone vulnerable to SIM Swap? Get a FREE scan now!
Please ensure your number is in the correct format.
Valid for US numbers only!
The Deletion Promise Was Built for a World That's Already Changing
A data broker's business model depends on holding a legible, searchable record: your name tied to your address, your phone number, your relatives, your purchase history. That legibility is exactly what makes deletion possible. You can find the row, and you can delete the row.
A large language model doesn't have rows. When a model is trained, the text, images, and records it ingests are compressed into billions of numerical parameters, statistical relationships rather than a filing cabinet. Regulators in Europe have now formally stated that this doesn't matter legally: the EDPB's mid-2026 guidance confirmed that GDPR applies in full to personal data scraped for AI training, with no special exemption for AI companies, and that data being publicly available online is not the same as someone consenting to have it used this way.
Legally, that's a meaningful statement. Technically, it runs into a wall.
Companies training these models have been candid, in filings and public statements, that once your data is baked into a trained model's weights, there is no clean way to surgically remove just your contribution. The data isn't stored. It's absorbed. Retraining an entire model to exclude one person's information is enormously expensive, and nobody is doing it on a per-request basis. What most AI companies offer instead is a promise to exclude flagged data from future training runs. That does nothing about the model that already exists, that has already learned your patterns, and that is very possibly already deployed and in use.
So the "right to erasure" exists on paper. In practice, once you've been crawled and trained on, there is no equivalent of the data broker's delete button. There's no row to remove. There's no confirmation email. There's no 45-day compliance window, because there's no known mechanism to comply with in the first place.
Once You're In, There's No Registry and No Address
When a data broker holds your information, there's an entity you can identify, a jurisdiction it operates in, and increasingly, a law that tells it what it has to do when you ask. That's uncomfortable, but it's at least a known relationship.
Once your data has been absorbed into a trained model, that relationship dissolves. Model weights get copied, fine-tuned, distilled into smaller models, redistributed as open-weight files, and licensed to companies who never touched your data directly and have no way of knowing it's in there. There's no registry of everywhere a given model, or a model derived from it, has ended up. There's no single company you can send a request to, because by the time a model is three or four generations of fine-tuning removed from the original training run, "who has my data" stops being an answerable question.
This is the part that's easy to underestimate: it's not just that deletion is hard. It's that the entire concept of a defined boundary, a company, a database, a jurisdiction, stops applying. Your information doesn't sit in one place anymore. It becomes a permanent, diffuse input into systems whose downstream copies nobody is tracking.
Not Every System Guarding That Data Has the Same Guardrails
Here's where it gets uneven in a way that matters. Some AI platforms are heavily filtered, audited, and restricted in what they'll surface or generate. Others are explicitly built and marketed around having few or no content restrictions at all. That's a selling point for some users, but a genuine liability for everyone whose data ended up in the training set regardless of consent.
That means the practical risk of having your information crawled isn't uniform. It depends entirely on which systems absorbed it, how those systems are governed, who has access to query them, and what those people or organizations are permitted, or willing, to do with what comes out. A well-governed model with strict use policies and a poorly governed one trained on overlapping data can produce very different outcomes from the exact same underlying information about you.
Multiply that by how many systems have already crawled the open internet. Training runs happen continuously, from companies you've heard of and dozens you haven't. The honest answer to "who has access to information about me through AI systems" is: nobody actually knows, including, in some cases, the companies that trained the models.
SIM Swap Protection
Get our SAFE plan for guaranteed SIM swap protection.
The Uncomfortable Bottom Line
Data broker deletion laws exist because brokers are identifiable, their data is structured, and regulators built enforcement mechanisms around both of those facts. None of those three conditions reliably hold for AI training data. The law is starting to reach for AI. Europe's guidance and a handful of narrow, new state-level rules are real, recent developments. But "the law technically applies" and "there is a working way to comply with it" are two very different sentences right now, and the gap between them is where your information currently sits.
This isn't a reason to assume nothing can be done. It's a reason to stop assuming that "I can request deletion" is still the whole strategy. The old playbook, find the company, ask them to delete your record, confirm it's gone, was built for a world of databases. We're not entirely in that world anymore, and pretending otherwise is the actual risk.
The more useful question isn't "how do I get this deleted." It's "what do I do about information that, realistically, isn't coming back out?" And that's a very different problem to solve.
Monthly
Yearly
Frequently Asked Questions
Can I request that an AI company delete my data from a trained model?
You can ask, but companies have said in filings and public statements that there's no clean way to surgically remove one person's contribution once it's baked into a model's weights. Most offer only to exclude flagged data from future training runs, which does nothing about the model already deployed.
Does GDPR's right to erasure apply to AI training data?
Legally, yes. The EDPB's mid-2026 guidance confirmed GDPR applies in full to personal data scraped for AI training, with no special exemption for AI companies. Technically, there's often no mechanism to comply with a deletion request once the data has been absorbed into a model.
Does California's Delete Act cover AI companies the way it covers data brokers?
The Delete Act's deletion platform reaches every registered data broker, which is checked every 45 days, with fines for non-compliance. That structure depends on a company storing data in a queryable, deletable database, the kind of structure AI training doesn't have.
If I can't get my data deleted from an AI model, what can I actually do?
The more useful question isn't "how do I get this deleted," it's what to do about information that, realistically, isn't coming back out, meaning reducing further exposure and hardening the accounts and numbers tied to your identity, rather than treating deletion as the whole strategy.

%20in%20Gmail.avif)


