Privacy Has a Hoarding Problem
Key Takeaways
- AI Is Changing the Value of Old Data: Information once kept largely “just in case” can now support analytics, enterprise search, retrieval systems and other AI applications, strengthening the incentive to retain it.
- Retention Carries Its Own Risk: Every dataset an organization keeps may create additional cybersecurity, privacy, legal, operational and governance obligations.
- Data Minimization Is Risk Management: Properly disposing of information that is no longer needed can reduce the amount of sensitive data available to attackers, employees, third parties and AI systems.
- Retention Policies Need to Result in Actual Deletion: Written schedules mean little if information continues to survive in backups, cloud applications, downstream systems and other copies after its intended retention period.
- Future Usefulness Is Not a Retention Strategy: AI may make historical data genuinely valuable, but the possibility that information could someday prove useful should not become a default justification for keeping it indefinitely.
Deep Dive
For a long time, deleting data could feel strangely reckless. Storage was cheap, and information was potentially valuable. The cost of keeping another year of customer records, internal correspondence, transaction histories or old documents was difficult to see, while the cost of deleting the wrong thing was easy to imagine. So companies kept it. Some of it was retained for legal or operational reasons, some because nobody was quite sure whether it could safely be destroyed, and some because of the most durable retention policy in corporate life: we might need it someday.
Someday has arrived with artificial intelligence. A body of information that spent years gathering dust on a server can now look quite different to an organization building AI systems. Historical records can support analytics. Documents can populate retrieval systems. Communications can contain knowledge that was previously difficult to extract at scale. Old datasets may be useful for developing or refining models, depending on their contents, provenance and the legal basis for using them. Generative AI and enterprise search can make information scattered across an organization far easier to find and interrogate.
The possibility of future usefulness was always part of the case for keeping data. AI has made the possibility considerably less hypothetical. It has also made the decision to delete considerably harder.
That is where a familiar privacy problem starts looking like a much broader problem of enterprise risk. The same archive that might contain tomorrow's useful insight is also something that must be secured today. If it contains personal information, its continued retention may have to be justified. Copies may have spread across production systems, backups, employee devices, cloud applications and third parties. Access has to be controlled. Retention requirements have to be applied. Litigation or an investigation may make portions of it relevant. And when a new AI application arrives, somebody has to decide whether that application should be allowed anywhere near it.
Data may be an asset. Assets, unfortunately, require looking after.
The Case for Keeping Everything Just Got Better
There are perfectly legitimate reasons to retain information, and AI does not make large datasets inherently suspect. Historical information can have operational and analytical value. Organizations have legal obligations that require certain records to be preserved. Researchers and data scientists may need sufficiently large and representative datasets to do useful work. AI systems themselves can require substantial quantities of data, depending on what they are designed to do.
The problem begins when those specific reasons collapse into a general presumption that more is better. That presumption was already tempting when the chief advantage of retention was optionality. AI strengthens it because information that seemed practically unusable can become usable again. A company does not necessarily need employees to know which folder contains a decade-old document if an enterprise search system can locate it. Nor does someone necessarily need to read thousands of records individually if software can analyze them at scale.
The dormant archive is not quite so dormant anymore. For privacy and risk teams, however, usefulness is not the same thing as necessity.
The UK Information Commissioner's Office makes the point rather plainly in its guidance on AI and data protection. Data minimization requires organizations to determine what personal data is actually needed for the relevant purpose. The regulator warns that discovering later that information is useful for prediction does not retroactively establish why an organization needed to keep it in the first place. Personal data should not be collected simply on the chance that it might prove useful later, although foreseeable future needs may justify retention where an organization can properly support that decision.
That creates a tension AI governance committees are likely to encounter repeatedly. A technologist may be perfectly correct that a dataset could have future value. A privacy officer may be equally correct that hypothetical future value does not, by itself, justify retaining personal information indefinitely. The hard part is deciding when possibility becomes purpose.
The Liability Sitting Quietly on the Server
Data retention has an unusual quality as a risk decision because doing nothing can look like doing nothing. Keeping another archive does not necessarily produce an immediate incident. A forgotten database can sit untouched for years. An old customer file can survive several system migrations without attracting much attention. A retention schedule can exist on paper while copies of the underlying information continue living elsewhere.
Deletion is visible, but retention often isn't. The risk nevertheless accumulates. NIST's treatment of privacy makes clear why. Its Privacy Framework approaches privacy as an enterprise risk-management issue and treats the data lifecycle as extending from collection through disposal. NIST has also explicitly identified records retention as something that can create privacy risk and has noted that retained personal information can become vulnerable to unauthorized access or use.
The cybersecurity logic is even less complicated. Information that has been legitimately destroyed is no longer sitting in a corporate system waiting to be stolen from it. That does not mean deletion is a substitute for cybersecurity. Organizations still need access controls, encryption, monitoring, segmentation and the rest of the security architecture appropriate to their risks. Nor does deleting one copy necessarily mean the information has disappeared from backups, downstream systems or third parties.
But reducing unnecessary holdings changes what there is to defend. The ICO made precisely that connection in cybersecurity guidance published in May. Among its recommendations for protecting personal information from AI-powered cyber threats were data minimization and storage limitation, observing that organizations should retain what they genuinely need and that holding less leaves less available to steal. It also recommended auditing what personal data organizations hold, where it resides and who can access it, with particular attention to AI systems that process or train on personal information.
This is where data minimization begins to look less like an austere privacy principle and more like an ordinary control on exposure. A company would not normally describe accumulating unnecessary privileged accounts as a strategy for preserving future optionality. It would recognize that every unnecessary account creates another avenue that must be governed and protected. Yet data can escape the same intuition because its potential value is easier to imagine than the incremental risk created by one more retained dataset.
The risk does not disappear because it is difficult to put on a balance sheet.
AI Has a Way of Finding What We Forgot
There is another reason the old approach to retention deserves reconsideration. For much of the information companies accumulated, obscurity provided a kind of accidental friction. The information existed, but using it required someone to know it existed, know where it was stored, obtain access and then make sense of it. A document buried seven folders deep on an old shared drive might technically have been available to hundreds of employees while remaining practically invisible to almost all of them.
Generative AI and modern enterprise search can weaken that friction. That can be tremendously useful. It is part of the attraction. An employee can find institutional knowledge without knowing who created it or where it was filed. Organizations can make previously fragmented information accessible to people who need it.
The governance question is whether everything that can be found should be findable.
An old file may contain personal information that is no longer required. A document may have been created under access assumptions that made sense for the system in which it originally lived. Historical information may be inaccurate, obsolete or stripped of the context that once made its meaning obvious. Permissions that appear adequate when humans retrieve records one at a time may deserve another look when an AI system can search across repositories and synthesize what it finds.
The ICO has begun considering the same issue in its work on agentic AI. Its guidance says organizations should not give AI agents access to personal information merely because the information might prove useful in the future. Instead, access should follow a defined purpose and a justifiable need, an approach the regulator compares to the familiar cybersecurity principle of least privilege.
There is an important idea buried in that comparison. Data minimization and least privilege are cousins. Least privilege asks why a person or system needs access. Data minimization asks why the organization needs the information in the first place. One limits the doors. The other asks whether everything behind those doors still needs to be in the building.
A Retention Schedule Is Not a Deletion Program
Most mature organizations do not need to be told that retention schedules exist. The more revealing question is what happens when the retention period expires. The ICO's storage-limitation guidance says organizations should establish standard retention periods where possible, have systems for ensuring those periods are followed in practice and review retention at appropriate intervals. It also advises organizations to reconsider retention when records are not actually being used.
“Followed in practice” is doing considerable work there.
Information rarely lives neatly in one place for its entire life. It is exported. Copied. Attached. Backed up. Migrated. Synced to cloud services. Shared with processors. Pulled into spreadsheets. Incorporated into other datasets. A policy can say five years while the technical environment quietly says forever.
That makes retention an awkwardly cross-functional problem. Privacy teams can establish requirements but usually do not control every system in which information resides. Records managers can define schedules without owning the infrastructure that executes deletion. CIOs inherit systems containing years of accumulated information. CISOs must defend whatever remains. Legal teams may impose holds that properly suspend ordinary destruction. Compliance functions need evidence that policies are operating rather than merely existing. AI governance committees may suddenly discover that a proposed application depends on datasets whose origins and retention history are poorly understood.
Nobody needs to have made an obviously bad decision for the organization to end up with too much data. Thousands of individually understandable decisions will do. AI merely makes the consequences easier to see.
A serious retention program therefore has to answer questions that are less glamorous than selecting an AI model and considerably more important than they sound. What information do we actually possess? Why do we still have it? Which system is authoritative? Where are the copies? Who can reach them? Which third parties have them? What prevents deletion? When a retention period expires, does anything actually happen?
And if the answer to the last question is no, the organization does not really have a retention schedule. It has a document describing one.
Deletion Has an Economic Value
Corporate discussions about data tend to give preservation the benefit of the doubt. Keeping information preserves possibilities. Deleting it extinguishes them. That accounting is incomplete because it assigns a value to optionality while treating continued possession as free. It isn't.
Every retained dataset creates some combination of security, privacy, legal, operational and governance obligations. Not every dataset creates the same ones, and the existence of those obligations does not mean deletion is appropriate. But indefinite retention is still a decision, even when nobody remembers making it.
The arrival of AI makes a disciplined approach more important, not less. Organizations have better tools for extracting value from information, which means the opportunity cost of deletion may genuinely be higher than it once was. That deserves to be recognized. So does the other side of the calculation.
Data minimization offers something that most security controls cannot: the possibility of removing part of the exposure altogether. An organization cannot suffer a breach of information it has properly disposed of and no longer holds. An AI application cannot unexpectedly retrieve a deleted archive. Employees cannot misuse information that is no longer available to them. The organization does not have to keep classifying, migrating and controlling data it has legitimately destroyed.
There are caveats to each of those propositions in a complicated technical environment. Copies matter. Backups matter. Legal obligations matter. De-identification and anonymization can alter the analysis. Retention requirements differ by information, purpose and jurisdiction. None of this supports indiscriminate deletion.
It supports treating deletion as an affirmative risk decision rather than the regrettable loss of a potentially useful asset. The better question is not whether organizations should keep more data or less. It is whether they can distinguish information they have chosen to retain from information they have merely failed to delete.
AI will make some old data unexpectedly valuable. It will probably make organizations grateful they preserved certain records. It may also tempt them to turn every uncertain future use into an argument for indefinite retention. That is a poor substitute for governance. The modern enterprise has spent years becoming very good at accumulating information and considerably less enthusiastic about throwing any of it away. Generative AI did not create that habit. It exposed how consequential the habit has become.
More data can mean more optionality. It can also mean more attack surface, more access decisions, more systems to govern and another body of information whose continued existence somebody will eventually have to defend. If an organization cannot explain why it still has a particular body of data, keeping it is difficult to call an asset strategy. It looks much more like unmanaged risk.
The GRC Report is your premier destination for the latest in governance, risk, and compliance news. As your reliable source for comprehensive coverage, we ensure you stay informed and ready to navigate the dynamic landscape of GRC. Beyond being a news source, the GRC Report represents a thriving community of professionals who, like you, are dedicated to GRC excellence. Explore our insightful articles and breaking news, and actively participate in the conversation to enhance your GRC journey.

