AI Data Privacy: Your Prompts and Business Knowledge Are Data Too 

Businesses usually think about AI privacy in terms of customer records, contracts, financial information and confidential documents. Those things matter, but AI introduces another category of valuable information that is much easier to overlook: the knowledge contained in the conversation itself. 

Employees are now using AI to work through strategy, solve technical problems, design products, improve pricing and challenge existing business processes. In some cases, the most commercially valuable information given to an AI system may never have existed in a company database. It is created during the conversation. 

In our companion article on AI Data Sovereignty, we look at where information travels depending on how the model is hosted. This article focuses on the next question: once an AI service receives information, what can happen to it? 

Processing is not the same as training 

When an AI model answers a question, it performs inference. It receives your prompt and any supporting information, processes it and generates an answer. That does not automatically mean the information becomes training data. 

This distinction is important because major AI providers currently give business customers much stronger protections than ordinary consumer users. OpenAI says ChatGPT Business, Enterprise and API inputs and outputs are not used for model training by default. Anthropic makes the same commitment for its commercial products and API. 

So it would be misleading to say that every document retrieved through RAG, or every customer record provided to a business AI service, automatically becomes part of a future model. 

However, how you access the AI matters. A personal account can operate under different rules from an enterprise account, and settings around model improvement, feedback and retention can also change the position. OpenAI, for example, currently enables model sharing by default on personal workspaces while business products are opted out by default. 

For a business, saying “we use ChatGPT” or “we use Claude” therefore tells you very little. You need to know which product, under which commercial terms, with which settings. 

Even feedback can change what happens to a conversation 

The humble thumbs-up or thumbs-down button is a good example of why AI governance can be more complicated than it first appears. 

Anthropic says commercial inputs and outputs are excluded from training by default, but if a user explicitly submits feedback, the related conversation can be retained and potentially used for training. OpenAI similarly provides business and API customers with optional mechanisms to share feedback and other data for model improvement; these are disabled by default for API organisations. 

That means governance is not simply about selecting an approved AI platform. Businesses also need to understand what users can do inside it. 

What if the valuable information is the idea? 

This is where the privacy discussion becomes much more important for businesses. 

Imagine an executive spends several hours working with an AI model to develop a completely new pricing model. During that conversation, they explain why existing approaches fail, provide insights gained over 20 years in the industry, test several ideas and eventually arrive at a commercial approach that competitors are not using. 

Removing the executive’s name does not remove the value from that conversation. 

The same could apply to: 

  • a new manufacturing method; 
  • a unique sales process; 
  • a better way of allocating labour; 
  • a new software architecture; 
  • a novel financial model; or 
  • a previously unknown solution to an industry problem. 

AI models do not simply store conversations as searchable documents, and supplying an idea does not mean another user will later receive a copy of it. Training is about learning patterns and relationships from information. Anthropic itself describes models as learning general patterns from training data rather than storing it like a database. 

That is precisely why de-identification and intellectual-property protection are not the same thing. 

If a conversation is eligible for training, removing someone’s name may protect their identity while preserving much of the useful reasoning, methodology or problem-solving information contained in the interaction. 

Commercial agreements matter, but they still require trust 

Enterprise terms provide meaningful protection. They give businesses contractual rights and make it much clearer what the provider is and is not permitted to do with company data. 

But a contract is not the same type of control as preventing the provider from receiving the information at all. 

There are good reasons for businesses to keep some caution here. The history of AI training data has already shown that the industry’s judgement about what information it is entitled to use does not always match the expectations of the people who created that information. 

Anthropic provides a concrete example. A US court found that while training on lawfully acquired books could qualify as fair use, Anthropic’s creation and retention of a library containing millions of books obtained from piracy sites was not protected in the same way. Anthropic subsequently reached a US$1.5 billion settlement covering hundreds of thousands of works. 

OpenAI is also facing major ongoing litigation over its training sources. Recent filings in litigation brought by The New York Times and other publishers allege that OpenAI knowingly copied copyrighted and paywalled material for model training. OpenAI disputes the infringement claims and argues that its use is protected by fair use, so this should not be presented as an equivalent final finding against OpenAI. 

Importantly, these cases are about how foundation-model training data was sourced, not evidence that either company secretly broke its current enterprise agreements and trained on business customers’ private API data. 

But they do demonstrate why businesses should be cautious about making trust their primary control. 

Even the US Federal Trade Commission has specifically warned AI providers that promises not to use confidential customer information for purposes such as model training must be honoured, including attempts to achieve the same result through workarounds or later changes to terms. 

The practical lesson is not that commercial agreements are worthless. They are valuable. It is that a contractual promise and a technical restriction are different levels of protection. 

Reduce how much trust is necessary 

For ordinary business activity, a properly configured enterprise AI service may provide completely reasonable protection. 

But if the conversation contains genuinely important intellectual property, the better question may be whether an external provider needs access to it at all. 

This is where the privacy discussion connects directly to AI sovereignty. If a workload can run effectively using an open-weight model hosted in Azure, AWS, another controlled cloud environment or on-premise infrastructure, the organisation can reduce its dependency on another company’s data-handling promises altogether. 

The decision moves from: 

“Do we trust them not to use this?” 

towards: 

“Do they actually need to receive it?” 

That is a much stronger form of control. 

Treat prompts as company information 

The old advice for employees was simply: don’t paste confidential information into AI. 

That is no longer enough. Modern AI can retrieve information automatically, participate in hours of business planning and become part of the process through which new company knowledge is created. 

Businesses should therefore have clear expectations around which AI products staff can use, when personal accounts are unacceptable, whether training and feedback options are enabled, and what kinds of intellectual property should remain within controlled infrastructure. 

A useful rule is to treat prompts and AI conversations as another form of company information. 

Before giving information to an AI system, ask: what knowledge is the AI seeing, under what terms, and does it genuinely need to leave our environment? 

That question becomes increasingly important as AI moves from simply helping employees write faster to actively participating in how businesses think, design and innovate. 

Know what your AI is seeing

Most businesses can’t say which AI tools their team uses, under which terms, or with which settings. We can help you find out, and work out which information should never leave your environment.

Book a 30-minute discovery call. No sales pitch, just a clear view of where you stand and what to fix first.