What actually happens to your business data when it enters an AI model
Sending a contract, a customer list, or internal financials into an AI tool feels like using a search engine. It is not, and the distinction carries real legal and competitive consequences.
Neo NeumannAI Practice LeadJuly 31, 2026Listen to the podcast
4 min
The concept at stake here istraining data leakage: the risk that information your organisation inputs into an AI model becomes part of that model's future behaviour, or is retained in ways you did not authorise and cannot easily reverse. This is not a theoretical threat from a distant future. It is a structural feature of how many commercial AI systems are designed, and most organisations using these tools in 2026 have not fully thought through the exposure.
The confusion stems from a simple mental model that most people carry over from search engines or SaaS tools: you put something in, you get something out, and then it's over. With many AI services, that is not how data flows.
Why it matters for business professionals specifically
If you are a consultant summarising client financials in ChatGPT, a lawyer drafting a contract with Copilot, or an HR lead generating performance review language with an AI writing tool, you are making a data governancedata governanceData governance is the set of policies, roles, and processes that ensure data is accurate, secure, well-defined, and used responsibly across an organization.View full definition → decision whether you realise it or not. The person doing the task usually does not see themselves as making that decision. That gap is where most incidents happen.
The legal exposure is concrete. Under GDPR, personal data cannot be transferred to a processor without a Data Processing Agreement (DPA) in place and, in many cases, without explicit legal basis. Under the EU AI Act, which applies from mid-2026 onward, providers of high-risk AI systems face additional obligations around data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition → and documentation. If an employee pastes employee performance data or customer PII into a free-tier AI tool that has no DPA with your organisation, that is a reportable breach in many EU jurisdictions, regardless of whether anything bad actually happened with the data.
Beyond legal risk, there is competitive exposure. Strategy documents, product roadmaps, M&A targets, pricing models: these represent the actual value of a firm. Sending them into a cloud-based model operated by a third party, without contractual protections, is the kind of decision that boards ask about after an incident, not before.
How it actually works
When you use a consumer or free-tier AI product, your inputs are typically logged, reviewed by the provider for safety and quality purposes, and in many cases used to fine-tune or improve future model versions. OpenAI's terms of service for its consumer product, ChatGPT Free and Plus, stated until relatively recently that conversations could be used to train models unless users opted out. The APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → and enterprise tiers work differently: OpenAI's enterprise terms, and similar terms from Anthropic and Google, generally exclude customer data from training by default, though the specifics vary and change over time.
The mechanics at a practical level look like this. You paste a 40-page supplier contract into a chat interface. That text is sent as a prompt to the provider's servers. The model processes it, returns a response, and the exchange is logged. Depending on the product tier and jurisdiction, that log may be retained for 30 days, 90 days, or longer. A human reviewer at the provider may read flagged conversations. The data may cross borders, landing in data centres in the US even if you are operating under EU law.
Fine-tuningFine-tuningFine-tuning adapts a pre-trained model to a specific task or domain by continuing training on a smaller, targeted dataset, improving accuracy and style for that use case.View full definition → adds another layer. If your organisation uses a custom fine-tuned model, data used during that fine-tuning process is literally baked into the model's weights. There is no "delete" function for a weight. If a model is fine-tuned on customer data and then that customer relationship ends, you cannot surgically remove what the model learned. This is a real problem that several enterprise teams discovered in 2024 and 2025 when they attempted to comply with GDPR's right to erasure.
A concrete example: in 2023, Samsung engineers pasted proprietary source code into ChatGPT to help debug it. The code became part of the conversation logs on OpenAI's servers. Samsung subsequently banned internal use of generative AI tools for several months. The incident did not require a malicious actor. The exposure happened through ordinary, well-intentioned use.
When using AI with sensitive data makes sense, and when it does not
There are configurations where the risk is genuinely low. Using an on-premise or private-cloud deployment of an open model (Llama 3, Mistral, Falcon) means data never leaves your infrastructure. Microsoft's Azure OpenAI Service, under the enterprise agreement, offers a contractual commitment that data is not used for training and stays within your specified Azure region. These are legitimate options, and for organisations with the technical capacity to implement them, they shift the risk profile materially.
Vendor claims deserve scrutiny here. Microsoft, Google, and Anthropic all make data protection commitments in their enterprise tiers, but those commitments are written by the vendors themselves and their strength depends entirely on contractual enforcement and your legal team's ability to negotiate and audit. Independent legal review of any AI provider's DPA is not optional if the use case involves personal data.
The situations where you should not be putting data into a cloud AI tool without protections in place:
- Any data covered by attorney-client privilege
- Customer PII, including names combined with any behavioural or financial information
- Non-public financial information, particularly around M&A or earnings
- HR records, performance data, health-adjacent information
- Competitive intelligence or IP you would not share with a vendor under a standard NDA
The honest tradeoff is productivity versus control. Employees using AI tools with unrestricted access to company data will work faster. That speed has real value. But it does not eliminate liability, and it does not disappear if something goes wrong. The framing that makes sense for most organisations is not "should we allow AI use" but "which data categories require which controls, and who owns that decision."
Getting that classification done, and communicating it clearly to employees, is currently the most underinvested capability in enterprise AI governance. Most organisations have an AI policy as a document. Fewer have translated it into concrete guidance at the task level, where the actual exposure occurs.
Finished reading?
Validate your read to earn XP and feed your radar.