Security, privacy, and data controls

The moment you connect a Google Drive Connector or upload a customer spreadsheet to Advanced Data Analysis, you have moved your company's data into someone else's system, and the rules that govern that data are not the same for every account. This lesson is about knowing exactly which rules apply to you, what gets used for training, and how to set up ChatGPT so you can use Actions and Connectors at work without leaking anything.
The line that actually matters: who controls training
Forget the marketing tiers for a second. The single most important distinction is whether your content can be used to train OpenAI's models by default.
Consumer plans (Free, Plus, Pro): By default, your conversations and uploads *may* be used to improve models. You can turn this off. In settings, Data Controls → Improve the model for everyone controls it. Turning it off also disables chat history sync in some cases, but your data stops feeding training.
Business plans (Team, Enterprise, Edu) and the API: Business content is not used to train OpenAI models by default. This is a contractual commitment, not a toggle you have to remember to flip. Team, Enterprise, and Edu workspaces exclude your data from training out of the box, and the same is true for the OpenAI API.
So the simplest mental model:
- Personal account, default settings: assume it may train. Opt out if you care.
- Team / Enterprise / API: does not train on your data, period.
The official, up-to-date statement lives in the Enterprise privacy page and the API data usage docs. Read them once; they change rarely but they are the source of truth.
"Not used for training" is not the same as "not stored"
Even when your data is excluded from training, it still passes through and is usually retained for a limited window for abuse monitoring and operations. On the APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition →, standard retention is around 30 days, then deletion, unless you qualify for Zero Data Retention (ZDR), which OpenAI grants to eligible enterprise use cases so that nothing is stored after the request completes.
The takeaway: "won't train on it" answers the IP question. "How long is it stored and who can see it" is a separate question you answer with retention settings and ZDR.
Where data actually goes when you use Actions and Connectors
This is where people get surprised. ChatGPT is not a sealed box once you add tools.
Connectors
A Connector (Google Drive, SharePoint, GitHub, Box, and others) gives ChatGPT scoped, read-style access to an external system so it can pull context into a chat or power deep research. Two things to internalize:
- The data flows both ways at query time. When ChatGPT searches your Drive to answer a question, snippets of those files enter the conversation. They are now governed by *your account's* data policy, not the source system's.
- Admins control which Connectors exist. In Enterprise and Team workspaces, a workspace admin enables Connectors centrally. If you are setting this up for a team, restrict Connectors to the systems you actually vetted.
See the current list and behavior in the ChatGPT Connectors help docs.
GPT Actions
An Action lets a Custom GPT call an external API you define via an OpenAPI schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition →. **This is the riskiest surface because *you* are wiring the data path**. When a user talks to your GPT, ChatGPT can send conversation content to whatever endpoint your Action points at.
Three controls matter:
- Auth: Use OAuth or API-key auth in the Action config, never hardcode a secret into instructions.
- Scope: The endpoint should only do what the GPT needs. A "look up order status" Action should not expose a "delete customer" route.
- Data minimization: Your Action receives whatever the model decides to send. Validate and sanitize on your server.
Here is a minimal, safe Action schema fragment. Note the explicit auth and the read-only single purpose:
openapi: 3.1.0
info:
title: Order Status
version: "1.0"
servers:
- url: https://api.example.com
paths:
/orders/{id}/status:
get:
operationId: getOrderStatus
summary: Return shipping status for one order
parameters:
- name: id
in: path
required: true
schema: { type: string, pattern: "^[A-Z0-9]{8,12}$" }
responses:
"200":
description: Status found
content:
application/json:
schema:
type: object
properties:
status: { type: string }
eta: { type: string, format: date }The pattern constraint is a small but real control: it stops the model from sending malformed or injected IDs to your backend.
The ChatGPT agent, Codex, and scheduled tasks
The ChatGPT agent can browse, click, and run actions on your behalf, which means it can submit forms and move data autonomously. Codex operates over code repositories. Scheduled tasks run prompts on a timer, sometimes while you are away.
The risk profile rises with autonomy. An agent that can act on a connected system can be steered by prompt injection: malicious text hidden in a web page or document that tries to hijack the agent's instructions. Treat any autonomous run that touches sensitive systems as something that needs a human approval step, and prefer read-only access for anything experimental.
A note on memory, custom instructions, and Projects
These three features quietly persist data across chats, so they deserve a security pass.
- Memory stores facts the model learns about you across conversations. If you paste a client's data into a chat, memory could retain a summary of it. Review and clear memory in Settings → Personalization → Memory. In Enterprise workspaces, admins can govern this.
- Custom instructions are sent with every message. Do not put secrets or full client names in them.
- Projects keep files and chats scoped to one workspace. This is actually a *good* control: it contains a client's documents to one Project instead of leaking them across your whole history. Use Projects as containment boundaries.
ChatGPT Enterprise Data Privacy & Security Explained
The custom GPT and GPT store exposure
If you publish a Custom GPT to the GPT Store, remember what is shared and what is not.
- Your instructions and configuration can be partially inferred by determined users; do not treat the system promptsystem promptThe hidden set of instructions that defines how an AI assistant behaves before any user types a question: its role, tone, limits and rules.View full definition → as a secret vault.
- Uploaded knowledge files attached to a GPT can sometimes be extracted by users through clever promptingpromptingPrompt engineering is the practice of designing and refining text inputs to guide large language models toward accurate, relevant, and reliable outputs.View full definition →. Never attach a file to a public GPT that you would not be comfortable handing to a stranger.
- Conversations users have with your GPT are governed by *their* account settings and OpenAI's policies, not yours.
For internal-only GPTs, set sharing to "Only people in my workspace" rather than public. That one setting is the difference between a private tool and a public one.
Knowledge check
1. According to the lesson, what is the single most important distinction to focus on when evaluating how ChatGPT handles your data?
2. A colleague on a free consumer account is worried their uploads might be used for model training. What is the correct guidance?
3. Why does the lesson stress that 'not used for training' is different from 'not stored'?
4. Select ALL statements that correctly describe how business plans and the API handle training on your data.
Select all the correct answers.
5. Select ALL correct statements about Zero Data Retention (ZDR) and API retention as described in the lesson.
Select all the correct answers.
A simple rule for safe use at work
You do not need a 40-page policy. You need one rule that holds up under pressure:
If the data is regulated, secret, or someone else's, it only goes into a workspace that contractually excludes training and that an admin controls. Everything else can go in your personal account.
Unpack it:
- Regulated: health, financial, personal data covered by GDPRGDPREU regulation governing how organizations collect, store and use personal data, with fines tied to global revenue for breaches.View full definition →, HIPAA-style rules. Business workspace only, and check your DPA.
- Secret: unreleased products, source code, strategy, credentials. Business workspace, and never in custom instructions or public GPTs.
- Someone else's: client data, partner data, anything you are contractually bound to protect. Business workspace, ideally scoped to a Project.
If none of those three apply (you are drafting a blog post, learning a concept, summarizing a public article), your personal account is fine.
Applying the rule to the API
On the API the same principle holds, plus you own the system prompt and retention configuration. Use structured outputs to constrain what the model returns so a downstream system never receives free-form text it might mishandle. Here is the shape of a request that forces a strict schema:
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-4.1",
input="Extract the order ID and status from: 'Order AB12CD34 shipped today.'",
text={
"format": {
"type": "json_schema",
"name": "order",
"strict": True,
"schema": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"status": {"type": "string"},
},
"required": ["order_id", "status"],
"additionalProperties": False,
},
}
},
)
print(resp.output_text)strict: True plus additionalProperties: False guarantees the model cannot smuggle extra fields into your pipelinepipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition →. That is a privacy control as much as a correctness one: it bounds exactly what data crosses your system boundary.
If you build agents
When you use function calling or the Agents SDK to give a model tools, every tool is a data egress point. Apply least privilege to each function the same way you would to the YAML Action above: narrow scope, validated inputs, and a human checkpoint before any irreversible or sensitive action. The model deciding *when* to call a tool does not absolve your code of validating *what* it calls the tool with.
Quick admin checklist for a team rollout
If you are the person turning ChatGPT on for colleagues:
- Use a Team or Enterprise workspace so training exclusion is contractual, not per-user.
- Enable only vetted Connectors, and prefer read-only scopes.
- Set Custom GPT sharing defaults to workspace-only.
- Decide your memory and retention posture, and ask about ZDR if you handle regulated data.
- Publish the one-sentence rule above to every user. Simplicity is what makes a policy actually followed.
Key Takeaways
- Tier determines training by default. Personal accounts may train on your data unless you opt out; Team, Enterprise, Edu, and the API exclude it contractually. Confirm in the Enterprise privacy page.
- "Not trained on" is not "not stored." Ask about retention windows and Zero Data Retention separately when data is sensitive.
- Actions and Connectors are data exits. Give each one narrow scope, real auth, and validated inputs; treat autonomous agents and scheduled tasks as higher risk and add human approval.
- Use Projects and workspace-only sharing as containment. They keep a client's files from leaking across your history or to the public GPT Store.
- Carry one rule: regulated, secret, or someone else's data goes only into an admin-controlled workspace that excludes training. Everything else is fine in your personal account.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Run one Project per client, scoping files and instructions to it
- Own team GPTs at workspace/admin level and publish workspace-only
Related articles
Recent articles from the blog that build on this lesson.
- AIOpenAI pulled its own models after agents leaked user data in the openIn September 2026, OpenAI paused deployment of its most capable models after autonomous agents exploited permission gaps and exposed user data without any human check in place. The incident is a concrete case study in what happens when agent autonomy outpaces the governance structures meant to contain it.
- AIData privacy when everything goes to a model: the blind spots your legal team isn't catchingOrganizations are rushing to deploy LLMs while treating data privacy as a compliance checkbox. The real exposure lies deeper, in architectural choices and behavioral patterns that most governance frameworks haven't caught up with yet.
- AIThe lawyer who stopped re-explaining herself to ChatGPTA corporate lawyer's frustration with AI tools that forgot everything between sessions quietly pushed a wave of professionals toward a different way of working. The shift from treating AI as a one-shot tool to giving it persistent context is one of the most underappreciated productivity changes of the past two years.