Sensitive Data in ChatGPT: Risk Management, GDPR Rules and Safe Alternatives
How to adequately protect personal data and trade secrets when using generative AI in the workplace.
Entering sensitive data or personal information into the public version of ChatGPT is a significant security risk and often a violation of the General Data Protection Regulation (GDPR). Data you share, such as customer names, social security numbers, or contracts, is stored on external (often US-based) servers and can be used to train future AI models. The Dutch Data Protection Authority explicitly warns against these AI data leaks. To process trade secrets and personal data, exclusively use secure corporate environments like ChatGPT Enterprise or Microsoft 365 Copilot, and always apply strict anonymisation.
Using generative AI tools offers enormous productivity benefits, but handling corporate data involves significant risks. When employees copy and paste customer information, financial forecasts, or patient records into the public version of ChatGPT, this data leaves the secure corporate environment. This not only leads to potential corporate espionage and reputational damage but can also result in hefty fines for violating the General Data Protection Regulation (GDPR).
This article provides a detailed analysis of the risks surrounding sensitive data in ChatGPT, the legal frameworks of the GDPR in relation to artificial intelligence, and the technical and policy alternatives for working safely with Large Language Models (LLMs).
Why sensitive data in public ChatGPT poses a risk
When a user enters a prompt in the standard, free version of ChatGPT (or ChatGPT Plus), this data is sent to OpenAI's servers, which are primarily located in the United States. OpenAI's terms explicitly state that entered conversations may be used to train future iterations of their language models.
This mechanism creates two fundamental security issues:
- Retention on external servers: The data is stored outside the sphere of influence of your IT administrators. You generally cannot selectively delete this data once it has been entered.
- Training data exposure (Memorization): Once models are trained on your data, there is a risk that this data—including trade secrets or personal information—will be inadvertently reproduced when another user outside your organisation enters a specific, related prompt.
Once sensitive data has been incorporated into the model via training, it is technically extremely complex, if not impossible, to make the AI "forget" this specific knowledge (machine unlearning). This creates a permanent risk of exposure. For more information on the broader societal and business implications of these types of AI issues, you can consult the other concerns about AI.
What constitutes an AI data leak under the GDPR?
The Dutch Data Protection Authority (DPA) explicitly warned in 2024 about the dangers of entering personal data into chatbots. According to the GDPR, a data breach occurs when personal data is accidentally or unlawfully destroyed, lost, altered, disclosed, or accessed.
When an employee places personal data (such as a client's name, a medical condition, or a combination thereof) into a public AI tool, this legally qualifies as an unauthorised disclosure to a third party. In such cases, a valid Data Processing Agreement (DPA) is lacking.
A well-known case that garnered national attention involved the municipality of Eindhoven. Here, civil servants were found to have entered citizen data into ChatGPT to edit and summarise texts. The municipality had to formally report this as a data leak to the DPA and immediately block the functionality on the internal network. Organisations risk fines of up to 20 million euros or 4% of global annual turnover for such violations.
Data you should absolutely never share with AI (Table)
To provide employees with clear guidelines, it is crucial to exactly define what falls under 'sensitive data'. Any information security policy regarding AI must explicitly prohibit at least the following categories in public AI tools.
| Category | Examples of specific data points | Potential impact of a leak |
|---|---|---|
| Personal Data (PII) | Social security numbers (BSN), name and address details, dates of birth, passport numbers. | Identity fraud, severe GDPR fines, reputational damage. |
| Special Category Personal Data | Medical records, political preferences, race, biometric data. | Violation of fundamental rights, direct DPA intervention. |
| Financial Business Information | Unpublished quarterly figures, budget plans, bank accounts. | Insider trading, competitive disadvantage, financial fraud. |
| Intellectual Property (IP) | Unpatented inventions, raw source code with API keys. | Theft of trade secrets, loss of competitive advantage. |
| Customer and Supplier Data | Pending contracts, price agreements, NDAs. | Breach of contract, loss of trust, loss of accounts. |
For organisations, a clear demarcation is crucial. If you want to legally and technically secure your business frameworks, supported by in-depth market research, consider using AI consultancy for a safe AI policy or review the best practices in the AI reports by ai.nl.
Safe AI alternatives for business use
If you want to leverage the functionality of generative AI without compromising data, the market offers various enterprise alternatives centred around privacy by design. The choice of a specific platform depends on your compliance requirements (such as ISO 27001 or NEN 7510) and budget.
1. ChatGPT Enterprise and Team
OpenAI offers business subscription plans (Enterprise and Team) that fundamentally differ from the public versions. In these subscriptions, the contracts explicitly state that customer input will not be used to train OpenAI's models. Furthermore, OpenAI provides Data Processing Agreements (DPAs) for this, Single Sign-On (SSO) is available, and you can mandate that corporate data is retained within Enterprise environments via specific retention settings.
2. Microsoft 365 Copilot and Azure OpenAI
For many organisations already using the Microsoft ecosystem, Microsoft 365 Copilot is the most logical step. This AI works with the same GPT models but is isolated within the organisation's tenant (secure corporate environment). Data processed by Copilot generally remains within the same geographical region (EU Data Boundary) and automatically inherits the existing access permissions of SharePoint and Teams. No model training takes place using tenant data.
3. European and Open-Source Models
For organisations with the highest security requirements—such as central government, defence, or banks—dependency on American tech giants is sometimes undesirable due to the US Cloud Act. European alternatives like Mistral (Le Chat), based in France, adhere more strictly to the EU's vision on data sovereignty. Additionally, an organisation can opt for Llama (Meta) or other open-source models and host them on-premises or within a self-managed European cloud environment. In this setup, you send absolutely no data to external parties.
Anonymising and pseudonymising data (Prompting techniques)
Even when using enterprise AI environments, best practice dictates applying the core principle of data minimisation. To summarise or process texts, an AI model often does not need to know specific names or unique identifiers. We distinguish two techniques to strip prompts of sensitivity:
- Anonymisation: Previously entered data can definitively no longer be traced back to a person. For example, changing "Patient Johan de Vries, BSN 12345678, has diabetes" to "An anonymous patient has diabetes".
- Pseudonymisation: Data is replaced by a synonym or code, with the key kept locally. For example, changing names to "[Customer A]" and "[Customer B]". The AI restructures the text, and after the output is generated, you replace "[Customer A]" back to the real name locally (outside the AI tool).
Training employees to recognise and redact this type of data is essential for a safe AI rollout. For targeted workshops and courses, platforms like ai.nl offer responsible ChatGPT training to teach teams these skills.
Adjusting privacy settings in standard ChatGPT
Although company-wide adoption of public ChatGPT is discouraged, millions of professionals still use it locally. If an employee—in the role of a freelancer, student, or for private use—nevertheless uses the regular (free or Plus) application for less critical data, the risks of training exposure can be manually minimised.
You can prevent conversations from being used for AI training in the ChatGPT settings:
- Click on your profile name in the bottom left.
- Navigate to Settings.
- Go to Data controls.
- Toggle off the Chat history & training option.
Note: By disabling this option, the chat history will no longer be saved to your account, making it difficult to navigate through old prompts. Furthermore, this is a personal setting and cannot be centrally enforced by an IT department in the free version.
Technical oversight and Data Loss Prevention (DLP)
In addition to informing and training staff, technical blocks are often necessary to definitively nip data leaks in the bud. Network administrators and CISOs are increasingly deploying Data Loss Prevention (DLP) tools and Cloud Access Security Brokers (CASB) to prevent data exfiltration to AI tools.
A DLP tool monitors network traffic and clipboard activity in real time. When an employee highlights a Social Security Number pattern (a specific nine-digit number) or a credit card number pattern, copies it, and attempts to paste it into the domain chat.openai.com or claude.ai, the system intervenes.
DLP systems can:
- Warn (Soft block): A pop-up appears asking: "Are you sure you want to send social security numbers to an external website? This is contrary to policy."
- Block (Hard block): The text output (paste action) is technically blocked and an automatic alert is immediately sent to the IT helpdesk.
- Domain block (DNS-level block): Access to unvetted AI websites is temporarily suspended or redirected to the approved internal corporate chatbot.
The impact of AI hallucinations on business continuity
When you feed AI with raw, internal data without adequate guardrails, there is an integrity risk alongside the risk of data leakage. Generative AI is trained to produce plausible, well-written pieces of text, not necessarily to guarantee factual truth. This phenomenon is known as a hallucination.
If an employee enters (pseudonymised) financial contracts with the prompt: "Calculate the total depreciation for the past four quarters", the LLM may generate correct-looking, but mathematically flawed amounts. Many standard text models are not optimised for calculations without the Advanced Data Analysis plugins.
When sensitive decisions—such as legal advice or credit assessments—are directly copied by people from an AI output without a "human-in-the-loop" review, this leads to operational errors. Data may or may not have been leaked; but if the output is incorrect, it still causes damage to business decisions.
Drafting a robust, internal AI policy (Checklist)
Prevention is better than cure. Every modern organisation must now have an explicit AI protocol or an AI Acceptable Use Policy (AUP), regardless of company size. This policy should cover not only AI chatbots but also AI plugins, browser extensions, and API integrations.
Include the following fundamentals in your policy:
- Clear whitelisting: Explicitly list which AI tools (and specific versions) are approved for organisational use.
- Classified data streams: The "Traffic Light Model". Green for public data; Orange for internal data that may carefully feed into corporate environments (Enterprise/Copilot); Red for highly personal or critical financial data that may exclusively feed self-hosted on-premise AI.
- Approach towards output: Guidelines stating that employees remain ultimately responsible for everything an AI generates and publishes under the organisation's name.
- Legal basis: Ensuring that the use complies with the GDPR, potentially supported by a Data Protection Impact Assessment (DPIA) for the structural deployment of AI on sensitive data.
By actively choosing enterprise accounts, proactively training employees to recognise sensitive data, and deploying technical solutions such as DLP, organisations maintain control. Artificial Intelligence can significantly accelerate your business-critical processes, as long as this acceleration does not compromise compliance and information security.
Veelgestelde vragen
Can I put personal data and customer data into ChatGPT?+
No, you may not enter personal data into the public or free versions of ChatGPT. The tools generally do not comply with the GDPR, as there is no processing agreement and data goes to US servers. For this purpose, exclusively use specific corporate environments (such as ChatGPT Enterprise or Copilot) where data is not used for training, backed by a data processing agreement.
What is an AI data leak according to GDPR rules?+
An AI data leak occurs when employees share sensitive information or personal data with a publicly accessible AI, or when this data is logged into an AI model's training data. Under the GDPR, this qualifies as an unauthorised data transfer and may result in a reportable incident to the Data Protection Authority (DPA).
Does ChatGPT use my conversations for training?+
Yes, unless you manually turn this off. In the free and Plus versions of ChatGPT, logged conversations are standardly shared with OpenAI's servers to optimise their algorithms and future versions. With paid Enterprise and Team contracts, it is standardly established that customer information will not substitute into training purposes.
Is Microsoft Copilot safer for corporate data than standard ChatGPT?+
For corporate data, Microsoft 365 Copilot is significantly safer as it integrates directly with your secure corporate environment (tenant). Data you share via Copilot does not leave your internal infrastructure for AI training by third parties. The authorisations and security settings of your existing Microsoft license are maintained.
What did the municipality of Eindhoven do wrong with ChatGPT?+
In the case of the municipality of Eindhoven, civil servants directly copied, pasted, and processed documents containing citizens' personal data into the public version of ChatGPT. The CISO classified this as a reportable data leak to the Data Protection Authority. The municipality subsequently blocked access to the AI tool on its internal network immediately.
How do I prevent employees from pasting customer data into AI without authorisation?+
Organisations can prevent data exposure through a combination of three pillars: 1) Purchasing a secure corporate license (like Enterprise), 2) Drafting and enforcing a clear AI code of conduct, and 3) Technically deploying Data Loss Prevention (DLP) tools that block the upload of categories like social security numbers (BSN) and IBANs to public chatbots.
Are there European or locally hosted AI alternatives to ChatGPT?+
Yes, European alternatives are gaining traction due to strict compliance needs (GDPR). The French AI company Mistral offers similar capabilities to ChatGPT via 'Le Chat' and heavily focuses on European data sovereignty. You can also opt for an open-source LLM, like Llama, and run it yourself on an internal European cloud server or local machine (on-premise) without an external data connection.
Blijf scherp op AI
Want to be sure your AI policy is correct?
Engage our AI specialists to make your data security and AI protocols legally and technically watertight.
Volgende stap
Bekijk ai.nl consultancy for secure AI usage