Avoid Data Leakage
Catch Me If You Can Series
Protect the data before you ask AI to process the text.
EvalCommunity Tutorial
Avoid Data Leakage When Using AI
A practical guide for protecting personal data, unpublished research, confidential reports, proposal drafts, and proprietary methods when using generative AI tools.
Quick Answer
AI data leakage happens when personal, confidential, unpublished, protected, or proprietary information is exposed through an AI tool without proper safeguards.
Do not upload personal data, unpublished research, confidential reports, proposal drafts, or proprietary methods into external AI tools unless you have permission, appropriate safeguards, and organisational approval.
What You Will Learn
- What AI data leakage means in research, evaluation, and proposal writing.
- Which information should not be uploaded into external AI tools without safeguards.
- How to check whether data is safe to use with AI.
- How to reduce risk through redaction, anonymisation, synthetic examples, or approved tools.
- How to manage AI meeting recorders and transcription tools responsibly.
- How to document AI use and data protection safeguards.
Main Sources and Related Tutorials
This tutorial is based on the ERA Living Guidelines on the responsible use of generative AI in research and related EvalCommunity guidance. The key issues are privacy, confidentiality, intellectual property protection, transparency, and human responsibility.
AI Data Protection Workflow
Use this workflow before uploading, pasting, summarising, analysing, or transcribing information with AI tools.
Step 1
Identify Data
Know what you plan to share.
Step 2
Check Risk
Public, personal, or confidential?
Step 3
Confirm Permission
Check consent and policy.
Step 4
Reduce Exposure
Redact, anonymise, or summarise.
Step 5
Use Safe Tools
Prefer approved systems.
Step 6
Document Safeguards
Record what was shared.
1. Why Data Leakage Matters
Generative AI tools can support writing, translation, summaries, checklists, outlines, and analysis. But every time you paste information into an external AI tool, you may be sharing data with a system outside your direct control.
The key question is not only whether AI can help. The key question is whether the information is safe and authorised to share.
Core Rule
If the data is not safe to share publicly, do not upload it into an external AI tool without safeguards.
2. What Is AI Data Leakage?
AI data leakage is the unintended or unauthorised exposure of personal, confidential, unpublished, protected, or proprietary information through an AI tool or workflow.
It can happen through public AI chatbots, document analysers, transcription tools, plugins, AI research platforms, or third-party systems with unclear privacy settings.
Practical Test
Would you be allowed to share this information with an external third party? If not, do not paste it into an external AI tool without safeguards.
3. What Not to Upload Without Safeguards
Do not upload the following types of information into external AI tools unless safeguards are in place.
Personal Data
- Names and contact details
- Photos or voice recordings
- Interview transcripts with identifiers
- Survey responses linked to individuals
- Case notes or identifying details
“`
Confidential Documents
- Internal reports
- Donor communications
- Financial information
- Procurement or legal documents
- Internal risk registers
Research and Evaluation Data
- Raw transcripts
- Raw survey datasets
- Unpublished findings
- Field notes
- Stakeholder feedback
Proposal and Donor Data
- Proposal drafts
- Budget logic
- Consortium strategies
- Partner roles
- Competitive positioning
Proprietary Material
- Internal tools
- Proprietary frameworks
- Unpublished methodologies
- Training manuals
- Consultancy templates
High-Risk Information
- Safeguarding case data
- Peer review material
- HR records
- Legal or contractual data
- Information about vulnerable groups
“`
4. Why Data Leakage Is Dangerous
| Risk | Why It Matters |
|---|---|
| Privacy breach | Personal data may be exposed without consent, approval, or a valid basis. |
| Confidentiality breach | Internal reports, donor files, or partner documents may be disclosed improperly. |
| Intellectual property loss | Proprietary methods, tools, or unpublished ideas may be exposed or reused. |
| Participant harm | Sensitive information may expose individuals, communities, or vulnerable groups. |
| Contractual risk | Client, donor, or partner agreements may prohibit external sharing. |
| Trust damage | Organisations may lose credibility if they mishandle sensitive information. |
5. The Data Safety Test
- Is this information public?
- Does it include personal data?
- Could someone be identified directly or indirectly?
- Is the document confidential or unpublished?
- Is it part of a proposal, donor process, or competitive submission?
- Does it include proprietary methods or internal tools?
- Do I have permission to upload it?
- Does the AI tool store, reuse, or train on inputs?
- Is the tool approved by my organisation?
- Can I use an anonymised, redacted, synthetic, or summarised version instead?
6. Safer Alternatives to Uploading Sensitive Data
Use Public Information
Use AI with public documents, public guidance, public templates, or approved public text.
“`
Use Redacted Text
Remove names, locations, identifiers, financial details, and sensitive quotations.
Use Synthetic Examples
Use fictional examples that reflect the task without exposing real people or organisations.
Use Summaries
Create a human-prepared summary that excludes sensitive details, then ask AI to organise or edit it.
Use Approved Tools
Prefer enterprise, internal, locally governed, or organisation-approved AI environments.
Use Internal Workflows
For highly sensitive material, keep analysis inside approved internal systems.
“`
7. Redaction Checklist Before Using AI
Before pasting text into AI, remove or replace:
- Names, email addresses, phone numbers, and IDs
- Locations, dates, and job titles that identify people
- Organisation names where sensitive
- Quotes that identify participants
- Case details, safeguarding information, and complaints
- Financial, contractual, donor, and negotiation details
- Partner strategies, internal IDs, metadata, and sensitive file names
8. Data Leakage in Common Workflows
| Workflow | Sensitive Content | Safer Approach |
|---|---|---|
| Evaluation | Interview transcripts, complaints, stakeholder feedback, draft findings. | Use anonymised summaries or approved protected tools. |
| Proposal writing | Proposal drafts, budget logic, partner strategy, donor feedback. | Use redacted excerpts or public criteria for checklists. |
| Research | Unpublished findings, participant data, peer review material, raw datasets. | Use protected environments or keep analysis internal. |
| Meetings | Confidential discussions, strategy, audio, transcripts, action points. | Inform participants, obtain consent, and review summaries manually. |
9. AI Meeting Recorders and Transcription Tools
AI data leakage can happen in meetings when AI recorders, transcription tools, or notetakers capture confidential discussions without clear consent or safeguards.
- Inform participants before using AI meeting tools.
- Ask for consent where required.
- Explain what will be recorded or transcribed.
- Explain where the data will be stored.
- Explain who will access the transcript or summary.
- Review AI-generated summaries before sharing.
- Delete recordings when no longer needed.
- Follow organisational rules.
Suggested Meeting Notice
This meeting may use an AI transcription or note-taking tool to support documentation. The transcript or summary will be reviewed by a human before being shared. Please let us know if you do not consent to AI-assisted transcription or summarisation.
10. Tool Safety Checklist
- Who owns or manages the tool?
- Is it public, enterprise, internal, or locally hosted?
- Does it use prompts or uploads for training?
- Can data be deleted?
- Where is the data stored?
- Who can access the data?
- Are plugins or third-party integrations enabled?
- Is there an enterprise or protected mode?
- Does the organisation approve this tool?
- Are audit logs and privacy settings clear?
11. Decision Table: Can I Upload This to AI?
| Information Type | Public AI Tool? | Safer Approach |
|---|---|---|
| Public guidance document | Usually yes | Verify the summary against the source. |
| Non-sensitive paragraph | Usually yes | Use AI for clarity or grammar. |
| Interview transcript with names | No | Anonymise or use approved protected tools. |
| Raw dataset with identifiers | No | Remove identifiers or use internal systems. |
| Confidential donor report | No | Use human review or approved secure AI. |
| Proposal draft | Use caution | Use redacted excerpts or approved tools. |
| Peer review material | No | Avoid external AI tools. |
| Safeguarding case data | No | Do not upload into external AI tools. |
12. Safer Prompt Examples
Public Text
Improve the clarity of the following public paragraph. Do not add new facts, claims, or recommendations.
“`
Redacted Summary
Organise the following anonymised summary into clear bullet points. Do not infer identities, locations, or missing details.
Synthetic Example
Using this fictional example, create a checklist for reviewing data privacy risks in an evaluation workflow.
Sensitive Data Check
Review the text below and flag any personal, confidential, or sensitive information that should be removed before external sharing.
“`
13. Risky Prompts to Avoid
- Summarise these identifiable interview transcripts.
- Analyse this confidential donor report.
- Upload this proposal draft and improve the strategy.
- Extract sensitive stakeholder concerns from this raw transcript.
- Use these HR records to identify staff performance issues.
- Summarise this safeguarding case.
- Analyse this unpublished peer review document.
- Improve this proprietary methodology and make it more competitive.
14. Simple Data Use Notes
For Non-Sensitive Editing
AI was used to support language editing of non-sensitive text. No personal, confidential, unpublished, or proprietary data was uploaded.
For Anonymised Data
AI was used to organise anonymised summary information. Identifiers and sensitive details were removed before use.
For Protected Internal Tools
AI support was provided through an organisation-approved tool with data protection safeguards. Outputs were reviewed by the responsible team.
15. Frequently Asked Questions
What is AI data leakage?
AI data leakage is the unintended or unauthorised exposure of personal, confidential, unpublished, protected, or proprietary information through an AI tool or workflow.
Can I upload interview transcripts into AI?
Only if the transcripts are properly anonymised or the tool, consent, ethics approval, and organisational safeguards allow it. Identifiable transcripts should not be uploaded into public AI tools.
Can I upload proposal drafts into AI?
Use caution. Proposal drafts may include confidential strategy, partner roles, budget logic, and proprietary methods. Use redacted excerpts or approved protected tools.
Can I upload public documents?
Usually yes, but still verify the output. AI summaries of public documents can be incomplete or incorrect.
What is the safest alternative?
Use public, anonymised, redacted, synthetic, or summarised information. For sensitive material, use organisation-approved protected tools or internal workflows.
16. Final Takeaway
Sensitive Information Is Not Ordinary Text.
AI can help with editing, structure, summaries, and checklists. But personal data, unpublished research, confidential reports, proposal drafts, and proprietary methods carry responsibilities.
If the data is not safe to share publicly, do not upload it into an external AI tool without safeguards.
