AI Safely with Sensitive M&E Data
EvalCommunity Academy Tutorial
How to Use AI Safely with Sensitive M&E Data
A risk-based workflow for confidential evaluations, monitoring data, qualitative evidence, and security-sensitive projects.
Artificial intelligence can support evaluators and M&E professionals with qualitative coding, evidence synthesis, report drafting, indicator analysis, knowledge products, and learning summaries.
However, many projects involve information that should not be entered into a public or unapproved AI chatbot. This is especially important when working with security-sector actors, vulnerable populations, confidential evaluations, complaints, protection data, politically sensitive findings, or unpublished institutional information.
Learning objectives
By the end of this tutorial, you will be able to:
- Identify sensitive information in M&E documents and datasets.
- Distinguish anonymization from pseudonymization.
- Classify an AI-supported task by risk level.
- Apply data minimization before using AI.
- Create a limited AI-ready data packet.
- Use synthetic data where appropriate.
- Document and review an AI-assisted workflow.
- Recognize when AI should not be used.
1. The scenario
Imagine that you are evaluating a project involving security-sector institutions. You have:
- Fourteen key-informant interviews
- Meeting notes from project partners
- An internal monitoring report
- A draft evaluation report
- References to ranks, institutions, locations, incidents, and operational challenges
You would like AI assistance to identify themes, compare perspectives, organize findings, develop a report structure, and improve the final document.
However, anonymizing all the material may take longer than completing the analysis manually.
2. What makes M&E data sensitive?
Sensitive information is not limited to names, email addresses, or telephone numbers. M&E professionals must also consider contextual, institutional, operational, and political risks.
Personal data
Personal data is information relating to an identified or identifiable person. It can include:
- Names and contact details
- Identification or staff numbers
- Photographs and signatures
- Exact professional roles
- Precise dates and locations
- Combinations of details that indirectly identify someone
Example of indirect identification
“The only female district commander appointed in 2024.”
No name is included, but colleagues familiar with the institution may still identify the person.
Sensitive non-personal data
Information may create harm even when it does not identify an individual. Examples include:
- Locations of supplies or programme activities
- Internal institutional weaknesses
- Unpublished evaluation findings
- Allegations involving an organization
- Partner-performance assessments
- Community tensions or political relationships
- Operational or security information
3. Anonymization and pseudonymization are different
Pseudonymization
Pseudonymization replaces identifying information with a code or label. A separate identification key may still connect that code to the original person, institution, or location.
| Original information | Pseudonymized information |
|---|---|
| Colonel Maria D. | Respondent 04 |
| Northern Security Directorate | Institution B |
| Municipality of Rivona | District 3 |
Pseudonymization reduces risk, but the information may still be linkable to its original source.
Anonymization
Anonymization aims to prevent people from being identified through the available information and other information that could reasonably be combined with it.
4. Start with the task, not the dataset
A common mistake is to prepare the complete dataset for AI before defining exactly what support is required. This creates unnecessary work and unnecessary risk.
Ask: What is the smallest amount of information the AI needs to complete this task?
Poorly defined task
“Analyze all the interviews and write the evaluation report.”
Better-defined task
“Develop a neutral structure for presenting three already validated evaluation findings.”
Other focused AI tasks
- Compare anonymized recommendations and identify overlap.
- Suggest headings for a summarized evidence matrix.
- Rewrite a generalized paragraph in clearer evaluation language.
- Review whether conclusions are supported by listed evidence references.
- Generate questions for human evaluator review.
5. The Sensitive M&E Data Decision
Complete the following decision process before providing project information to an AI system.
Question 1
Is the tool and environment approved?
Check:
- Organizational AI policy
- Data-protection and security policies
- Confidentiality clauses
- Donor or government restrictions
- Evaluation ethics requirements
- Commitments made through informed consent
Approved: Continue to Question 2.
Unclear: Obtain guidance before proceeding.
Not approved: Do not upload project information.
Question 2
Does the task require real project data?
Many tasks can use:
- Fictional examples
- Blank templates
- Generalized scenarios
- Synthetic evidence matrices
- Aggregated summaries
- Approved public information
No real data required: Use synthetic or public information.
Some real data required: Continue to Question 3.
Substantial raw data required: Reconsider the task or environment.
Question 3
What could happen if the information were disclosed?
Consider possible harm to:
- Participants and communities
- Staff and implementing partners
- Institutions and programme access
- Evaluator independence and credibility
- Organizational reputation
- Operational security
- Future relationships and participation
6. Classify the AI-supported task
Green: Lower-risk task
- Creating a generic evaluation-report outline
- Improving a blank template
- Editing public, non-sensitive text
- Analyzing synthetic data
- Generating example indicators or questions
Approach: AI use may be reasonable, subject to organizational rules and human review.
Amber: Controlled task
- Analyzing an anonymized evidence matrix
- Organizing generalized interview themes
- Reviewing pseudonymized monitoring information
- Comparing non-identifiable stakeholder groups
- Summarizing aggregated results
Proceed only when:
- The tool is approved.
- The information is minimized.
- Identifiers have been reviewed.
- The output will receive qualified human review.
- The activity is documented.
Red: High-risk task
- Uploading identifiable transcripts
- Uploading protection or safeguarding case files
- Sharing allegations about recognizable individuals
- Providing precise security locations
- Uploading restricted documents
- Sharing passwords or access keys
- Asking AI to make consequential decisions about individuals
Approach: Do not use a general-purpose external AI tool. Consider manual analysis, synthetic data, or an approved secure environment.
7. Build an AI-ready data packet
An AI-ready data packet is not a cleaned copy of the complete dataset. It is a deliberately limited package containing only the information required for one approved task.
The packet may contain
- A precise task instruction
- A generalized project description
- An anonymized evidence table
- Selected extracts
- Evidence identifiers
- Output instructions
The packet should exclude
- Names and contact details
- Precise addresses and locations
- Unnecessary dates and titles
- Identifiable quotations
- Document metadata
- Access credentials
- The identification key
Example transformation
Original record
During an interview on 14 March 2026, Deputy Commander Arben K. from the West Regional Border Police stated that an incident near Kelmara checkpoint had not been documented because the liaison officer feared disciplinary action.
Minimized narrative
A regional security official reported that an operational incident had not been formally documented because a responsible staff member feared disciplinary consequences.
| Evidence ID | Stakeholder group | Generalized finding | Confidence |
|---|---|---|---|
| INT-07 | Regional security actor | Fear of disciplinary consequences may discourage incident reporting. | Moderate |
8. Apply data minimization
Data minimization means limiting the information processed to what is genuinely necessary for the intended purpose.
| Level | Action | Example |
|---|---|---|
| Records | Reduce the number of records. | Use five evidence entries instead of 30 transcripts. |
| Variables | Remove irrelevant fields. | Exclude contact details and unrelated demographics. |
| Detail | Reduce precision. | Replace an exact date with a project phase. |
| Output | Narrow the requested output. | Ask for themes rather than a final conclusion. |
9. Separate sensitive analysis from AI support
AI does not need to participate in every stage. A safer workflow preserves human control over the most sensitive evidence and decisions.
- Step 1: Collect information using approved systems.
- Step 2: Store original records securely.
- Step 3: Review and code sensitive information manually.
- Step 4: Create a minimized evidence matrix.
- Step 5: Validate the matrix against the sources.
- Step 6: Provide only the approved packet to AI.
- Step 7: Use AI to organize themes or test structures.
- Step 8: Verify every AI-generated interpretation.
- Step 9: Approve and take responsibility for final conclusions.
10. Use synthetic data for methodological tasks
Synthetic data is invented information designed to represent the structure of a situation without reproducing real cases.
- Design a qualitative coding framework
- Test an AI prompt
- Develop a reporting template
- Compare analytical approaches
- Build an evidence-matrix structure
- Train team members without exposing project information
Example
Create fictional extracts representing a positive institutional perspective, a critical community perspective, and a mixed implementation perspective. Ask AI to demonstrate a coding method using those examples.
Synthetic information must never be presented as real evidence or used to manufacture findings.
11. Practical prompts
Prompt 1: Develop a report structure without sharing evidence
You are supporting the organization of an evaluation report. Do not generate findings or assume evidence that has not been provided. Create a proposed structure for presenting: 1. Institutional capacity findings 2. Implementation constraints 3. Stakeholder trust 4. Sustainability risks For each section, suggest: - A neutral heading - Questions the evaluator should answer - Types of evidence required - Possible limitations to disclose Use professional evaluation language. Do not invent project facts.
Prompt 2: Analyze a minimized evidence table
Analyze only the information in the evidence table. Requirements: - Do not infer participant identities. - Do not infer locations, institutions, or events. - Do not add external facts. - Cite the relevant evidence ID. - Distinguish evidence from interpretation. - Identify contradictory evidence. - Flag findings supported by one source. - State “insufficient evidence” where appropriate. - Do not draft a final evaluation conclusion. Output: 1. Emerging themes 2. Supporting evidence IDs 3. Contradictory evidence 4. Evidence gaps 5. Questions for evaluator review
12. Apply the efficiency test
Before investing significant time in data preparation, compare the cost, benefit, and risk of each approach.
- How long will safe preparation take?
- How long would manual completion take?
- What additional value will AI provide?
- What new risks and review requirements will AI introduce?
| Option | Preparation | Analysis | Review | Total |
|---|---|---|---|---|
| Manual review | 0 hours | 3 hours | 1 hour | 4 hours |
| AI transcript review | 5 hours | 1 hour | 2 hours | 8 hours |
| AI evidence matrix | 1.5 hours | 45 minutes | 1 hour | 3 hours 15 minutes |
13. Document the AI-assisted workflow
A short AI-use record improves accountability, traceability, and institutional learning.
Project: [Project name or internal code]
Task: [Specific AI-supported task]
Tool and environment: [Approved environment]
Date: [Date]
Data provided: [Minimized information]
Data excluded: [Original records and identifiers]
Preparation measures: [Anonymization or aggregation]
Output: [Themes, structure, questions, or revision]
Human review: [Reviewer role]
Corrections: [Summary]
Final decision: [Accepted, revised, or rejected]
Retention or deletion: [Required action]
14. Know when not to use AI
Manual analysis may be preferable when:
- The dataset is small.
- The context is highly specific.
- The information is extremely sensitive.
- Anonymization removes essential meaning.
- The AI output requires extensive verification.
- The tool is not approved.
- Consent or contractual conditions do not permit processing.
- The consequences of disclosure are serious.
- The task depends heavily on cultural, political, or interpersonal nuance.
15. Practical exercise
Scenario
You have ten interviews with justice-sector actors containing:
- Names and professional roles
- Institutions and locations
- Descriptions of incidents
- Critical observations about management
- Direct quotations
Your objective is to identify the main barriers to institutional learning.
Complete these steps
- Write the precise AI-supported task.
- Classify the information by sensitivity and analytical value.
- Choose manual analysis, synthetic data, an evidence matrix, or a secure environment.
- Create an AI-ready packet with no more than ten evidence entries.
- Test whether a project colleague could identify a person, institution, location, or event.
- Document whether the expected benefit is proportionate to the remaining risk.
16. Responsible AI checklist
☐ The AI task is clearly defined.
☐ The tool and environment are approved.
☐ Relevant policies and restrictions have been reviewed.
☐ The information has been classified by sensitivity.
☐ Real project data is genuinely necessary.
☐ Only the minimum necessary information is included.
☐ Direct identifiers have been removed.
☐ Indirect and contextual identifiers have been reviewed.
☐ Sensitive non-personal information has been assessed.
☐ The identification key is stored separately.
☐ Evidence references have been retained.
☐ A qualified person will review the output.
☐ The activity will be documented.
☐ Retention and deletion requirements are understood.
☐ Manual completion has been considered.
☐ The expected benefit justifies the remaining risk.
17. Frequently asked questions
Is replacing names with participant numbers enough?
Usually not. Roles, locations, dates, quotations, incidents, and demographic characteristics may still reveal identities.
Can I upload data when model training is disabled?
Disabling training does not automatically resolve authorization, retention, access, confidentiality, contractual, legal, or ethical requirements.
Is an enterprise AI account automatically appropriate?
No. Enterprise controls may reduce some risks, but the environment must still be approved and correctly configured.
Is AI useful for a small qualitative dataset?
Yes, but it may provide more value for coding frameworks, evidence matrices, alternative interpretations, or writing support than for processing all original records.
What if anonymization removes important context?
Reduce the scope of the AI task and conduct the contextual analysis manually.
Is deciding not to use AI a negative outcome?
No. A documented decision not to use AI can demonstrate professional judgment and responsible data management.
Further reading
NIST: De-Identification of Personal Information
European Data Protection Board: Anonymisation and Pseudonymisation
