
The Use of AI at UNFPA
EvalCommunity Academy Case Study
How AI Supported the Evaluation of UNFPA’s Humanitarian Action
A practical case study for evaluators, M&E specialists, humanitarian teams, evaluation offices, and development practitioners exploring responsible AI, multilingual evidence processing, and human-led evaluation judgment.
Last updated: May 2026 · 9 min read · EvalCommunity Academy case study
Introduction
UNFPA’s use of AI in the independent evaluation of its humanitarian action provides a practical example of how artificial intelligence can support large-scale evaluation work without replacing professional evaluation judgment. The evaluation of UNFPA’s capacity in humanitarian action from 2019 to 2025 covered 15 countries, six humanitarian contexts, more than 1,500 documents and publications, and over 600 interviews and focus group discussions.
The scale of the evidence base created a methodological challenge: processing the material manually at the required pace and depth would have limited both the breadth and rigour of the evaluation. The evaluation team therefore explored whether AI could support experienced evaluators responsibly, particularly in evidence processing, multilingual analysis, synthesis, triangulation, and cross-country comparison.
This case study is relevant for evaluators because it shows a carefully governed, ethics-first approach to using AI in a sensitive humanitarian evaluation. It demonstrates how AI can increase efficiency and improve evidence coverage while keeping human validation, evaluator expertise, transparency, and responsible data protection at the centre of the process.
Case Background
Humanitarian evaluations often require teams to synthesize large volumes of documentary, interview, and focus group evidence across multiple countries and crisis contexts. In the UNFPA case, the independent evaluation of humanitarian capacity from 2019 to 2025 included extensive multilingual evidence, with a significant proportion of source material in French and Spanish.
The evaluation team saw an opportunity to test whether AI could improve the breadth, pace, and rigour of evidence processing. However, the team did not begin with technology. It first developed a deliberate ethics-first strategy, then selected tools that could meet evaluation needs while respecting security, privacy, multilingual inclusion, and human oversight requirements.
The case is also part of a broader institutional commitment. UNFPA’s Independent Evaluation Office had already developed a strategy for a GenAI-powered evaluation function, including principles and a phased roadmap for responsible use. The humanitarian evaluation became one of the first centralized evaluations to put those principles into practice at scale.
The Evaluation Problem
The core problem was scale. The evaluation team needed to process more than 1,500 documents and publications and over 600 interviews and focus group discussions across 15 countries. A purely manual process would have required weeks of sequential review and could have constrained the depth and consistency of analysis.
The second challenge was language inclusion. Four of the sampled countries were francophone and three were Spanish-speaking. If multilingual evidence were more difficult to process, there was a risk that evidence from these contexts could be underused or deprioritized in the synthesis.
The AI Strategy
The evaluation team tested two tools: InsightWise and Google NotebookLM. NotebookLM was selected for most of the analytical work because it was easy to use, integrated with Google Translate, and supported processing of French and Spanish materials alongside English-language documents.
The selection also reflected UNFPA IEO’s institutional preference for tools within the Google ecosystem, which met information security and data privacy requirements. This was especially important because the evaluation handled sensitive humanitarian evidence, including anonymized interview and focus group material.
Scale
15 countries, six humanitarian contexts, more than 1,500 documents, and over 600 interviews and focus groups.
Tool choice
NotebookLM was used for most analytical work after testing against evaluation needs and security requirements.
Human oversight
Evaluators remained responsible for validation, interpretation, triangulation, and final findings.
How AI Was Used
1. Data analysis
Before any interview or focus group data were ingested into the platform, the team removed all personal identifiers. AI then helped generate concise summaries and identify recurring themes, patterns, and emergent evidence across a very large transcript dataset. This did not replace manual coding; it complemented the evaluation team’s coding against a pre-agreed evaluation matrix.
2. Secondary data review
AI supported structured scanning and extraction across the 1,500-document evidence base. It helped draw out relevant passages and organize them by country, evaluation assumption, and timeframe.
3. Synthesis
AI assisted with cross-country comparisons and flagged points of convergence and divergence between data sources. These signals helped evaluators identify areas of agreement, contradiction, or tension that required deeper professional scrutiny.
Governance and Safeguards
The evaluation team embedded human oversight at every stage. AI-generated outputs were not accepted automatically. Evaluators reviewed outputs, checked them against primary and secondary sources, and corrected or discarded results that were inaccurate, superficial, or reflected bias.
Two workshops strengthened external quality control. An analysis workshop with evaluation managers took place immediately after data collection, and a co-creation workshop enabled the full Evaluation Reference Group to critically review AI-supported findings.
The process was governed by UNFPA IEO’s GenAI-powered evaluation strategy and the UNEG Ethical Principles for Harnessing AI in United Nations Evaluations. Data protection was also supported by UNFPA’s enterprise agreement with Google, which helped ensure that sensitive evaluation data, even in anonymized form, was not retained or used beyond the immediate analytical task.
Transparency was built into the evaluation report through an annex documenting how AI was applied, which tools were used, and how ethical and responsible use was upheld.
Results and Limitations
Tangible efficiency gains
AI reduced the time needed to process and structure a dataset that would otherwise have required weeks of sequential manual review. The process delivered a verified cost saving of $12,000 and gave evaluators more time to focus on interpretation, validation, and professional judgement.
Stronger multilingual inclusion
The multilingual processing capability helped ensure that evidence from francophone and Spanish-speaking contexts contributed to the synthesis on equal terms. This strengthened evidence coverage and reduced the risk that language constraints would shape what counted in the final analysis.
Clear limits of AI-generated analysis
The experience showed that AI tools can be highly efficient at extracting and summarizing information, but the resulting analysis can remain superficial. Complex interpretation, contextual judgement, and meaningful insight still depend on human evaluators with institutional and methodological expertise.
Evaluation Framework for Responsible AI Use in Humanitarian Evaluations
EvalCommunity Academy users can adapt the following framework when considering AI support for complex evaluations.
Purpose and necessity questions
- What evaluation problem will AI help solve?
- Is AI needed, or would conventional qualitative and quantitative analysis be sufficient?
- Does AI improve the quality, breadth, or timeliness of the evaluation?
- Will AI free evaluators to focus on interpretation and judgement?
- How will the team know whether AI added value?
Data protection and ethics questions
- Have all personal identifiers been removed before data are processed?
- Does the tool meet institutional information security and privacy requirements?
- Are sensitive data retained, reused, or exposed beyond the evaluation task?
- Are AI use, tools, limitations, and safeguards documented transparently?
- Is the approach aligned with relevant ethical principles for evaluation and AI?
Quality and validation questions
- How will all AI-generated summaries and classifications be checked?
- Are findings traced back to primary and secondary sources?
- How are bias, superficial analysis, and inaccuracies identified?
- Does AI complement an evaluation matrix or replace it?
- Who has final authority over interpretation and findings?
Inclusion and language questions
- Does the tool support all languages represented in the evidence base?
- Are non-English materials included equitably in the analysis?
- Does translation preserve meaning, context, and nuance?
- Are country-specific findings protected from being flattened into generic themes?
- Does AI strengthen or weaken representation of affected communities?
Practical Lessons for M&E and Development Professionals
First, use AI only where it improves evaluation quality and efficiency. UNFPA’s example shows a demand-driven approach: AI was used because the evidence base was large, multilingual, and complex, not because AI adoption was an end in itself.
Second, protect sensitive data before using AI tools. Removing personal identifiers before ingestion is essential, especially in humanitarian evaluations involving staff, partners, and community members.
Third, treat multilingual processing as an inclusion issue. AI can help prevent evidence from non-English contexts from being deprioritized, but evaluators must still verify whether translation and synthesis preserve meaning.
Fourth, document AI use transparently. An annex describing tools, stages of use, and safeguards strengthens credibility and creates a learning resource for other evaluation teams.
Finally, keep humans in the lead. AI can accelerate extraction, summarization, and pattern identification, but evaluators remain responsible for interpretation, triangulation, ethical judgement, and final findings.
FAQ
What evaluation did UNFPA use AI to support?
AI supported the independent evaluation of UNFPA’s capacity in humanitarian action from 2019 to 2025, covering 15 countries and six humanitarian contexts.
Which AI tool was used?
The team tested InsightWise and Google NotebookLM, then selected NotebookLM for most analytical work because it was easy to use, supported multilingual processing with Google Translate, and aligned with institutional security requirements.
How was AI used in the evaluation?
AI was used for transcript summarization and theme identification, structured scanning of the document evidence base, and synthesis tasks such as cross-country comparisons and flagging convergence or divergence between data sources.
How did UNFPA manage ethical risks?
The team removed personal identifiers before data ingestion, used tools aligned with institutional security requirements, validated outputs manually, followed UNFPA and UNEG ethical guidance, and documented AI use transparently in the evaluation report.
What were the results?
AI reduced the time needed to process and structure the evidence base, delivered a verified cost saving of $12,000, strengthened multilingual evidence inclusion, and allowed evaluators to focus more on validation, interpretation, and judgement.
Can AI replace evaluators in humanitarian evaluations?
No. The case shows that AI can support extraction, summarization, and comparison, but complex interpretation, contextual understanding, ethical judgement, and final findings depend on human evaluators.
Conclusion
UNFPA’s humanitarian evaluation shows how AI can support complex evaluation work when it is introduced with clear purpose, ethical safeguards, human oversight, multilingual inclusion, and transparent documentation. The case is not a story of replacing evaluators. It is a story of helping evaluators handle scale while preserving the professional judgement required for credible findings.
For EvalCommunity Academy users, the main lesson is practical: responsible AI in evaluation requires more than choosing a tool. It requires a governance strategy, secure data handling, validation routines, quality control workshops, and a clear commitment to keeping human evaluators in the lead.
