How UNICEF Used AI to Analyze 160 Evaluations and Improve Humanitarian Response
- Categories AI, Case Studies
- Date April 19, 2026
How UNICEF Used AI to Analyze 160 Evaluations and Improve Humanitarian Response
This case study is based on the official UNICEF review document:
The following write-up integrates insights from this official source for the M&E and international development community.
Why This Case Matters for M&E
Most M&E systems produce reports—but fail to drive organizational change. This case from UNICEF's Office of Emergency Programmes (EMOPS) shows how a large organization systematically analyzed 160 evaluation management responses to answer two critical questions: Is UNICEF actually learning from its evaluations? And can Generative AI make this analysis faster and more effective?
The Approach: Hybrid AI-Human Evaluation
UNICEF tested a hybrid evaluation model combining three Generative AI methods with traditional qualitative coding using Atlas.ti. The dataset included 160 management responses from humanitarian evaluations conducted between 2017 and 2024, split into two periods (2017-2020 and 2021-2024) to track changes over time.
Machine learning platform for rapid theme identification (3 days, $500-$2,500)
Chatbot interface using ChatGPT/Claude for queryable document repository
UNICEF's internal AI chatbot (free to staff, requires careful prompting)
Manual thematic analysis with Atlas.ti (~54 hours, $3,000-$5,000)
Key Findings: What UNICEF Learned
1. Organizations Do Learn—But Slowly
Of nearly 2,500 actions reviewed, approximately 68% were marked as completed, and nearly a quarter were underway. Management responses showed that UNICEF teams largely agreed with evaluation recommendations. However, some recommendations continued to repeat across years, signaling systemic issues rather than isolated gaps.
2. Strong Progress in Three Areas
Accountability to Affected Populations (AAP): Before AAP guidance was released, recommendations called for a common approach. After release, recommendations shifted to implementation and building systems to support AAP. This shows clear policy-to-practice progression.
Coordination Systems: The cluster coordination model operates more clearly in recent years. Recent challenges focus on proper resourcing and staffing, not basic functionality.
Humanitarian-Development-Peace Nexus (HDPN): Better knowledge of HDPN as an approach and better integration of guidance into existing UNICEF tools is now evident.
3. Weak Progress in Core M&E Areas
Themes that showed limited or unclear progress included: partnerships with implementing partners (localization remains a policy commitment but operational gaps persist), organizational capacity and culture (ongoing need to standardize risk appetite), support and operations functions (challenges remain in multi-year funding and rapid personnel placement), gender and equity (conversations are continuing but no clear progress), and planning, monitoring, evaluation, learning and KM (ongoing focus but no clear progress).
4. AI is Fast—But Not Enough Alone
Generative AI was extremely helpful in identifying key themes in the data for further inquiry and analysis. AILyze identified key themes within three business days at relatively low cost. However, AI struggled with identifying trends over time, anomalies in the data, and grasping nuance. Only human researchers could answer the research questions fully.
5. Data Quality Remains a Challenge
Nearly 10% of the dataset was not substantive information (waivers, non-humanitarian evaluations). AI failed to flag these problematic data that were immediately obvious to a human researcher. This highlights the continued need for human oversight in evaluation data analysis.
Comparing AI and Traditional Methods
| Method | Estimated Time | Estimated Cost | Best Use Case |
|---|---|---|---|
| Traditional Coding | ~54 hours | $3,000-$5,000 | Nuanced understanding, trend analysis |
| AILyze (ML) | ~24 hours | $500-$2,500 | Rapid theme identification |
| Exchange Design | ~80 hours setup | $3,000-$5,000 + monthly | Repeated queries on fixed corpus |
| UNIBot (Internal) | ~20 hours prep | Free to UNICEF staff | Single document or small dataset queries |
Lessons for M&E Professionals
Use AI as a "research assistant"—all results must be verified, reviewed, and validated by a human researcher with appropriate skills.
Some management responses showed 100% completion rates that may not reflect reality. Incentives for accurate reporting matter.
When the same recommendations appear year after year, it indicates systemic issues, not isolated gaps.
Each AI platform required data in a slightly different format. Time for data cleaning and preparation should not be underestimated.
Challenges and Caveats
- Data Asymmetry: Some evaluations had many recommendations while others had few, skewing the analysis.
- Self-Reporting Bias: Status of completion may not have been updated upon completion, and there may be incentives to report actions as completed or underway even with only incremental progress.
- AI Cannot See Time: Generative AI platforms could not identify trends over time without manual data restructuring.
- Non-Replicable Results: AI-generated results change as the corpus of data changes, raising concerns for research replicability.
- Privacy Concerns: UNICEF only used publicly available data with platforms that do not train models on UNICEF data. Personally identifiable information was never provided to AI systems.
Recommendations for Organizations
Use AI for initial data exploration; improve tracking of action completion; ensure data quality before analysis.
Standardize evaluation data structures; strengthen feedback loops between evaluations and management responses; train staff on prompt engineering.
Build hybrid AI-human evaluation systems; focus on learning—not just reporting; create queryable repositories of evaluation documents.
Key Takeaway for M&E Professionals
The future of M&E is not AI versus humans—it is AI plus human judgment working together. Generative AI can rapidly identify themes across hundreds of documents, but only human evaluators can grasp nuance, identify trends over time, and make meaning of findings in context. Organizations that build hybrid systems will learn faster and improve more effectively than those relying on either method alone.
References and Resources
Primary Source: UNICEF Office of Emergency Programmes (2024). Review of Management Responses to Humanitarian Evaluations, 2017–2024; Generative Artificial Intelligence (AI) and Traditional Qualitative Coding Compared.
Citation: Bruce, K. (2024). Review of Management Responses to Humanitarian Evaluations, 2017-2024; Generative Artificial Intelligence (AI) and traditional qualitative coding compared. UNICEF Office of Emergency Programmes, New York.
Advance your skills in AI-enhanced evaluation
Learn how to integrate AI-assisted qualitative methods into your own M&E practice.
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
