UN AI evaluation evidence mapping
EvalCommunity Academy Case Study
UN AI Evaluation Evidence Mapping
A practical case study for evaluators, M&E specialists, knowledge managers, UN teams, and development practitioners exploring how AI and large language models can improve access to evaluation evidence across the UN system.
AI-Assisted Evidence Workflow
From scattered UN evaluations to usable evidence maps
A visual interpretation of the OECD-described workflow: scoping, framework development, eligibility criteria, search strategy, LLM-assisted extraction, classification, interactive mapping, and evidence summaries.
Scoping
Concept note and consultation with evaluation and synthesis experts.
Framework
Taxonomy aligned with UN priorities and report typologies.
Eligibility
Selection of strategic evaluations linked to SDGs.
Search
Screening selected evaluations from the UNEG repository.
Last updated: May 2026 · 9 min read · EvalCommunity Academy case study
Introduction
UN AI evaluation evidence mapping is a practical example of how artificial intelligence can support evidence use across a large multilateral system. The OECD Development Co-operation TIPs case Enhancing evidence use with AI: The UN system-wide approach describes how the UN Sustainable Development Group System-Wide Evaluation Office conducted an AI-assisted initiative in 2024 to map and summarize evaluations across the United Nations system.
The initiative responded to a persistent challenge: the UN system produces a large amount of evaluation evidence, but that evidence is often scattered, difficult to access, hidden in lengthy reports, and underused by decision makers and Member States. The goal was to make evidence more accessible and useful for strategic decision-making, UN programme effectiveness, and progress toward the 2030 Agenda and the Sustainable Development Goals.
This case study is highly relevant for monitoring and evaluation because it shows how AI and large language models can support evidence extraction, classification, synthesis, knowledge management, and evaluation coverage analysis while still requiring human oversight, clear protocols, and improved data governance.
Case Background
Across the UN system, evaluations contain evidence, examples, lessons, and recommendations that can inform policy and practice. However, the case notes that evaluations are often not well known or widely used, either by the UN system itself or by Member States assessing the UN contribution to development results.
To address this problem, the UNSDG System-Wide Evaluation Office launched an AI-assisted evaluation mapping initiative. The work served two purposes: it aimed to make evaluation evidence more accessible, and it acted as a proof of concept for system-wide mapping of UN evaluations.
The initiative also tested the use of AI for evidence extraction, classification, and synthesis, making it a strong example of digital transformation in evaluation systems and learning-oriented knowledge management.
The Evidence-Use Challenge
The main challenge was not a lack of evaluation evidence. The UN system had thousands of reports. The problem was that evidence was dispersed across repositories, embedded in long documents, inconsistently tagged, and difficult for decision makers to use at the right moment.
This is a common problem for large institutions: knowledge exists, but it does not automatically become usable evidence. Without accessible synthesis and clear classification, evaluation findings can remain hidden and underutilized.
The AI-Assisted Approach
The UN evaluation mapping initiative was implemented between April and October 2024. It followed a structured process that combined consultation, taxonomy development, eligibility criteria, search strategy, LLM-assisted coding, interactive mapping, and evidence summary production.
Scoping and framework
The team consulted senior evaluation officers and evidence synthesis experts, then developed a taxonomy based on UN development system priorities and report typologies.
Selection and search
Eligibility criteria focused on evaluations of strategic importance that clearly contributed to the SDGs, including country-level, regional, thematic, strategic, joint, pooled funding, and synthesis evaluations.
LLM classification
Large language model coding and data extraction were piloted on a sample of reports, then scaled to classify all selected reports.
The process resulted in a geographical map on ArcGIS and several interactive maps using EPPI Mapper, a free open-source software. The team also shortlisted topics, selected five priorities, sampled relevant evaluations, and drafted five 10–15 page evidence summaries.
Outputs and Results
1. Approximately 1,000 evaluations mapped and summarized
The initiative mapped and summarized approximately 1,000 UN evaluations issued from 2021 to 2024. This provided a clearer picture of the state of evaluative evidence across the UN system and demonstrated the usefulness of AI and LLMs for evidence classification and extraction.
2. Evidence became more visible and usable
The mapping made evaluation evidence easier to access and helped show that useful insights are often hidden inside lengthy reports. This points to the need for stronger system-wide knowledge management solutions.
3. Thematic priorities and evidence gaps were identified
The mapping showed that much evaluative evidence focused on national priorities, gender equality, climate, education, and food security. It also identified gaps in areas such as results-based management, disability inclusion, SDG financing, funding quality, and UN system coordination.
4. LLM performance varied by topic complexity
The case notes that LLM performance was not uniform across all topics. Some complex issues were more difficult to classify, reinforcing the need for human review, expert prompts, and careful interpretation.
Lessons Learned
Mature LLMs
Commercially available LLMs supported extraction, classification, and abstract generation without custom model training.
Human in the loop
Expert review was essential for prompt iteration, quality control, and classification refinement.
Focused teams
Three core team members produced the mapping in 8–10 weeks by combining evaluation, synthesis, and AI expertise.
Better data inputs
Standardized formats, tagging, and metadata are essential for effective automation and analysis.
Relevance for M&E and International Development
This case is directly relevant to M&E because it addresses one of the most important challenges in evaluation systems: how to turn large volumes of evaluation reports into usable, timely, and decision-relevant evidence.
For international development, the case matters because the mapping was designed to support UN decision makers and Member States in improving the effectiveness and efficiency of UN programmes and supporting progress toward the 2030 Agenda and the SDGs.
For humanitarian and development evaluators, the case also offers practical guidance on AI governance, evidence synthesis, knowledge management, and the need to combine technical tools with expert human judgement.
Evaluation Framework for AI-Assisted Evidence Mapping
EvalCommunity Academy users can adapt the following framework when assessing AI-supported evaluation evidence mapping initiatives.
Purpose and use
- What decision moments will the evidence map support?
- Who are the intended users?
- How will evidence use be tracked?
Data and classification
- Are reports consistently formatted and tagged?
- Is the taxonomy aligned with priorities?
- How is LLM performance assessed?
Governance and quality
- Who validates AI-generated outputs?
- What review protocols are used?
- How are limitations documented?
FAQ
What is UN AI evaluation evidence mapping?
It refers to the use of AI and large language models to classify, extract, map, and summarize evaluation evidence across the UN system so that it can be more easily accessed and used for decision-making.
Who led the initiative?
The initiative was conducted by the UNSDG System-Wide Evaluation Office, drawing from the repository maintained by the United Nations Evaluation Group.
How many evaluations were mapped?
The initiative mapped and summarized approximately 1,000 UN evaluations from 2021 to 2024.
What did the mapping reveal?
It revealed important thematic priorities, such as national priorities, gender equality, climate, education, and food security, as well as gaps in areas such as results-based management, disability inclusion, SDG financing, funding quality, and UN coordination.
What is the main lesson for evaluators?
AI can help make evaluation evidence more visible and usable, but effective results require human oversight, clear taxonomies, strong metadata, prompt iteration, interdisciplinary teams, and data governance.
Conclusion
The UN AI evaluation evidence mapping case shows how AI can help large institutions turn scattered evaluation reports into more accessible and decision-relevant knowledge. The initiative demonstrated the value of LLMs for evidence extraction, classification, and synthesis, while also showing the limits of automation when topics are complex or data inputs are inconsistent.
For EvalCommunity Academy users, the practical lesson is clear: AI can strengthen evidence use in M&E when it is embedded in a strong knowledge management system, supported by human expertise, and aligned with real decision-making needs.
