Applying Evaluation Criteria Thoughtfully – OECD
- Categories Guides
- Date April 17, 2026
Applying Evaluation Criteria Thoughtfully: An EvalCommunity Analysis
What is this guidance about?
The OECD Development Assistance Committee (DAC) first laid out five evaluation criteria in 1991. After nearly three decades of learning and a wide-ranging consultation process (2017-19) informed by the 2030 Agenda for Sustainable Development and the Paris Agreement, the criteria were revised and expanded to six, with new definitions endorsed in December 2019. This 2021 guidance — Applying Evaluation Criteria Thoughtfully — is the first comprehensive document to help evaluators operationalize these definitions.
The guidance addresses a critical gap identified during the consultation process: while the criteria are widely used and understood, their practical application is often mechanistic and not tailored to context. The document unpacks each criterion's definition, provides elements for analysis, highlights common challenges with practical solutions, and offers real-world examples from evaluations across sectors and regions.
What the guidance does well
✅ Thoughtful application as a core principle
The guidance explicitly rejects mechanistic "tick-box" application. It emphasizes that criteria should be adapted to context, purpose, and stakeholders — a critical corrective to how criteria are often misused in practice.
✅ Introduction of coherence
The addition of coherence as a sixth criterion responds to the 2030 Agenda's emphasis on policy integration, synergies, and trade-offs between sectors. This encourages evaluators to look beyond individual interventions to systems-level dynamics.
✅ Emphasis on inclusion and equity
Effectiveness now explicitly calls for analysis of "differential results across groups." Sustainability includes social dimensions alongside financial and environmental. A dedicated gender lens table (Table 3.1) provides practical guidance.
✅ Practical challenges addressed
Each criterion chapter includes a table of common challenges (e.g., data gaps, poorly articulated objectives) with concrete strategies for evaluators and evaluation managers — highly useful for real-world application.
✅ Distinction between effectiveness and impact
The guidance clarifies that effectiveness addresses achievement of stated objectives (outputs and outcomes), while impact examines higher-level, potentially transformative effects — addressing the "so what?" question that is often neglected.
✅ Real-world examples
Boxes throughout illustrate how criteria have been applied in evaluations of budget support, peacebuilding, electoral assistance, land-use planning, and maternal health — across diverse contexts and methodologies.
Critical gaps: What the guidance leaves underdeveloped
⚠️ Power and politics remain undertheorized
While the guidance mentions power dynamics, it does not provide systematic tools for analyzing how power shapes which criteria are prioritized, whose perspectives count, and how evaluative judgements are made. The "leave no one behind" agenda requires deeper engagement with structural inequalities.
⚠️ Participatory application is assumed but not operationalized
The guidance calls for stakeholder involvement in deciding which criteria to use and how to interpret them, but offers limited practical guidance on how to do this in contexts with power imbalances, resource constraints, or where beneficiaries are not organized.
⚠️ Local and national evaluation contexts underrepresented
The Secretariat notes it is "particularly interested in gathering more examples from beyond the EvalNet membership" — a recognition that most examples come from bilateral and multilateral agencies, not from country-led or civil society evaluations. This limits relevance for many evalcommunity members.
⚠️ Trade-offs between criteria are underexplored
While the guidance notes that criteria may conflict (e.g., efficiency vs. equity), it does not offer frameworks for making transparent, principled trade-offs. Evaluators are left to weigh competing criteria without methodological support.
Deep dive: The coherence criterion
The addition of coherence as a sixth criterion is a significant innovation. It addresses the 2030 Agenda's call for policy integration and the recognition that interventions do not operate in isolation. Coherence has two dimensions:
- Internal coherence: Synergies and interlinkages between the intervention and other interventions by the same institution/government, including alignment with international norms and standards.
- External coherence: Consistency with other actors' interventions in the same context, including complementarity, harmonization, coordination, and adding value while avoiding duplication.
The guidance includes practical examples, such as the evaluation of Norwegian efforts to ensure policy coherence for development (Box 4.4), which examined dilemmas between Norway's business engagement in Myanmar and its peacebuilding commitments. This kind of analysis is essential for understanding trade-offs across policy areas — but it also raises a challenge: evaluation offices are rarely mandated to assess policy areas beyond their immediate scope.
Deep dive: Effectiveness and differential results
The revised definition of effectiveness explicitly requires analysis of "differential results across groups." This is a direct response to the 2030 Agenda's "leave no one behind" principle. The guidance explains that this includes examining whether an intervention contributes to or exacerbates equity gaps — a significant departure from earlier definitions that focused only on aggregate achievement of objectives.
Box 4.7 (Evaluation of Australian electoral assistance) illustrates this well, showing how the evaluation used qualitative findings to reveal that social norms and legal barriers led to disproportionately lower voter participation among women and persons with disabilities. This analysis allowed the evaluation to make recommendations focused on increasing equity of results — not just overall effectiveness.
However, the guidance acknowledges a persistent challenge: data disaggregation is often lacking. The challenges table for effectiveness notes that baseline data may be absent, and evaluators may need to reconstruct baselines or use triangulation. For many evalcommunity members working in low-resource settings, this is a daily reality that the guidance addresses only superficially.
Applying a Gender Lens to the Evaluation Criteria
Table 3.1 from the guidance (adapted)
| Criterion | Guiding Questions for Applying a Gender Lens |
|---|---|
| Relevance | Was the intervention designed to respond to the needs of all genders? Does it reflect the rights of persons of all genders? |
| Coherence | Is the intervention coherent with international commitments to gender equality (CEDAW, Beijing Declaration, 2030 Agenda)? |
| Effectiveness | Were there differential results for different genders? Was the theory of change informed by gender analysis? |
| Efficiency | Were resources allocated in ways that considered gender equality? Was differential resource allocation appropriate? |
| Impact | Were there equal impacts for different genders? How did gendered norms and barriers affect outcomes? |
| Sustainability | Did the intervention contribute to greater gender equality within wider systems? Will achievements persist? |
Key takeaways for the evalcommunity
📋 Criteria are lenses, not checklists
The most important message: criteria should be applied thoughtfully and adapted to context, not used mechanistically. Evaluation managers should resist institutional pressure to "cover all criteria" regardless of relevance.
🔗 Coherence enables systems thinking
The new coherence criterion encourages evaluators to look beyond individual interventions to how policies and actors interact — essential for understanding complex change processes and the 2030 Agenda.
⚖️ Differential results matter
Effectiveness and impact evaluations must analyze how results are distributed across groups. Aggregate success can hide exclusion. This requires intentional design of data collection and disaggregation.
⚠️ Trade-offs are inevitable
Efficiency may conflict with equity. Short-term effectiveness may undermine long-term sustainability. Evaluators need frameworks for making transparent, principled trade-offs — not just noting tensions.
🏛️ Institutional adaptation is key
The meso level (institutional guidance) mediates between global norms and individual evaluations. Evaluation policies should encourage thoughtful adaptation, not rigid templates.
📊 Evaluability is a prerequisite
Before applying criteria, assess whether an intervention has clearly defined objectives, a plausible theory of change, and adequate data. If not, be transparent about limitations and manage expectations.
Recommendations for the evalcommunity
- Develop power-conscious criteria application guides — The guidance mentions power dynamics but does not operationalize them. EvalCommunity members could develop practical tools for analyzing whose priorities shape criterion selection and interpretation.
- Document and share local and national evaluation examples — The OECD Secretariat explicitly requests more examples from beyond EvalNet membership. EvalCommunity is well-positioned to collect and disseminate country-led and civil society-led applications of the criteria.
- Create participatory frameworks for stakeholder engagement — The guidance calls for stakeholder involvement in deciding which criteria to use and how to interpret them, but offers limited practical guidance. EvalCommunity members could develop and test participatory methods for this purpose.
- Develop trade-off analysis tools — When criteria conflict (e.g., efficiency vs. equity), evaluators need structured approaches to making transparent, principled trade-offs. This is a priority area for methodological innovation.
- Integrate the criteria with AI and emerging technologies — As the 2026 humanitarian AI research shows, individual adoption of AI tools is outpacing organizational governance. The criteria could provide a framework for assessing AI use in evaluation — ensuring relevance, coherence, effectiveness, efficiency, impact, and sustainability of AI-enhanced evaluation practices.
Frequently asked questions
What are the six OECD DAC evaluation criteria?
The six criteria are: Relevance (doing the right things), Coherence (how well the intervention fits), Effectiveness (achieving objectives), Efficiency (how well resources are used), Impact (what difference the intervention makes), and Sustainability (whether benefits last). They provide a normative framework for determining the merit or worth of an intervention.
Why was coherence added as a new criterion?
Coherence was added in response to the 2030 Agenda and the SDGs, which emphasize policy integration, synergies between sectors, and the need to understand trade-offs. It encourages evaluators to look beyond individual interventions to how policies and actors interact in complex systems.
How is impact different from effectiveness?
Effectiveness focuses on whether an intervention achieved its stated objectives (typically outputs and outcomes). Impact examines higher-level, potentially transformative effects — answering the "so what?" question. Impact looks at ultimate significance, changes in systems or norms, and unintended effects, both positive and negative.
What is the main critique of this guidance?
While the guidance is a significant contribution, it undertheorizes power and politics, offers limited practical guidance for participatory application, and relies heavily on examples from bilateral and multilateral agencies rather than local or national evaluations. The trade-offs between competing criteria are also underexplored.
How can I apply these criteria in my evaluation work?
Start by clarifying the purpose of your evaluation and who will use the findings. Select criteria that are most relevant to that purpose — do not apply all six mechanistically. For each selected criterion, interpret it in light of your specific context, stakeholders, and intervention type. Use the "elements for analysis" as prompts, not sub-criteria. Be transparent about limitations and trade-offs.
Main reference and original source
OECD (2021). Applying Evaluation Criteria Thoughtfully. OECD Publishing, Paris.
https://www.evalcommunity.com/wp-content/uploads/2026/04/OECD-Applying-Evaluation-Criteria-.pdf
Related references:
OECD (2019). Better Criteria for Better Evaluation: Revised Evaluation Criteria Definitions and Principles for Use.
OECD (2002). Glossary of Key Terms in Evaluation and Results Based Management.
Resources for further learning
- EvalCommunity – AI in M&E Course
- OECD DAC Evaluation Network
- Download PDF: Applying Evaluation Criteria Thoughtfully
Advance your skills in evaluation methodology
Learn how to apply the OECD DAC criteria thoughtfully, integrate gender and equity lenses, and navigate trade-offs in complex evaluations. Access courses, templates, and expert guidance.
The courses and articles have been developed by an experienced team of evaluators and software developers under the guidance of Fation Luli. The EvalCommunity Academy combines practical expertise in Monitoring & Evaluation with cutting-edge AI technologies to provide high-quality, accessible learning experiences for professionals around the world.
