How to Use AI to Improve Evaluation Design
EvalCommunity Tutorial
How to Use AI to Improve Evaluation Design
A practical tutorial for evaluators and M&E professionals on using AI to sharpen evaluation questions, test assumptions, build stronger evaluation matrices, and design responsible AI-assisted workflows.
Tutorial Summary
Many evaluation problems begin before data collection. They begin when evaluation questions are unclear, theories of change contain untested assumptions, indicators are weak, or evidence needs are unrealistic.
This tutorial shows how to use AI practically and responsibly during the design stage: not to replace the evaluator, but to support sharper thinking, better structure, stronger evidence planning, and more useful evaluation products.
Who This Tutorial Is For
- Evaluators and evaluation consultants
- M&E, MEAL, MEL, and results-based management professionals
- Programme managers and technical advisers
- Researchers, learning specialists, and knowledge managers
- NGO, donor, government, UN, and development teams using AI responsibly
What You Will Learn
Clarify Questions
Use AI to improve evaluation questions so they are specific, answerable, and useful.
Test Assumptions
Surface hidden theory of change assumptions and decide which ones should be tested.
Build AI Workflows
Create a simple AI-assisted design workflow using prompts, tools, and optional agents.
Protect Judgement
Know what AI can support and what must remain human-led.
1. Why AI Can Help at the Evaluation Design Stage
Evaluation design is where the quality of an evaluation is shaped. Before fieldwork begins, the evaluator has already made important decisions about questions, scope, criteria, theory of change, indicators, methods, data sources, ethics, and intended use.
AI can help at this stage because it can review structure, detect gaps, compare options, identify assumptions, and organize information quickly. But AI should be treated as a design support tool, not as the evaluator.
Key Principle
Use AI to challenge and improve your design. Do not use AI to replace ethical judgement, stakeholder engagement, contextual knowledge, or professional accountability.
2. The Evaluation Design Tasks AI Can Support
AI can support several design tasks when the evaluator provides enough context and reviews the output critically.
Evaluation Question Review
AI can identify questions that are too broad, duplicated, output-focused, unanswerable, or not clearly linked to decisions.
Theory of Change Testing
AI can make implicit assumptions visible and help evaluators decide which assumptions are most important to test.
Evaluation Matrix Review
AI can check whether questions, indicators, data sources, methods, and analysis plans are logically connected.
Risk and Feasibility Checks
AI can help flag risks related to data quality, access, ethics, sensitivity, inclusion, and use of findings.
Stakeholder Use Mapping
AI can help map which evaluation questions are most relevant to donors, implementers, communities, policymakers, or learning teams.
Reporting Structure
AI can suggest how findings could be organized around decisions, learning needs, criteria, or theory of change pathways.
3. Before You Use AI: Prepare the Evaluation Context
The quality of AI support depends on the quality of the context you provide. Generic input produces generic output. Before using AI, prepare a short design brief.
Evaluation Context Brief Template
- Evaluation purpose: What decision, learning need, or accountability requirement will this evaluation support?
- Programme description: What intervention, policy, project, or system is being evaluated?
- Target groups: Who is expected to benefit, participate, or be affected?
- Geographic and institutional context: Where does the programme operate and under what conditions?
- Stakeholders: Who will use the findings?
- Theory of change: What change pathway is assumed?
- Available data: What monitoring, administrative, survey, qualitative, or secondary data exists?
- Constraints: What time, budget, access, ethical, security, or data limitations exist?
- Sensitive issues: Are there risks related to harm, privacy, power, exclusion, conflict, or politics?
4. Prompt 1: Improve Evaluation Questions
Use this prompt when you already have draft evaluation questions and want AI to review them.
Prompt Template: Evaluation Question Reviewer
Copy, adapt, and paste:
You are supporting an evaluator during the evaluation design stage. Treat all information I provide as context, not as final conclusions.
Review the evaluation questions below and assess them against five criteria: clarity, answerability, usefulness, feasibility, and ethical sensitivity.
For each question, tell me:
- Whether the question is clear or needs refinement.
- Whether it can realistically be answered with available evidence.
- What data sources would be needed.
- Whether the question duplicates another question.
- Whether the wording contains hidden assumptions.
- How the question could be improved.
Here is the evaluation context: [paste short context brief]. Here are the draft evaluation questions: [paste questions].
Example of a Weak Question and Better Versions
Weak question: Did the programme work?
Better version 1: To what extent did the programme contribute to changes in knowledge, behaviour, access, or service use among the intended participants?
Better version 2: Which programme components appear most strongly associated with observed outcomes, and what evidence supports this contribution?
Better version 3: For which groups did the programme appear more or less effective, and what contextual factors may explain these differences?
5. Prompt 2: Identify Hidden Assumptions in a Theory of Change
Use this prompt when you want AI to examine the logic behind a theory of change, results chain, logframe, or programme strategy.
Prompt Template: Theory of Change Assumption Mapper
You are supporting an evaluator to review a theory of change. Identify the assumptions that must hold for the programme logic to work.
Please organize the assumptions into the following categories:
- Assumptions about participant behaviour.
- Assumptions about institutional capacity.
- Assumptions about service quality or implementation quality.
- Assumptions about inclusion and equity.
- Assumptions about political, social, or economic context.
- Assumptions about sustainability after support ends.
For each assumption, provide:
- A plain-language statement of the assumption.
- Why it matters for the theory of change.
- What evidence could test it.
- Whether it appears low, medium, or high risk.
Here is the theory of change or programme logic: [paste theory of change, logframe, results framework, or description].
Practical Tip
Do not try to test every assumption. Prioritise assumptions that are central to the change pathway, uncertain in the context, and likely to weaken the evaluation findings if ignored.
6. Prompt 3: Build or Review an Evaluation Matrix
The evaluation matrix should connect questions, sub-questions, indicators, evidence sources, methods, analysis, and limitations. AI can help review whether the matrix is coherent and realistic.
Prompt Template: Evaluation Matrix Builder
You are helping me build an evaluation matrix. Use the evaluation questions below and propose a structured matrix.
For each evaluation question, include:
- Sub-questions.
- Potential indicators or evidence needs.
- Possible data sources.
- Suitable methods.
- Possible disaggregation.
- Known limitations.
- Ethical considerations.
- How the finding could be used.
Here is the evaluation purpose: [paste purpose]. Here are the evaluation questions: [paste questions]. Here are the data and access constraints: [paste constraints].
7. Prompt 4: Review Feasibility, Risks, and Ethics
A design can look strong on paper but fail in practice if data, access, ethics, or politics are not considered early.
Prompt Template: Feasibility and Risk Reviewer
You are reviewing an evaluation design for feasibility and ethical risk. Identify risks that could affect data quality, access, respondent safety, stakeholder trust, and use of findings.
Please organize the review under these headings:
- Data availability and data quality risks.
- Access and sampling risks.
- Ethical and safeguarding risks.
- Inclusion and equity risks.
- Political or institutional sensitivity.
- Risks to use of findings.
- Mitigation actions.
Here is the evaluation design summary: [paste summary]. Here are the planned methods: [paste methods]. Here are the known constraints: [paste constraints].
8. Suggested AI Agents for Evaluation Design
Instead of using one general AI assistant for everything, evaluators can create a small set of specialised AI agents. Each agent has a clear role, boundaries, and output format.
Agent 1: Question Reviewer
Reviews evaluation questions for clarity, answerability, duplication, usefulness, and hidden assumptions.
Output: question quality table, revised questions, evidence needs.
Agent 2: Assumption Mapper
Identifies explicit and implicit assumptions in theories of change, logframes, and results chains.
Output: assumption map, risk ranking, evidence to test.
Agent 3: Matrix Builder
Connects questions, sub-questions, indicators, data sources, methods, and limitations.
Output: draft evaluation matrix.
Agent 4: Ethics and Risk Checker
Reviews sensitive topics, privacy risks, power dynamics, inclusion risks, and possible harms.
Output: risk register and mitigation checklist.
Agent 5: Evidence Planner
Suggests evidence sources, triangulation options, and possible data gaps.
Output: evidence plan and data gap list.
Agent 6: Use and Reporting Planner
Links evaluation questions to decision-makers, reporting products, and learning moments.
Output: use plan, reporting outline, stakeholder products.
9. Agent Instruction Template
Use this template when creating a custom GPT, Claude Project, Gemini Gem, OpenAI agent, n8n AI Agent, Make AI Agent, Zapier Agent, or internal evaluation assistant.
Reusable Agent Instructions
You are an AI assistant supporting evaluation design. Your role is to help evaluators think more clearly, not to make final decisions.
Always follow these rules:
- Ask for or use the evaluation context before giving recommendations.
- Separate facts, assumptions, risks, and recommendations.
- Flag uncertainty instead of pretending to know the context.
- Do not invent evidence, citations, stakeholders, indicators, or findings.
- Do not process confidential or personal data unless the user confirms the tool is approved for that purpose.
- Provide structured outputs that evaluators can review and adapt.
- Keep final judgement, ethics, and accountability with the human evaluator.
When reviewing evaluation content, use headings, checklists, and practical suggestions. Always include limitations and questions for human follow-up.
10. Tools and Connectors for AI-Assisted Evaluation Design
Different tools can support different parts of the evaluation workflow. The choice depends on your organization’s data protection rules, budget, technical capacity, and risk level.
AI Assistants
Useful for drafting, reviewing, structuring, and challenging evaluation design content.
- ChatGPT or Custom GPTs
- Claude Projects
- Gemini or Gems
- Microsoft Copilot
Research and Evidence Tools
Useful for literature scanning, evidence extraction, source comparison, and briefing notes.
- Elicit
- NotebookLM
- Consensus
- Perplexity
Field Data Systems
Useful when AI-supported design needs to connect with survey or field data systems.
- KoboToolbox
- ODK Central
- SurveyCTO
- Google Forms or Microsoft Forms
Automation and Agents
Useful for connecting forms, documents, spreadsheets, databases, AI agents, and notifications.
- n8n AI Agent node
- Make AI Agents
- Zapier Agents
- OpenAI Agents SDK for technical teams
Knowledge and Storage
Useful for storing terms of reference, evaluation matrices, guidance notes, and templates.
- Google Drive
- SharePoint or OneDrive
- Airtable
- Notion
Dashboards and Reporting
Useful for turning structured evaluation data into visual summaries and decision products.
- Looker Studio
- Power BI
- Tableau
- Google Sheets or Excel
11. Example AI-Assisted Evaluation Design Workflow
Below is a practical workflow that an evaluation team can adapt. Start simple before building automation.
Manual Version: No Automation Needed
- Prepare a non-confidential evaluation context brief.
- Paste the draft evaluation questions into an approved AI tool.
- Use the Evaluation Question Reviewer prompt.
- Revise the questions manually.
- Paste the theory of change into the AI tool.
- Use the Assumption Mapper prompt.
- Select the assumptions that matter most.
- Use the Matrix Builder prompt to create a draft evaluation matrix.
- Use the Feasibility and Risk Reviewer prompt.
- Validate the revised design with stakeholders and technical reviewers.
Semi-Automated Version: Using Forms, Sheets, and AI
- Create a design intake form for programme teams.
- Collect programme purpose, target groups, context, theory of change, draft questions, data sources, and constraints.
- Send form responses to Google Sheets, Excel, Airtable, or another approved database.
- Use an automation tool such as Zapier, Make, or n8n to trigger an AI review.
- Route the review through specialised agents: question reviewer, assumption mapper, matrix builder, and risk checker.
- Save AI outputs in a structured review document.
- Notify the evaluator or team lead for human review.
- Require manual approval before any output is added to the final evaluation design.
Important Safeguard
Do not connect AI agents directly to sensitive datasets, beneficiary data, confidential proposals, or decision systems unless your organization has approved the tool, access controls, security settings, and human review process.
12. Practical Exercise: Build an Evaluation Design Review Workflow
This exercise helps evaluators practice using AI across the full design process. Use a non-confidential example or a simplified version of a real evaluation.
Exercise Goal
Create a reviewed evaluation design package that includes improved questions, visible assumptions, a draft evaluation matrix, a risk checklist, and a human validation plan.
Step 1
Choose a Case
Use a non-confidential project, programme, policy, or training initiative.
Step 2
Write the Purpose
State why the evaluation is needed and what decision it will inform.
Step 3
Draft Questions
Write 5 to 8 draft evaluation questions before using AI.
Step 4
Run Question Review
Use the prompt to identify weak, duplicated, or unclear questions.
Step 5
Map Assumptions
Use AI to identify the assumptions behind the change pathway.
Step 6
Rank Assumptions
Select the assumptions that are most important and uncertain.
Step 7
Build Matrix
Use AI to create a first draft evaluation matrix.
Step 8
Check Risks
Review data, ethics, access, inclusion, and use risks.
Step 9
Revise Manually
Accept, reject, or rewrite AI suggestions based on professional judgement.
Step 10
Validate with People
Review the design with stakeholders, field teams, and technical experts.
13. Example Exercise Inputs
Sample Evaluation Case
Programme: A two-year capacity-building initiative supporting local organizations to improve service delivery and community feedback systems.
Evaluation purpose: To understand whether the initiative strengthened organizational capacity, improved responsiveness to community feedback, and generated lessons for a possible second phase.
Draft question: Did the project improve local capacity?
Possible improved question: To what extent did participating organizations improve their ability to collect, analyze, and respond to community feedback, and what factors supported or limited this change?
14. Suggested Output Format for AI Reviews
Ask AI to produce structured outputs so the evaluation team can review them quickly.
Prompt Add-On: Output Format
Please present your response in this structure:
- Summary of main design strengths.
- Top design weaknesses.
- Questions that should be revised.
- Assumptions that should be tested.
- Evidence gaps.
- Ethical or feasibility risks.
- Suggested revisions.
- Issues requiring human judgement.
15. Responsible AI Rules for Evaluation Teams
- Do not upload confidential, identifiable, or sensitive data into tools that are not approved by your organization.
- Use AI outputs as drafts, not as final evaluation products.
- Check all AI suggestions against programme documents, stakeholder input, and contextual knowledge.
- Record when and how AI was used in the design process.
- Keep final design decisions with human evaluators.
- Be especially careful with evaluations involving children, survivors of violence, refugees, conflict-affected communities, marginalized groups, or sensitive political contexts.
16. Frequently Asked Questions
Can AI write evaluation questions?
AI can help draft and improve evaluation questions, but evaluators must decide which questions are relevant, ethical, feasible, and useful for decision-making.
Can AI build an evaluation matrix?
AI can create a draft matrix and check alignment between questions, indicators, data sources, and methods. Human review is required to validate feasibility and quality.
Can AI agents be connected to field data systems?
Yes, but only with appropriate approvals, security controls, and safeguards. Field data may include sensitive information, so human oversight and data protection rules are essential.
What is the safest first step?
Start with non-confidential documents and manual AI-assisted review. Build automation only after the team has clear rules, approved tools, and a human validation process.
17. Final Checklist
- The evaluation purpose is clear.
- The intended users and decisions are identified.
- The evaluation questions are specific and answerable.
- The theory of change assumptions are visible.
- High-risk assumptions are prioritised for testing.
- The evaluation matrix connects questions, evidence, methods, and analysis.
- Data, access, ethics, inclusion, and use risks are reviewed.
- AI outputs are reviewed by humans.
- Confidential data is protected.
- Final decisions remain with accountable evaluators.
Conclusion
AI can make evaluation design stronger when evaluators use it to test logic, improve questions, surface assumptions, plan evidence, and identify risks.
But AI should not replace the evaluator. Evaluation remains a human, ethical, professional, and contextual practice.
The strongest use of AI is not to automate judgement. It is to support better judgement.
