
Use AI to Think Better — Not Just Work Faster
How M&E Professionals Can Use AI to Think Better — Not Just Work Faster
A practical guide to using AI as a research assistant, tutor, critical reader, red team and thinking partner across monitoring, evaluation, learning and international development work.
Before you use AI: follow the applicable client, organisational and professional requirements for confidentiality, data protection, safeguarding, research ethics and disclosure of AI use. Do not upload sensitive or identifiable information unless the relevant controls and approvals are in place.
What you will learn
- How to use AI to improve the questions you ask, not just the speed of your outputs.
- How to use AI for evaluation design, evidence analysis, learning and reporting.
- How to stress-test assumptions, findings and recommendations.
- How to keep evidence, interpretation and professional judgement distinct.
- How to introduce AI into an evaluation team without making it an invisible or uncontrolled workflow.
1. Start with the right question
A common starting point is:
For an evaluator, a better question is:
That shift changes the role of AI. It becomes useful for scanning, comparison, explanation, questioning, red-teaming and drafting, while the evaluator remains responsible for methodological choices, interpretation, causal reasoning, evaluative judgement and final recommendations.
| AI can assist with | Evaluator remains responsible for |
|---|---|
| Scanning and organising information | Deciding what evidence is fit for purpose |
| Explaining unfamiliar concepts | Methodological choices |
| Generating alternative explanations | Interpretation and evaluative judgement |
| Finding possible gaps or contradictions | Causal and contribution claims |
| Drafting and restructuring | Final conclusions and recommendations |
2. Improve your information diet and learn just in time
M&E professionals work across donor guidance, evaluation reports, research, policy documents, monitoring dashboards, learning products, situation reports and programme data. The problem is often not a lack of information. It is too much information.
Use AI to curate information and connect learning to an immediate task instead of collecting resources indefinitely.
Reusable prompt
Identify the three developments in M&E, evaluation methodology or international development that are most relevant to my current work. For each, explain what changed, why it matters for evaluators, what evidence supports the claim, and one question I should investigate further.
The last instruction matters. A useful briefing should create a better question, not close the thinking process.
3. Turn AI into an M&E tutor
If you want to learn, do not let AI do all the cognitive work. Ask it to question you, wait for your response, identify what you missed and gradually increase the difficulty.
Act as my M&E tutor. I am learning [TOPIC]. Ask me one question at a time. Do not give me the answer before I respond. After each response: 1. Tell me what I got right. 2. Identify what I missed. 3. Challenge weak reasoning. 4. Give a short explanation where needed. 5. Ask a harder question next. Use practical examples from international development and evaluation.
Try this with Theory of Change, contribution analysis, process tracing, DAC criteria, outcome harvesting, qualitative coding, sampling, mixed methods or gender-responsive evaluation.
4. Red-team the evaluation before fieldwork
Once an evaluation question, Theory of Change or hypothesis feels convincing, it becomes easy to look mainly for supporting evidence. Use AI as a critical reviewer before you commit time and resources.
Act as a critical evaluation adviser. Here is my evaluation design: [PASTE DESIGN] Do not praise it. Identify: 1. The three weakest assumptions. 2. The most important unanswered evaluation question. 3. Where the methodology may not answer the stated questions. 4. Where the design risks confirmation bias. 5. Which stakeholder perspective is missing. 6. What evidence could falsify the emerging Theory of Change. For each issue, explain why it matters and propose one way to test it.
This is more useful than asking, “Is my evaluation design good?” The latter invites reassurance; the former asks for scrutiny.
5. Interrogate the Theory of Change and its assumptions
A Theory of Change can look convincing because the arrows connect neatly. The harder question is what assumptions sit underneath those arrows.
Review this Theory of Change as a critical evaluator. For every major causal link: 1. State the assumption required for the link to hold. 2. Identify what could make the assumption false. 3. Identify alternative explanations. 4. Identify evidence that would support the link. 5. Identify evidence that would weaken or falsify it. Pay particular attention to power, gender, institutional incentives, context, implementation quality, unintended effects and external shocks.
Then decide which assumptions actually deserve investigation. AI can widen the field of view; the evaluator decides what is substantively important.
6. Stress-test findings, causal claims and recommendations
One of the most common analytical problems in evaluation is moving too quickly from evidence to judgement.
Evidence: 82% of participants completed the training.
Potential judgement: The training was effective.
The second statement does not automatically follow from the first. Ask AI to separate direct evidence, interpretation, inference, causal claims and the additional evidence needed for a defensible judgement.
Here is an evaluation finding: [INSERT FINDING] Separate: 1. What is directly supported by evidence? 2. What is interpretation? 3. What is inference? 4. What is a causal claim? 5. What alternative explanations exist? 6. What additional evidence would strengthen the conclusion? 7. What would be needed for a defensible evaluative judgement?
Apply the same discipline to recommendations.
Act as a sceptical programme director reviewing these recommendations. For each recommendation: 1. What evidence supports it? 2. What evidence does not support it? 3. What assumption does it depend on? 4. Who would need to act? 5. What resources might implementation require? 6. What unintended consequences could result? 7. What could make it fail? 8. How could it be made more specific and actionable? Critique first. Do not rewrite until the critique is complete.
7. Improve evaluation questions before collecting data
AI can act as a design critic before you spend time and money on fieldwork.
Review these proposed evaluation questions. For each question assess: - clarity - answerability - evaluability - evidence requirements - methodological implications - overlap with other questions - stakeholder usefulness - risk of producing description rather than evaluation. Identify problems in the original wording before proposing improvements.
The important step is making AI explain why a question is weak, rather than simply rewriting it.
8. Use AI to look for equity, inclusion and humanitarian blind spots
AI can help identify questions you may have overlooked, but equity analysis cannot be reduced to asking for a paragraph about gender or inclusion.
Ask whether the design could miss differences in access, outcomes, participation, voice or unintended effects across relevant groups. Then assess whether your sampling, indicators and methods can actually support the analysis.
Review this evaluation design for potential blind spots affecting gender, disability, age, geography, socioeconomic status and other relevant groups. Identify: - whose access may differ; - whose outcomes may differ; - whose voice may be missing; - which indicators may hide unequal effects; - what assumptions may affect different groups differently. Do not assume every category is relevant. Explain why each suggested line of inquiry may or may not matter.
In humanitarian settings, the same principle applies alongside stronger safeguards for protection information, personally identifiable information, survivor information and sensitive community feedback.
9. Use AI across the M&E cycle
| Activity | Useful AI role | Human check |
|---|---|---|
| MEL framework | Identify ambiguities and overlaps | Indicator validity and programme logic |
| Evaluation design | Red-team assumptions and methods | Methodological appropriateness |
| Data collection | Generate probes and interview questions | Ethics, cultural and contextual fit |
| Analysis | Surface patterns and contradictions | Meaning, context and causal reasoning |
| Reporting | Structure and draft | Findings, judgements and recommendations |
| Learning | Identify recurring themes and questions | Decide what should change in practice |
10. Keep the evidence–interpretation boundary visible
Evidence: What does the source directly show?
Interpretation: What could the evidence mean?
Inference: What conclusion is being drawn beyond the direct observation?
Judgement: How does the evidence perform against agreed criteria?
Recommendation: What action follows, and why?
Keeping these layers separate makes it easier to spot when an AI-generated draft has moved beyond what the evidence can support.
11. Create a simple team protocol for AI use
For commissioned evaluations, AI use should not be an invisible individual habit. Agree the boundaries before the work begins.
| Question | What to define |
|---|---|
| Purpose | Which tasks will AI support? |
| Data | What information may be processed? |
| Boundaries | Which decisions remain human-only? |
| Quality assurance | How will AI-assisted work be checked? |
| Disclosure | How will AI use be communicated to the client? |
A good protocol should also specify who reviews AI-assisted outputs, where prompts or processing records are retained when appropriate, and how sensitive information is handled.
12. Measure whether AI actually improved the work
Do not evaluate AI only by hours saved. Ask whether it improved the quality of the evaluation process.
- Did the team identify more relevant evidence?
- Did it find contradictions earlier?
- Did evaluation questions become sharper?
- Were hidden assumptions surfaced?
- Were alternative explanations considered?
- Did the reasoning become clearer?
- Did the team learn something it would otherwise have missed?
- Did findings and recommendations become more useful?
13. A practical weekly AI routine for M&E professionals
| Day | Practice | Goal |
|---|---|---|
| Monday | Scan | Improve your information inputs |
| Tuesday | Learn | Practise one concept through questions |
| Wednesday | Challenge | Test an assumption in current work |
| Thursday | Analyse | Stress-test an emerging finding |
| Friday | Reflect | Identify where AI improved or weakened thinking |
14. The EvalCommunity AI thinking model
1. INPUT — Improve the quality of information entering your workflow.
2. LEARN — Use AI as a tutor rather than an answer machine.
3. CHALLENGE — Ask AI to argue against your assumptions.
4. TEST — Stress-test findings, theories and recommendations.
5. CREATE — Use AI after the problem, evidence and audience are clear.
6. REFLECT — Ask whether AI actually improved your thinking.
15. Final exercise: ask a better question
Take one task from your current M&E work: an evaluation question, Theory of Change, interview guide, emerging finding, recommendation, MEL framework or research proposal.
Before asking AI to produce anything, ask:
Then ask:
- What assumption am I making?
- What evidence would prove me wrong?
- Whose perspective is missing?
- What alternative explanation should I investigate?
- What would a sceptical evaluator challenge?
- What still requires my professional judgement?
EvalCommunity takeaway:
Don’t make AI your answer machine. Make it your research assistant, tutor, critical reader, red team, evidence organiser and sparring partner — while remaining the evaluator.
EvalCommunity Academy · Practical resources for monitoring, evaluation, learning, research, humanitarian M&E and international development professionals.
