
A 10-Step Workflow for M&E Professionals
AI-Assisted Evaluation: A 10-Step Workflow for M&E Professionals
How to use Claude and similar AI tools to reduce evidence-management work while keeping evaluation methodology, judgement, interpretation and accountability firmly with the evaluator.
Important: AI use should be agreed with the client and governed by applicable data-protection, confidentiality, safeguarding, contractual and organisational requirements. Do not upload sensitive evaluation material to an AI service unless the workflow is permitted and appropriately controlled.
Why this workflow matters
Evaluation teams spend substantial time moving information between documents, spreadsheets, interview notes, datasets, reports and presentation decks. That work is necessary, but much of it is not the same thing as evaluation judgement.
AI can potentially reduce repetitive evidence-management work so evaluators can spend more time on:
- methodological decisions and triangulation;
- causal and contribution reasoning;
- contextual, gender, equity and inclusion analysis;
- stakeholder engagement and learning;
- defensible evaluative judgements and recommendations.
The goal is not to make the evaluator disappear from the workflow. It is to make the evaluator’s time more valuable.
The 10-step workflow at a glance
| Step | Workflow | Primary responsibility |
|---|---|---|
| 1 | Design evidence architecture | Evaluator + client |
| 2 | Build evidence base | Evaluation team |
| 3 | Extract and code with AI | AI + human oversight |
| 4 | Check coding | Evaluator |
| 5 | Find gaps and contradictions | AI + evaluator |
| 6 | Validate and close gaps | Evaluator + client |
| 7 | Design report from evidence | Evaluator + client |
| 8 | Generate first-pass draft | AI + evaluator |
| 9 | Interpret and make judgements | Evaluator |
| 10 | Validate, disclose and take responsibility | Evaluator + client |
1. Design the evaluation evidence architecture
Do not begin with the AI tool. Begin with the evaluation design.
Agree with the client:
- evaluation questions;
- DAC criteria or another agreed framework;
- judgement criteria;
- indicators and definitions;
- key assumptions and evidence requirements;
- data sources, methods and sampling;
- triangulation and validation arrangements.
Then create one shared Evaluation Evidence Matrix.
| Evaluation question | Criterion | Judgement criterion | Evidence | Strength / gap |
|---|---|---|---|---|
| Was the intervention relevant? | Relevance | Responded to priority needs | Needs assessment + interviews | Strong / none |
| Did implementation produce intended results? | Effectiveness | Outputs contributed to outcomes | Monitoring + outcome data | Moderate / 2025 data needed |
| Are results likely to continue? | Sustainability | Local capacity can maintain results | Partner interviews + financing evidence | Weak / more evidence |
2. Build the evidence base
Collect the qualitative, quantitative and documentary evidence required by the methodology.
Proposal, logframe, theory of change, workplans and implementation reports.
Indicator data, dashboards, monitoring reports, baseline and endline studies.
KIIs, FGDs, observation, case studies and community feedback.
Policy, market, conflict, political economy and humanitarian evidence.
Keep a clear distinction between raw evidence, analysis, findings and judgements. That distinction becomes more important once AI enters the workflow.
3. Let Claude extract and code evidence
Give the AI a bounded role. Instead of asking it to “analyse the evaluation,” ask it to perform a defined evidence-management task.
Example instruction
Extract evidence relevant to the evaluation questions and judgement criteria in the evidence matrix. For every extract provide the source, location where available, relevant criterion, evaluation question and indicator. Do not infer findings that are not supported by the source. Preserve qualifiers, uncertainty and contradictory evidence.
Useful AI tasks include extracting passages, locating figures, classifying evidence against predefined categories, standardising terminology, identifying duplicates and creating structured evidence records.
Do not remove provenance. Every important extract should be traceable to its original source.
4. Check the coding
AI classification is not self-validating. Review the output.
- Is the extract accurate and in context?
- Is the source location correct?
- Does the evidence belong under the selected question?
- Have qualifications or limitations disappeared?
- Has correlation been turned into causation?
- Has contradictory or minority evidence been omitted?
5. Find gaps, contradictions and evidence tension
Once evidence is coded, AI can act as a second pair of eyes.
Which questions have strong evidence?
Which rely on only one source?
Which remain unanswered?
Where do sources appear to disagree?
Tell the AI not to resolve contradictions automatically. It can surface competing evidence; the evaluator interprets why the evidence differs.
6. Go back to the client and close evidence gaps
Do not wait until the final report to announce that important evidence is missing. Use the matrix as a live evaluation workspace.
| Question | Status | Missing | Action |
|---|---|---|---|
| EQ1 Relevance | Strong | None identified | Proceed |
| EQ2 Effectiveness | Moderate | Latest outcome dataset | Request data |
| EQ3 Sustainability | Weak | Partner perspectives | Additional interviews |
This changes the relationship with the client: they can see what the evaluation knows, what it does not know, and where their input could materially improve the assessment.
7. Design the report from the evidence
Build the report around the evaluation questions rather than starting with a blank document.
Every major finding should answer: What evidence supports this? Every judgement should answer: What does that evidence mean against the agreed criteria?
8. Ask Claude to draft the first-pass report
Once evidence has been validated and the structure agreed, AI can help turn coded evidence into a draft.
Example instruction
Draft the findings section using only the approved evidence matrix. Follow the agreed evaluation structure. Do not introduce new facts or unsupported interpretations. Preserve uncertainty and contradictory evidence. Cite substantive claims to their sources. If evidence is insufficient, state that explicitly.
The output is a drafting aid, not an approved finding, judgement or conclusion.
9. The evaluator interprets and makes the judgement
Read the draft as if you were trying to disprove it.
- What does the evidence actually establish?
- What remains uncertain?
- Are there plausible alternative explanations?
- Does the evidence support attribution, contribution, association—or none?
- Are different groups experiencing different results?
- What unintended effects are visible?
- Does the evidence support the theory of change?
- What contextual factors alter the interpretation?
10. Validate, disclose and take responsibility
Before the evaluation leaves your hands, complete a final quality-assurance pass.
| Check | Question |
|---|---|
| Sources | Can important factual claims be traced to evidence? |
| Numbers | Do figures match the underlying data? |
| Causality | Have causal claims been justified? |
| Equity | Have differences between groups been preserved? |
| Uncertainty | Have limitations and competing explanations been retained? |
| AI use | Has AI use been documented according to the agreed protocol? |
The client should know what role AI played, especially where confidential material or AI-assisted evidence processing was involved. Follow the contract and your organisation’s AI/data-governance requirements.
What AI should do—and what it should not do
| Good AI use | Do not delegate blindly |
|---|---|
| Extract evidence from long documents | Final evaluation ratings |
| Classify evidence against predefined codes | Causal attribution |
| Identify possible gaps | Sensitive safeguarding judgements |
| Surface apparent contradictions | Interpretation of community evidence without context |
| Standardise formatting and terminology | Final recommendations with major resource implications |
| Draft from an approved evidence base | Final conclusions without human verification |
A worked M&E example: livelihoods
Suppose the evaluation question is:
The evidence includes baseline and endline income data, employment status, business survival, training completion, participant interviews, market information, monitoring reports and context on inflation or economic shocks.
AI can organise the evidence. The evaluator still has to ask:
- Did income increase in real or nominal terms?
- Did the programme contribute to the change?
- Were other interventions operating in the same communities?
- Did results differ by gender, disability, geography or livelihood group?
- Could market conditions explain some of the change?
- What evidence supports the mechanism in the theory of change?
This is the difference between evidence organisation and evaluation reasoning.
Where this workflow is especially useful
Monitoring and MEL
Use AI to help reconcile reporting formats, identify missing indicator information, standardise terminology and flag unusual changes for human review. Do not let it silently redefine indicators.
Humanitarian M&E
AI can help organise monitoring reports, post-distribution monitoring, needs assessments and feedback, but protection, safeguarding and personally identifiable information require stronger controls.
Gender, equity and inclusion
AI can help find disaggregated evidence, but evaluators must check whether the underlying data represents different groups adequately and whether averages hide unequal effects.
Research and evidence synthesis
AI can support literature organisation, extraction and comparison, but methodological quality, study design, transferability and strength of evidence require research judgement.
Portfolio and meta-evaluation
Consistent evidence structures can help identify recurring findings, recommendations, implementation challenges and evidence gaps across multiple evaluations.
Build a simple AI evaluation protocol
Before using Claude or another AI tool, agree five things with the client:
- Purpose: What tasks will AI support?
- Data: What information will be processed?
- Boundaries: What decisions remain human-only?
- Quality assurance: How will AI-assisted work be checked?
- Disclosure: How will AI use be communicated?
Prompt patterns evaluators can reuse
Evidence extraction
Extract only evidence relevant to EQ2 and the agreed judgement criteria. For each extract provide source, location, evidence, indicator, evaluation question and limitations. Do not infer a finding.
Contradiction scan
Identify claims that appear to conflict. Show both claims, their sources and context. Do not decide which claim is correct.
Evidence-gap scan
Review the matrix against the evaluation questions and judgement criteria. Identify unanswered questions, single-source findings, missing indicators and claims lacking adequate supporting evidence.
Evidence-based drafting
Draft from the approved evidence matrix only. Preserve uncertainty, limitations and contradictory evidence. Do not introduce facts not contained in the matrix. Cite substantive claims to their sources.
How to quality-assure AI-assisted evaluation work
| QA dimension | Test | Failure to watch for |
|---|---|---|
| Factual accuracy | Check claims against source | Invented or altered facts |
| Completeness | Compare with matrix | Important evidence omitted |
| Provenance | Trace claim to source | Unverifiable statements |
| Interpretation | Compare with evaluator analysis | AI inference presented as evidence |
| Equity | Review subgroup evidence | Averages hide unequal effects |
| Causality | Check design and alternatives | Correlation presented as attribution |
What changes for the evaluator?
Less time: moving text between documents, repetitive extraction, formatting, basic classification and first-pass comparison.
More time: methodology, triangulation, causal reasoning, context, stakeholder engagement, equity analysis, interpretation and judgement.
That is the strategic reason to experiment with AI in evaluation.
A 30-minute M&E team exercise
- 5 minutes: choose one completed evaluation or small evidence folder.
- 5 minutes: identify three evaluation questions and judgement criteria.
- 10 minutes: ask AI to extract and classify evidence without interpreting it.
- 5 minutes: compare the AI output with a human-coded sample.
- 5 minutes: identify where AI saved time and where human judgement was essential.
The objective is to find out where AI is useful in your particular evaluation workflow, not to prove that AI is useful everywhere.
A maturity model for AI-assisted evaluation
| Level | Practice | What changes |
|---|---|---|
| 1 | Ad hoc AI | Individual prompts and limited governance |
| 2 | Defined tasks | Clear boundaries for extraction and drafting |
| 3 | Shared evidence matrix | AI operates against an agreed evaluation architecture |
| 4 | Reusable workflows | Prompts, QA and disclosure become standard practice |
| 5 | Evidence intelligence | Connected evidence supports portfolio learning |
The client-facing benefit
A strong AI-assisted workflow can make evaluation more transparent.
versus
The evidence matrix becomes a transparency mechanism: the client can see what is well supported, what is uncertain and where additional evidence could change the assessment.
The bigger professional shift
The valuable skill is not simply knowing how to ask AI to write a report. It is knowing how to design an evidence architecture, test AI outputs, recognise weak causal reasoning, preserve uncertainty, make defensible judgements and communicate findings to decision-makers.
EvalCommunity’s 10-step model
1. DESIGN — Build the evidence architecture.
2. COLLECT — Assemble qualitative, quantitative and documentary evidence.
3. ORGANISE — Use AI for defined extraction and coding tasks.
4. CHECK — Human-review the AI classification.
5. CHALLENGE — Use AI to surface gaps and contradictions.
6. VALIDATE — Close important evidence gaps with the client and stakeholders.
7. STRUCTURE — Build the report around questions, evidence, findings and judgements.
8. DRAFT — Use AI to create a first-pass evidence-based draft.
9. INTERPRET — The evaluator makes the judgement.
10. ACCOUNT — Validate, disclose AI use and take responsibility for the final evaluation.
Take this question into your next evaluation:
“Which parts of this workflow consume evaluator time without requiring evaluator judgement?”
Those are the tasks worth testing with AI first.
EvalCommunity Academy · Practical AI, evidence, monitoring, evaluation, learning and research resources for M&E, MEL/MEAL, humanitarian and international development professionals.
