
AI Research Protocols
AI Research Protocols for M&E and Development Professionals
A practical guide to using notebook-based AI for evaluation, evidence synthesis, programme learning, research, reporting, and development decision-making without allowing speed to replace professional judgement.
1. Why M&E Professionals Need Research Protocols for AI
AI notebook tools can make it much easier to work with evaluation reports, research papers, programme documents, monitoring documentation, interview materials, policy papers, and other sources.
But a polished answer is not automatically a defensible answer. An AI system can produce a confident, well-written synthesis that is problematic because the underlying research workflow was weak.
For example, you may have:
- asked an ambiguous evaluation question;
- combined sources with different definitions or reporting periods;
- treated repeated claims as independent evidence;
- mixed primary evidence with secondary commentary;
- uploaded several versions of the same report;
- allowed an AI-generated summary to become a de facto source;
- accepted an interpretation because it confirmed an existing assumption.
These are not necessarily hallucination problems. They are research process problems.
2. Prompt vs Protocol
A conventional AI workflow often looks like this:
Upload reports
↓
Ask a question
↓
Receive a summary
↓
Use the summaryA protocol-based workflow adds deliberate checks:
Define the question
↓
Architect the source set
↓
Check source quality and independence
↓
Analyse claims and evidence
↓
Search for contradictions
↓
Test alternative explanations
↓
Stress-test the conclusion
↓
Review uncertainty and provenance
↓
Human judgement
↓
Use or publish the resultThat extra friction is useful in evaluation and development work because it creates opportunities to catch weak assumptions before they become findings, recommendations, or decisions.
3. Protocol 1: Deep Drill
Use when: You need to move beyond a document’s surface summary and examine what supports a claim.
What it fixes: It reduces the risk of accepting a conclusion without examining assumptions, evidence, qualifications, and limitations.
Five-phase workflow:
- Identify the claim. What exactly is the document asserting?
- Trace the evidence. What evidence is presented?
- Surface assumptions. What needs to be true for the claim to hold?
- Identify qualifications. What limitations or conditions are stated?
- Define the boundary. What does the evidence not establish?
M&E example
Instead of:
Summarise whether the programme was effective.
use a structured request:
Identify the main effectiveness claim. Trace the evidence presented for that claim. Identify methodological assumptions and limitations. Identify qualifications and competing explanations. State what the source does not establish.
4. Protocol 2: Radar
Use when: You have many sources and need to distinguish evidence from commentary, repetition, and contradiction.
What it fixes: It reduces the tendency of a notebook to produce a smooth synthesis that hides disagreement between sources.
Ask the AI to classify the material into:
- Primary evidence: direct data, evaluation evidence, documented observations, or original research.
- Secondary interpretation: claims or conclusions drawn from underlying evidence.
- Repeated claims: statements appearing across several sources.
- Independent corroboration: claims supported independently rather than merely repeated.
- Contradictions: sources that disagree in definitions, results, or interpretation.
- Evidence gaps: questions that cannot be answered from the available material.
5. Protocol 3: Crucible
Use when: You are about to publish, recommend, present, or make a decision based on an AI-assisted conclusion.
What it fixes: It challenges the conclusion before another stakeholder, reviewer, or decision-maker does.
- State the conclusion.
- Identify its three weakest logical points.
- Identify evidence that could challenge it.
- Identify alternative explanations.
- Identify assumptions that may not hold.
- Rewrite the conclusion so it matches the strength of the evidence.
Example: avoiding an unsupported causal claim
Initial statement: “The programme improved employment outcomes.”
A Crucible review might ask:
- Is there a credible counterfactual?
- Are baseline and follow-up populations comparable?
- Could labour-market changes explain the result?
- Was exposure to the intervention consistent?
- Does the evaluation design support a causal conclusion?
The resulting statement might instead be: “Employment outcomes increased between baseline and follow-up among programme participants, although the available evidence does not establish that the programme alone caused the change.”
6. Protocol 4: Control Problem
Use when: You are concerned that the AI’s interpretation is becoming your interpretation without enough independent review.
Decision gates:
| Gate | Question | Possible result |
|---|---|---|
| 1 | Is the analytical question clearly defined? | GO / HOLD / BLOCKED |
| 2 | Are the relevant sources identified? | GO / HOLD / BLOCKED |
| 3 | Are definitions and populations comparable? | GO / HOLD / BLOCKED |
| 4 | Is the conclusion supported by the evidence? | GO / HOLD / BLOCKED |
| 5 | Does the conclusion stay within methodological limits? | GO / HOLD / BLOCKED |
If a gate is BLOCKED, resolve the underlying issue instead of asking the AI to continue anyway.
7. Protocol 5: Architect
Use before: Uploading a large collection of material to a notebook or AI workspace.
What it fixes: It prevents weak source architecture from becoming a hidden limitation of the analysis.
Before uploading, define:
- the decision or evaluation question;
- the relevant time period;
- the population and context;
- which sources are primary;
- which sources are secondary;
- which sources are authoritative;
- which sources are duplicates or versions of the same document;
- which materials are AI-generated or analyst-generated;
- which documents are outdated or superseded.
Suggested source architecture
| Layer | Examples | How to treat it |
|---|---|---|
| Primary evidence | Survey results, evaluation analysis, interview findings, administrative data documentation | Core evidence; trace important claims here |
| Programme documentation | Monitoring reports, results frameworks, implementation records | Context and programme evidence |
| External evidence | Research papers, policy studies, sector reports | Context, comparison and triangulation |
| Derived material | Summaries, draft reports, AI-generated notes | Useful working material; not automatically independent evidence |
8. Protocol 6: Compass
Use when: The information environment is changing faster than your team can reasonably track.
Development-sector applications include:
- new AI tools and capabilities;
- changes in donor requirements;
- new evaluation methods;
- new research findings;
- policy changes;
- emerging risks to programme delivery.
For each new development, ask:
- Does it affect our current work?
- Does it change a methodological choice?
- Does it create a material opportunity?
- Does it create a material risk?
- Does it require a workflow change?
- What can we safely ignore for now?
The purpose is not to know everything. It is to identify what is relevant enough to deserve professional attention.
9. Protocol 7: Unblock
Use when: You have collected enough information but cannot begin producing analysis.
A 25-minute approach:
- Choose one question.
- Ask the AI to propose an initial analytical structure.
- Produce one tangible output before searching for more material.
- Mark uncertainty rather than trying to eliminate it immediately.
- Review the output against the original evidence.
Your first output could be:
- three emerging themes;
- a draft evidence matrix;
- five competing explanations;
- a preliminary theory-of-change gap analysis;
- a list of evidence gaps;
- a draft evaluation question hierarchy.
10. Protocol 8: The Advocate
Use when: You already have a strong opinion about a finding, intervention, policy, or recommendation.
What it fixes: It reduces the risk of criticising a weaker version of an argument than the one actually being made.
- State your current verdict.
- Ask the AI to construct the strongest evidence-supported alternative interpretation.
- Identify evidence supporting that interpretation.
- Identify its strongest weaknesses.
- Only then conduct the critique.
This can be especially useful for evaluation recommendations, programme design reviews, policy analysis, and contested findings.
11. Protocol 9: The Threshold
Use when: You are deciding whether a source deserves a permanent place in your research or knowledge base.
Ten questions:
- What specific claim makes this source useful?
- Which evaluation or programme question does it support?
- What population and context does it cover?
- What method produced the evidence?
- What are the main limitations?
- Is it primary or secondary evidence?
- Does it duplicate another source?
- Is it sufficiently current for the intended use?
- What decision could it inform?
- Why should it remain in the knowledge base?
12. Protocol 10: The Ledger
Use when: A notebook or knowledge base contains many generated summaries, reports, or other derived material.
What it fixes: It helps distinguish real corroboration from repeated interpretations that all trace back to the same source.
Claim ↓ Interpretation ↓ Document ↓ Original source ↓ Underlying evidence
This matters when teams repeatedly upload AI-generated summaries, draft reports, or previous syntheses back into the same workspace.
13. Applying the Protocols Across the M&E Cycle
| M&E stage | Potential AI support | Useful protocols |
|---|---|---|
| Programme design | Review assumptions, compare programme logic, identify gaps | Architect, Advocate, Crucible |
| Indicator design | Check definitions, calculation rules, and consistency | Architect, Deep Drill, Control Problem |
| Data collection | Review instruments, metadata, and data-quality documentation | Radar, Deep Drill, Ledger |
| Monitoring | Identify patterns, anomalies, and possible explanations | Radar, Crucible, Control Problem |
| Evaluation | Synthesise findings and assess alternative explanations | Deep Drill, Advocate, Crucible |
| Reporting | Trace claims and draft evidence-supported narratives | Ledger, Deep Drill, Control Problem |
| Learning | Connect findings, lessons, and recommendations | Radar, Advocate, Threshold |
| Adaptive management | Assess emerging evidence and decision relevance | Compass, Radar, Crucible |
14. A Combined Workflow for Evaluation Evidence Synthesis
Suppose your team needs to answer:
- Architect: define the question and organise the source set.
- Radar: distinguish primary evidence, interpretation, repetition, and contradiction.
- Deep Drill: examine the strongest findings and trace their supporting evidence.
- Ledger: trace important claims back to original sources.
- Advocate: construct the strongest alternative explanation.
- Crucible: stress-test the emerging conclusion.
- Control Problem: run decision gates before accepting the conclusion.
- Human review: assess methodological validity, uncertainty, and implications.
Notice what is missing: a step called “Ask the AI whether the conclusion is correct.” AI can support the investigation, but it should not become the final authority on the validity of an evaluation conclusion.
15. AI Research Protocols for Qualitative Evidence
Notebook-based AI can help teams work across interview transcripts, focus group notes, case studies, and qualitative reports. The same protocols apply, but qualitative evidence requires particular care around context and interpretation.
Before analysis
- Define the research question.
- Document the sampling approach and population.
- Identify the coding framework if one exists.
- Separate participant quotations from analyst interpretations.
- Document relevant contextual information.
During analysis
- Ask AI to propose themes, not automatically declare them final.
- Trace themes back to underlying excerpts.
- Look for negative cases and disconfirming evidence.
- Preserve differences between participant groups.
- Check whether a theme is genuinely recurring or simply prominent in a small number of sources.
Before reporting
- Review the interpretation against the original excerpts.
- Check whether context has been lost.
- Check whether quotations support the interpretation being made.
- Distinguish participant perspectives from evaluator conclusions.
16. AI Research Protocols for Quantitative and Monitoring Evidence
AI can assist with quantitative analysis, but structured checks are essential. A notebook should not be allowed to infer meaning from a number without understanding the metric and the dataset that produced it.
Before interpreting a result, verify:
- indicator definition;
- numerator and denominator;
- unit of measurement;
- data grain;
- population;
- reporting period;
- missing-data treatment;
- disaggregation;
- data-quality flags;
- calculation method.
17. AI Research Protocols for Programme and Policy Documents
Development professionals often work with strategies, logframes, implementation plans, donor reports, policy documents, and programme reviews. These documents may use overlapping terminology while carrying different levels of authority.
A useful protocol asks AI to identify:
- the document’s purpose and intended audience;
- the decision or policy question addressed;
- definitions of key terms;
- stated assumptions;
- commitments versus aspirations;
- evidence-backed statements versus assertions;
- implementation responsibilities;
- monitoring and reporting requirements;
- conflicts with other authoritative documents.
18. Provenance: Make the Path Back to the Evidence Visible
Provenance means being able to understand where information came from and how it was transformed. For AI-assisted M&E, provenance is a practical quality-control mechanism.
For an important finding, record where possible:
- original source;
- source date and version;
- dataset or document identifier;
- data collection method;
- cleaning or transformation performed;
- analysis method;
- analyst or responsible team;
- approval or review status;
- report in which the finding was used.
For teams that want a formal provenance model, the W3C PROV-O recommendation provides a standard way to represent provenance relationships. Most M&E teams do not need to implement PROV-O directly; the practical lesson is to keep important results traceable to their sources.
19. FAIR Principles and AI-Ready Evidence
The FAIR principles focus on making data and metadata Findable, Accessible, Interoperable, and Reusable. For M&E teams, that translates into practical questions:
- Can another team understand an indicator without asking its creator?
- Can a dataset be connected to its metadata?
- Can an AI system identify the source of an important result?
- Are important terms defined consistently?
- Can findings be traced to evidence?
- Can another system reuse the information responsibly?
You can learn more from the GO FAIR Principles. You do not need to implement every FAIR principle before using AI. Start by improving the quality, consistency, findability, and traceability of the information your workflows depend on.
20. A Reusable AI Protocol Prompt
The following can be adapted for a notebook-based M&E research workflow:
You are supporting an M&E professional with evidence analysis. Do not treat repeated statements as independent evidence. Distinguish: 1. Primary evidence 2. Secondary interpretation 3. Generated or derived material For important claims: - identify the original source; - identify the reporting period and population where relevant; - identify important methodological limitations; - distinguish evidence from interpretation. Do not: - treat missing information as evidence of absence; - infer causality from descriptive evidence; - silently resolve contradictions; - treat AI-generated summaries as independent corroboration; - force a conclusion when evidence is insufficient. Before producing a final synthesis: 1. State the main evidence-supported findings. 2. Identify the evidence supporting each finding. 3. Identify important conflicting evidence. 4. Identify alternative explanations. 5. Identify what cannot be concluded. 6. Identify remaining evidence gaps. 7. State where human review is required. If the evidence is insufficient, return: HOLD — additional review or evidence required. If the question cannot responsibly be answered from the supplied material, return: BLOCKED — insufficient evidence.
21. A 30-Minute Practice Exercise
Use one real M&E or development question that your team is currently working on.
| Time | Activity | Output |
|---|---|---|
| 0–5 min | Choose the question | A specific, decision-relevant question |
| 5–10 min | Architect the sources | Primary, secondary, external, and derived source categories |
| 10–15 min | Run Radar | Repeated claims, contradictions, and evidence gaps |
| 15–20 min | Run Deep Drill | Claim → evidence → method → limitation → uncertainty |
| 20–25 min | Run Advocate | Strongest alternative explanation |
| 25–30 min | Run Crucible + Control Problem | GO, HOLD, or BLOCKED decision |
22. A Protocol Selection Guide
| Your situation | Start with |
|---|---|
| One important report needs deeper analysis | Deep Drill |
| 20+ documents are producing a blurry synthesis | Radar |
| You are about to publish a major finding | Crucible |
| You are unsure whether you are accepting AI reasoning too quickly | Control Problem |
| You are preparing a new notebook or evidence workspace | Architect |
| Too many new developments are competing for attention | Compass |
| You are stuck in research collection | Unblock |
| You strongly disagree with a finding or policy | The Advocate |
| Your knowledge base contains too many sources | The Threshold |
| Your AI workspace contains many generated summaries | The Ledger |
23. What These Protocols Do Not Solve
A good protocol cannot turn weak evidence into strong evidence.
It cannot compensate for a poorly designed evaluation, a biased sample, an invalid indicator, missing data, an inappropriate analytical method, or inadequate contextual understanding.
It can, however, make those limitations easier to see and harder to overlook.
AI research protocols also do not remove the need to consider:
- confidentiality and data protection;
- participant privacy;
- research ethics;
- bias and representation;
- methodological validity;
- reproducibility;
- human oversight;
- appropriate governance for consequential decisions.
24. The Bigger Lesson for AI-Assisted Development Work
AI readiness is often framed as a question of which model or tool an organisation should adopt.
For many M&E and development teams, a more fundamental question comes first:
If the answer is no, adding a more capable AI system may simply make the existing workflow faster.
Better AI-assisted research starts with better research architecture: clear questions, organised sources, explicit definitions, provenance, verification points, and human judgement.
25. Practical Next Step
Pick one recurring workflow this week: an evaluation synthesis, donor report, literature review, data-quality investigation, programme review, policy analysis, or learning exercise.
Do not start by asking the AI for the final answer. Start with:
- Architect: What information should be in the workspace?
- Radar: What is evidence, interpretation, repetition, and contradiction?
- Deep Drill: What supports the most important claim?
- Advocate: What is the strongest alternative explanation?
- Crucible: What could make this conclusion fail?
- Control Problem: Is the work ready to proceed?
The goal of an AI research protocol is not to make AI sound more confident. It is to make the path from source → evidence → interpretation → conclusion more visible, reviewable, and defensible.
Further EvalCommunity Resources
Continue building practical AI and M&E capability with these resources:
Explore the AI Toolkit for Evaluators
Browse the M&E Framework Library
Review the Principles for AI Use in M&E
Explore the AI for M&E Professional Bundle
EvalCommunity Academy · Practical AI workflows for monitoring, evaluation, learning, evidence, and international development.
