How to Commission Evaluations in the AI Era
EvalCommunity Tutorial
How to Commission Evaluations in the AI Era
A practical guide for evaluation commissioners, M&E managers, procurement teams, donors, and consultants on managing responsible AI use with contractors.
Tutorial Summary
AI tools, especially large language models, are becoming part of evaluation work. They can support drafting, summarising, translation, evidence organisation, coding support, data cleaning scripts, and report preparation.
But AI use in evaluation also creates risks for validity, confidentiality, stakeholder trust, cultural interpretation, transparency, accountability, and human judgement. This tutorial explains how commissioners and contractors can agree on responsible AI use before and during an evaluation assignment.
Why this matters
Evaluation commissioners often work with external consultants, research teams, or service providers. In the AI era, commissioners need to know not only what the contractor will deliver, but also how the contractor may use AI during the evaluation process.
Hidden or unmanaged AI use can affect evaluation quality. It may create polished but unsupported text, weak citations, unclear analysis, privacy risks, or outputs that are difficult to audit. These risks are especially important when evaluations involve public funds, sensitive information, affected communities, vulnerable groups, or policy decisions.
Core message
AI use in commissioned evaluations should be discussed, classified, documented, reviewed, and disclosed. It should not remain a hidden part of the contractor’s workflow.
Who should use this tutorial?
Evaluation commissioners
Public institutions, donors, NGOs, foundations, and agencies commissioning evaluation studies or managing external evaluation contracts.
M&E and MEAL managers
Professionals responsible for monitoring, evaluation, accountability, learning, evidence synthesis, reporting, and programme quality.
Evaluation consultants and contractors
Individuals, firms, and research teams using or considering AI tools in evaluation design, data analysis, reporting, quality assurance, and communication.
Procurement and quality assurance teams
Teams responsible for drafting terms of reference, selecting suppliers, reviewing inception reports, checking deliverables, and managing data protection responsibilities.
What AI can support in evaluation work
AI can be useful when tasks are well-defined, when data is appropriate, and when outputs are reviewed by humans with evaluation expertise.
Examples of supportive uses
- Improving readability of non-final text.
- Summarising public or non-sensitive documents.
- Translating working materials for internal review.
- Drafting spreadsheet formulas or simple scripts for data cleaning.
- Creating outlines for discussion.
- Organising evidence against an evaluation matrix.
- Preparing workshop materials or communication drafts.
The key condition is that AI should support the evaluation process, not replace evaluative reasoning. Human evaluators remain responsible for the evaluation design, evidence interpretation, conclusions, recommendations, and final report.
What AI should not replace
Evaluation is not only a process of organising information. It also involves judgement. Evaluators must decide what matters, what evidence is credible, what findings mean in context, and what conclusions are fair and useful.
AI should not replace human responsibility for:
- Defining evaluation purpose and scope.
- Developing evaluation questions.
- Choosing methods and sampling strategies.
- Collecting empirical data from people in sensitive contexts.
- Interpreting sensitive or contested findings.
- Making evaluative judgements.
- Writing final conclusions and recommendations without human validation.
- Processing confidential, personal, or sensitive data in public AI tools.
Step-by-step process for commissioners
The following process can be used before, during, and after an evaluation assignment. It is designed for practical use by commissioners, M&E managers, donors, procurement teams, and quality assurance reviewers.
Step 1: Add AI expectations to the Terms of Reference
Include a short section in the Terms of Reference explaining that contractors must disclose planned AI use, protect data, verify AI outputs, and avoid using AI for restricted tasks.
Example wording for a Terms of Reference:
“The evaluator must disclose any planned use of AI tools during the evaluation. AI tools must not be used to process personal, confidential, or sensitive data unless explicitly approved in writing and supported by appropriate safeguards. All AI-assisted outputs must be reviewed by the evaluation team, and final findings, conclusions, and recommendations remain the responsibility of human evaluators.”
Step 2: Ask bidders to describe planned AI use
During proposal review, ask bidders to explain whether AI will be used, for which tasks, with which tools, and under what safeguards. This helps prevent hidden AI use later in the assignment.
Step 3: Classify AI use during the inception phase
During inception, agree which AI uses are low-risk, which require supervision, and which are restricted. This classification should be documented in the inception report or quality assurance plan.
Step 4: Create an AI-use agreement
The commissioner and contractor should agree on approved tools, allowed tasks, restricted data, review responsibilities, disclosure requirements, and procedures for reporting AI-related errors.
Step 5: Keep an AI-use log
The contractor should maintain a simple log showing what AI was used for, what data was entered, what outputs were produced, who reviewed them, and how the outputs influenced the evaluation.
Step 6: Review AI-supported outputs before final reporting
Do not wait until the final report. Review AI-supported work during implementation, especially literature summaries, coding outputs, evidence tables, draft findings, synthesis notes, and report sections.
Step 7: Require final disclosure
The final report should disclose how AI was used, how outputs were reviewed, and which parts of the evaluation remained under human responsibility.
Practical AI-use classification
Commissioners can use this practical classification during the inception phase. It helps decide whether an AI use is acceptable, requires additional safeguards, or should be avoided.
Level 1: Routine support
These uses are generally lower risk when they involve public or non-sensitive material and human review.
- Editing grammar or improving readability.
- Summarising public documents.
- Drafting formulas, code snippets, or document outlines.
- Preparing non-final communication materials.
Required action: Declare AI use and check the output before use.
Level 2: Supervised analytical support
These uses may be helpful, but the quality of the result depends strongly on human expertise, method quality, transparent review, and source verification.
- Supporting literature screening.
- Assisting with preliminary coding of non-sensitive qualitative data.
- Organising evidence against an evaluation matrix.
- Testing consistency of draft findings.
- Supporting synthesis across large document sets.
Required action: Agree on oversight, validation, documentation, and a fallback plan before use.
Level 3: Restricted or avoid-use tasks
These uses can create serious risks for validity, confidentiality, accountability, or stakeholder protection.
- Uploading confidential, personal, or sensitive data into public AI tools.
- Replacing human interviews or stakeholder engagement in sensitive contexts.
- Generating final evaluation questions without expert review.
- Selecting samples without methodological justification.
- Making judgements, ratings, conclusions, or recommendations automatically.
- Using AI to create evidence that cannot be audited or traced.
Required action: Avoid unless there is a specific approved, documented, legally compliant, and strongly supervised process.
Practical example: how this works in an evaluation assignment
The following example is illustrative. It is not a real case. It shows how a commissioner and contractor could manage AI use in practice.
Scenario
A donor commissions a mid-term evaluation of a youth employment programme. The evaluation includes document review, interviews with programme staff, focus group discussions with young participants, analysis of monitoring data, and a final report with recommendations.
Contractor’s proposed AI use
The contractor proposes to use AI for:
- Summarising public programme documents.
- Improving readability of draft report sections.
- Generating draft Excel formulas to clean monitoring data.
- Supporting preliminary coding of anonymised interview notes.
- Drafting possible wording for findings after human analysis.
Commissioner’s review
The commissioner reviews the proposed AI use during inception and classifies the tasks:
- Approved: summarising public documents, improving readability, and drafting formulas.
- Controlled use: preliminary coding of anonymised interview notes, only if the method is documented and coding is reviewed by a human evaluator.
- Restricted: uploading raw interview transcripts, names, personal information, focus group recordings, or sensitive participant data into a public AI tool.
- Not allowed: using AI to produce final conclusions or recommendations without human judgement and evidence review.
Agreed safeguards
- The contractor keeps an AI-use log.
- Only public documents and approved anonymised notes can be used with AI.
- All AI-supported summaries are checked against original documents.
- All citations are manually verified.
- The coding framework is reviewed by the evaluation team.
- Findings are discussed in a validation workshop before final reporting.
- The final report includes an AI-use disclosure statement.
What this changes in practice
AI is not banned. It is managed. The contractor can use AI for useful support tasks, but the commissioner retains visibility over quality, data protection, methods, and final judgement.
What commissioners should ask contractors
Proposal stage
- Will AI tools be used in this assignment?
- For which tasks will AI be used?
- Which tools or systems are expected to be used?
- Will any personal, confidential, or sensitive data be processed by AI?
- How will AI outputs be checked and validated?
- Who is responsible for final analytical quality?
Inception phase
- Which AI uses are approved, controlled, or restricted?
- How will AI use be documented?
- How will the commissioner review AI-supported outputs?
- What is the fallback plan if AI-supported work is unreliable?
- How will confidentiality and data protection be ensured?
- How will stakeholders be informed if AI use affects the evaluation process?
Deliverable review
- Which sections or products were AI-assisted?
- Were AI-generated outputs checked against original sources?
- Are all claims supported by traceable evidence?
- Were citations verified manually?
- Were any errors, limitations, or uncertainties identified?
- Is the final judgement clearly made by the evaluation team, not by AI?
AI-use agreement template
The commissioner and contractor can use the following template during inception.
| Item | Agreement |
|---|---|
| Approved tools | List AI tools or platforms approved for the assignment. |
| Approved tasks | List low-risk tasks where AI use is allowed. |
| Controlled tasks | List tasks requiring human review, documentation, and commissioner approval. |
| Restricted tasks | List tasks where AI use is not allowed or requires specific written approval. |
| Data protection | Define which data must never be entered into public AI tools. |
| Human review | Name who reviews AI-supported outputs and how review is documented. |
| Disclosure | Define how AI use will be disclosed in deliverables or final reports. |
| Error reporting | Explain how AI-related errors, concerns, or limitations will be reported. |
AI-use log template
An AI-use log is a simple record showing when, why, and how AI was used. Its purpose is to make AI use transparent and reviewable.
| Log item | What to record |
|---|---|
| Date | When AI was used. |
| Task | What AI was used for, such as editing, summarising, coding support, translation, or synthesis. |
| Tool | The AI tool, model, software, or platform used. |
| Input type | Whether the input was public, internal, confidential, personal, anonymised, or sensitive. |
| Output use | How the output was used, such as internal draft, analysis support, report text, or workshop material. |
| Human review | Who reviewed the output and what checks were performed. |
| Limitations | Any known errors, uncertainties, assumptions, or concerns. |
| Disclosure | Whether and how AI use was disclosed to the commissioner, stakeholders, or readers. |
Example AI-use disclosure statement
Disclosure should be clear, proportionate, and honest. It should match the real use of AI in the evaluation.
“AI tools were used to support selected non-final tasks during this evaluation, including document organisation, readability improvements, and preliminary synthesis support. All AI-assisted outputs were reviewed by the evaluation team. Final findings, conclusions, and recommendations were developed and approved by human evaluators.”
Adapt this statement to match the real use of AI. If AI was used for analysis, coding, translation, evidence synthesis, or report drafting, disclose this more specifically.
Quality assurance checkpoints
Commissioners should avoid relying only on final report review. AI-related quality assurance should happen throughout the evaluation.
- Before contracting: define AI-use expectations in the Terms of Reference.
- During proposal review: ask bidders to disclose planned AI use.
- During inception: classify AI uses and agree on safeguards.
- During data collection: protect personal, confidential, and sensitive data.
- During analysis: check AI-supported outputs against original evidence.
- During reporting: verify citations, claims, and source traceability.
- Before approval: confirm that findings, conclusions, and recommendations are human-owned.
Common mistakes to avoid
- Assuming AI use is harmless because it is “only drafting”.
- Allowing contractors to use AI without disclosure.
- Uploading interview transcripts or beneficiary data to public AI tools without approval.
- Accepting AI-generated citations without checking them.
- Using AI to make judgements that should be discussed by humans.
- Ignoring cultural and language limitations in AI-generated summaries.
- Reviewing AI-related risks only at the final report stage.
Download the guide
Use the PDF version as a practical reference when preparing Terms of Reference, inception reports, AI-use agreements, or evaluation quality assurance plans.
Main takeaway
Responsible AI use in commissioned evaluations is not only a technical question. It is a question of evaluation quality, trust, confidentiality, accountability, and human judgement.
FAQ
Should AI use be included in the Terms of Reference?
Yes. If AI may be used, the Terms of Reference should clarify expectations around disclosure, data protection, quality assurance, and human responsibility.
Can contractors use AI for editing and readability?
This can be acceptable when the material is not sensitive and the output is reviewed. It should still be declared according to the commissioner’s requirements.
Can AI be used to analyse qualitative data?
Only with caution. Commissioners and contractors should agree on data protection rules, method, validation process, human review, and limitations before using AI for qualitative analysis.
Can AI write evaluation conclusions?
AI may support drafting, but final conclusions and recommendations should remain the responsibility of human evaluators who have reviewed the evidence and context.
Why is AI disclosure important?
Disclosure helps protect trust. It allows commissioners, stakeholders, and readers to understand where AI was used and how human review was applied.
What is the safest rule for confidential data?
Do not enter personal, confidential, sensitive, or vulnerable-group data into public AI tools unless there is a clear legal basis, approved safeguards, and a secure environment.
Related EvalCommunity resources
- Download this guide as a PDF
- AI in Monitoring & Evaluation Certificate
- UNESCO Ethical Impact Assessment: Practical Framework for Responsible AI
- AI for International Development and Humanitarian Practitioners Certificate
External sources and further reading
- Evaluation Helpdesk of the Cohesion Policy. Commissioning evaluation in the AI era: working with contractors and colleagues. Guidance note by T. Delahais.
- UNESCO Ethical Impact Assessment
- European Commission Assessment List for Trustworthy Artificial Intelligence
- CGIAR IAES Technical Note on Using AI in Evaluations
Source and attribution
This EvalCommunity tutorial is informed by the guidance note Commissioning evaluation in the AI era: working with contractors and colleagues, authored by T. Delahais for the Evaluation Helpdesk of the Cohesion Policy, and by broader responsible AI guidance relevant to evaluation practice.
It adapts the topic into an original practical workflow for evaluation commissioners, M&E managers, procurement teams, and evaluation consultants.
Note: This tutorial is for educational and professional learning purposes. It does not replace legal, procurement, data protection, or institutional AI governance advice.
