How to Define Artificial Intelligence in Monitoring, Evaluation, and International Development
EvalCommunity Academy Tutorial
How to Define Artificial Intelligence Accurately in Monitoring, Evaluation, and International Development
A practical guide for distinguishing artificial intelligence from digital tools, automation, analytics, machine learning, generative AI, and AI agents.
Artificial intelligence is now frequently mentioned in international development strategies, funding calls, programme proposals, research papers, evaluation reports, and technology platforms.
However, many interventions described as AI-powered are primarily based on conventional digital tools, fixed workflows, automated reminders, dashboards, or pre-written response systems.
This creates an important challenge for monitoring, evaluation, accountability, and learning professionals. Before evaluating whether an AI intervention works, evaluators must first determine what technology is being used, which part of the system contains an AI component, and how that component is expected to influence programme results.
Central evaluation question:
What specific system was introduced, what does it do, and what evidence shows that it contributed to the reported outcome?
Learning Objectives
By the end of this tutorial, you should be able to:
- Distinguish artificial intelligence from conventional digital systems and automation.
- Identify the specific AI component within a larger digital platform.
- Improve the accuracy of AI-related indicators and theories of change.
- Assess whether results are linked to technology or surrounding human support.
- Recognise exaggerated or unclear AI claims.
- Apply a practical technology-classification framework.
- Develop more credible AI evaluation questions.
1. Why Accurate Definitions Matter
Technology labels influence how programmes are designed, funded, implemented, monitored, evaluated, and communicated.
When a conventional digital intervention is described as an AI programme, teams may develop indicators that do not measure the intended concept. Donors may believe they are financing advanced modelling when the programme mainly uses automation. Evaluators may attribute results to AI when training, mentoring, or organisational support provides a stronger explanation.
Accurate definitions support:
Measuring the intervention that was actually delivered.
Explaining clearly what the system does and does not do.
Matching technology choices to programme needs and risks.
Comparing AI systems with simpler alternatives.
Applying safeguards that match the actual risks.
Avoiding comparisons between fundamentally different systems.
A programme can be useful, innovative, and effective without using artificial intelligence. Accurate language strengthens rather than weakens its credibility.
2. A Practical Technology Classification
Evaluators do not need to be software engineers, but they should be able to classify a system’s main function and identify where AI is present.
Category 1: Conventional Digital Systems
These systems store, display, organise, transmit, or retrieve information through predefined functions.
Examples: Digital forms, document repositories, electronic ledgers, basic dashboards, cloud storage, and mobile data-collection tools.
Category 2: Rule-Based Automation
These systems perform an action when a predefined condition is met.
Examples: Appointment reminders, automated emails, threshold alerts, field-validation warnings, and complaint-routing workflows.
The system follows rules established in advance and does not necessarily learn from new information.
Category 3: Statistical and Analytical Systems
These systems analyse data using formulas, statistical procedures, or structured analytical methods.
Examples: Trend calculations, regressions, forecasting formulas, indicator dashboards, and descriptive statistics.
Category 4: Machine-Learning Systems
These systems use patterns found in historical or labelled data to make classifications, estimates, recommendations, or predictions.
Examples: Dropout prediction, fraud detection, image classification, anomaly detection, and outcome-probability estimates.
Outputs are generally probabilistic and may include false positives, false negatives, or performance differences across groups.
Category 5: Generative AI Systems
These systems produce new text, images, audio, video, code, or other content in response to instructions and contextual information.
Examples: Drafting evaluation reports, summarising interviews, generating survey questions, translating materials, producing code, and answering questions.
Generative systems can produce fluent outputs that are inaccurate, unsupported, incomplete, or inappropriate.
Category 6: AI Agents
AI agents perform several connected actions in pursuit of an objective. They may retrieve information, use external tools, analyse evidence, make intermediate choices, and prepare outputs for review.
An evaluation agent might search programme documents, extract evidence, compare findings with an evaluation matrix, identify gaps, draft a summary, and submit the output for human review.
Quick Classification Reference
Stores or displays information: Conventional digital system
Follows a predefined condition: Rule-based automation
Uses formulas or statistical procedures: Analytics
Learns patterns to classify or predict: Machine learning
Creates new content: Generative AI
Plans and completes connected actions: AI agent
3. Focus on the Function, Not the Product Name
A single software platform may contain a conventional database, rule-based notifications, an analytical dashboard, a recommendation model, and a generative writing assistant.
It would therefore be inaccurate to classify every activity performed through that platform as an AI activity.
Instead of asking:
Is this an AI platform?
Ask:
Which feature is being used, what does it do, and what technical process produces its output?
4. The Four-Part AI Description
Every AI-related programme should be able to describe its system through four elements.
1. The Task
What is the system expected to classify, generate, predict, recommend, translate, detect, or prioritise?
2. The Mechanism
Does it use fixed rules, formulas, machine learning, a language model, computer vision, or several components?
3. The Data
What monitoring data, documents, records, images, transcripts, databases, or user instructions does it process?
4. The Decision Role
Does the output inform a person, recommend an action, trigger a workflow, prioritise a case, or make a decision?
Risk principle: The more influence an AI system has over important decisions, the stronger the requirements should be for validation, documentation, transparency, appeal, and human oversight.
5. Common AI Classification Errors
- Treating all automation as AI. Scheduled messages and workflow triggers are often based on fixed rules.
- Classifying the entire platform as AI. A platform may contain only one AI feature among many conventional functions.
- Relying on vendor terminology. Terms such as “smart” and “AI-powered” require a technical explanation.
- Assuming every recommendation uses AI. Recommendations may result from fixed rules or simple ranking formulas.
- Treating AI access as AI use. Participants may have access to a platform without using its AI feature.
- Attributing all outcomes to technology. Training, facilitation, trust, connectivity, and support may explain the result.
6. Why This Matters for Indicators
Poorly defined technology produces poorly defined indicators.
Weak Indicator
Percentage of participants using AI tools.
This does not clarify the relevant feature, level of use, reporting period, frequency, or relationship to the programme objective.
Improved Indicator
Percentage of trained participants who used the programme’s generative advisory feature at least twice during the reporting period and applied at least one output to a documented decision.
Additional indicators may measure:
- Human review of AI-generated outputs
- Outputs requiring correction
- User understanding of system limitations
- Recommendations accepted, modified, or rejected
- Time saved compared with the previous process
- Errors affecting programme decisions
- Users able to challenge or appeal an output
7. Improve the Theory of Change
Weak pathway: AI tool introduced → users adopt it → outcomes improve
A stronger pathway recognises the technical, human, organisational, and contextual conditions required for change.
- The programme selects an appropriate system.
- Relevant and sufficiently reliable data are available.
- Staff understand the system’s purpose and limitations.
- Participants receive accessible training and support.
- Users trust the process enough to engage with it.
- Outputs are reviewed and corrected where necessary.
- Users apply appropriate outputs to their work.
- Organisational processes support the new practice.
- The intervention improves decisions or services.
- Improved decisions contribute to programme outcomes.
Important Assumptions
- Users have adequate digital access.
- Training is understandable and relevant.
- The system performs adequately across groups.
- The data reflects the intended population.
- Users do not over-trust the outputs.
- Human reviewers have sufficient time and expertise.
- The organisation can act on the information produced.
Potential Negative Pathways
- Digital exclusion
- Biased or inaccurate recommendations
- Privacy or confidentiality breaches
- Automation bias
- Reduced professional judgement
- Increased verification workload
- Vendor dependency
- Weak accountability for decisions
8. Separate the Technology Effect from the Support Effect
Digital and AI programmes often combine software, devices, connectivity, training, mentoring, technical support, peer learning, and organisational supervision.
Evaluators should not assume that technology explains the outcome simply because it is the most visible component.
Useful approaches include:
- Contribution analysis
- Process tracing
- Comparison groups
- Dose-response analysis
- Qualitative interviews
- Usage analytics
- Implementation-fidelity assessment
- Realist evaluation
- Outcome harvesting
Questions to Investigate
- Which component did participants find most useful?
- Would the outcome have occurred without the training?
- Did participants receiving more support achieve better results?
- Were the AI features used regularly?
- Did participants apply the outputs independently?
- What alternative explanations are supported by the evidence?
9. Recognising AI Washing
AI washing occurs when an organisation exaggerates, misrepresents, or communicates unclearly about the role of artificial intelligence.
Warning signs may include:
- Repeated use of “AI-powered” without identifying the model
- Claims that a system learns when only fixed rules are documented
- No clear description of the data used
- No explanation of how outputs are produced
- No connection between the AI feature and a programme activity
- Outcomes based on platform access rather than actual feature use
- Strong impact claims without causal analysis
- Little information about validation or error rates
Not every inaccurate AI claim is intentionally misleading. The problem may result from limited technical understanding, vendor terminology, or imprecise communication. The claim should still be clarified before it is used in evaluation or reporting.
10. Questions for Programme Designers
- What problem are we trying to solve?
- Why is AI needed?
- Could a simpler digital system achieve the same result?
- What specific task will the AI perform?
- What data will it require?
- Is the data accurate and representative?
- Who will review the outputs?
- What happens when the system is wrong?
- How will users challenge or correct decisions?
- How will performance be monitored?
- What risks may affect vulnerable groups?
- What resources are needed to maintain the system?
11. Questions for Donors
Funding applications involving AI should specify:
- The exact AI function and model category
- The data requirements and intended users
- The decision being supported
- The expected benefit
- The non-AI alternatives considered
- The validation plan
- The human-oversight process
- The risk-management framework
- The monitoring indicators
- The sustainability or exit strategy
12. Questions for Technology Providers
Technical
- Which feature uses AI?
- What type of model is used?
- How was it developed or configured?
- What data was used?
- Does the model change over time?
- What are its known limitations?
Performance
- How was performance tested?
- Which metrics were used?
- Were results disaggregated?
- How often does the system make errors?
- What error types are most common?
- Was it tested in the local context?
Governance
- Where is user data stored?
- Is programme data used for training?
- Can the AI feature be disabled?
- Who is responsible for harm?
- Can users appeal decisions?
- What audit documentation is available?
13. AI Evaluation Matrix
Use the following dimensions during the evaluation inception phase.
What part of the intervention uses AI? Review system documentation, demonstrations, workflow maps, and technical specifications.
Is AI appropriate for the development problem? Review needs assessments, stakeholder views, and simpler alternatives.
Does the AI component improve programme performance? Use outcome data, comparison data, analytics, and interviews.
How often does the system produce acceptable outputs? Review validation samples, error logs, and expert assessments.
Does the system perform differently across groups? Use disaggregated results and subgroup error analysis.
Are outputs reviewed before they influence decisions? Examine operating procedures and approval workflows.
Can affected people understand and challenge decisions? Review appeals, complaints, and user guidance.
Can the organisation maintain the system responsibly? Review budgets, skills, contracts, staffing, and infrastructure.
14. Practical Exercise: Classify the Technology
Classification: Rule-based automation.
Classification: Machine learning or predictive analytics, depending on the method.
Classification: Rule-based system.
Classification: Generative AI.
Classification: AI agent or agentic workflow.
Reflection Questions
- What evidence would you request before confirming each classification?
- Which examples create the greatest programme risk?
- Which outputs require mandatory human review?
- Which examples could work effectively without AI?
15. Practical Exercise: Rewrite a Weak Indicator
Number of beneficiaries empowered through AI.
This does not define empowerment, the AI function, expected behaviour, measurement method, timeframe, or causal pathway.
Percentage of participating caseworkers who used the AI-supported document-review function during the reporting quarter and reported reduced time identifying missing information.
Percentage of information gaps identified by the AI system that were confirmed as valid by a qualified case supervisor.
16. Practical Exercise: Test the Causal Claim
Imagine a programme reports that an AI advisory platform improved the confidence of local entrepreneurs.
Before accepting the claim, investigate:
- Whether confidence was measured before and after the programme
- How confidence was defined
- Which platform features were used
- How frequently the AI feature was used
- How much training and mentoring participants received
- Whether more support was associated with stronger outcomes
- Whether participants applied the recommendations
- Whether similar improvements occurred among non-users
Participants reported increased confidence after receiving structured digital training and advisory support. The contribution of the AI feature could not be separated from the wider support package.
17. Minimum Documentation for an AI Intervention
Name, provider, version, model category, intended function, users, and operating context.
Sources, ownership, consent, cleaning procedures, gaps, representation risks, and retention.
Validation method, metrics, error categories, subgroup analysis, and known limitations.
Human oversight, approvals, escalation, complaints, incidents, and accountability.
Model updates, prompt changes, workflow modifications, new data sources, and emerging risks.
18. Pre-Evaluation Checklist
Intervention Clarity
☐ The specific AI function is identified.
☐ The technology category is documented.
☐ The description does not rely only on vendor terminology.
☐ AI and non-AI components are separated.
Theory of Change
☐ The causal role of the AI component is explicit.
☐ Human support and organisational conditions are included.
☐ Alternative explanations are identified.
☐ Potential negative pathways are documented.
Indicators
☐ AI use is operationally defined.
☐ Output quality is measured.
☐ Error rates are monitored.
☐ Results are disaggregated.
☐ Human review is tracked.
Ethics and Accountability
☐ Data-protection arrangements are documented.
☐ Affected people can raise concerns.
☐ Staff can override the system.
☐ High-risk outputs receive human review.
☐ Vulnerable groups are considered.
Evaluation Design
☐ Access is distinguished from actual use.
☐ AI contribution is separated from training where possible.
☐ Qualitative and quantitative evidence are included.
☐ Technical documentation is available.
☐ System and workflow changes will be recorded.
19. Key Principles for M&E Professionals
- Describe before you evaluate. Begin with the task, mechanism, data, and decision process.
- Evaluate the feature actually used. Focus on the function participants use in practice.
- Measure quality, not only adoption. High use does not prove that outputs are accurate, useful, fair, or safe.
- Keep human factors visible. Training, trust, facilitation, accessibility, and support may explain the outcome.
- Match the claim to the evidence. Do not claim causation when only satisfaction or access was measured.
- Prefer the simplest suitable solution. A conventional tool may be more reliable, transparent, affordable, and sustainable.
- Treat terminology as an accountability issue. Accurate language helps stakeholders understand capabilities and risks.
Conclusion
Monitoring and evaluation professionals have an important role in improving the quality of evidence surrounding artificial intelligence.
Digital platforms, automation, statistical analysis, machine learning, generative AI, and AI agents should not be treated as interchangeable. Each has different capabilities, limitations, risks, and evaluation requirements.
The Better Evaluation Question
What specific system was introduced, how was it expected to contribute, what evidence shows that it influenced the outcome, and what other factors may explain the result?
Final Assignment
Assess a Digital or AI Intervention
Select one digital or AI-related intervention from your organisation or professional context. Prepare a two-page assessment covering:
- The development problem being addressed
- The specific technology function
- The most appropriate classification
- The data used by the system
- The system’s role in decision-making
- The human-support components
- Three potential benefits
- Three potential risks
- Five evaluation questions
- Three indicators covering use, quality, and accountability
Conclude by explaining whether artificial intelligence is necessary or whether a simpler solution could achieve the intended result.
Continue Learning
AI in Monitoring & Evaluation Certificate Course
Build practical skills in AI-supported evaluation planning, data analysis, qualitative research, evidence synthesis, reporting, responsible AI, validation, and human oversight.
