
Designing a MEL Framework with Generative AI: GPT
EvalCommunity Academy · Student Case Study
Designing a MEL Framework with Generative AI: GPT
Using generative AI to support Monitoring, Evaluation, and Learning Framework design while retaining human evaluative judgment.
Case Overview
A small think tank focused on international development policy in Australia and the Indo-Pacific produces analysis, policy briefs, and strategic dialogues intended to inform decision-making across government and development actors.
The organization wanted to strengthen its approach to understanding whether its outputs contributed to changes in policy decisions, policy debates, or learning. Developing a Monitoring, Evaluation, and Learning Framework (MELF) was selected as the first step.
The case examines how GPT-4 was used as part of this design process. The purpose was not to delegate evaluative reasoning to AI, but to explore how generative AI could support the evaluator while identifying where human judgment remained essential.
Core question
How can evaluators use generative AI to accelerate MELF design without allowing AI-generated outputs to replace evaluative reasoning, contextual knowledge, stakeholder engagement, and professional judgment?
1. Study Context
The case was undertaken as a formative design exercise by a solo evaluator. Approximately two days of work were invested over a period of six weeks.
The organization wanted the MELF to provide an overarching evaluative structure, establish standards, and develop a theory of change that could later underpin implementation.
The exercise was designed to test how ChatGPT-4 could support MELF development, explore methodological adaptations, and examine the evaluator competencies required for responsible integration of AI into evaluation design.
2. Why Use Generative AI?
Designing a MELF requires bringing together theoretical literature, organizational priorities, evaluative standards, and information about the organization’s work.
Generative AI offered potential advantages in several areas:
- Accelerating synthesis of diverse evidence.
- Generating initial framework prototypes.
- Processing large amounts of organizational text.
- Identifying patterns and recurring themes.
- Supporting alternative formulations of evaluation concepts.
- Providing material that could subsequently be challenged and refined by the evaluator.
The purpose was explicitly not to delegate evaluative reasoning. Instead, the exercise examined how AI might extend, and sometimes constrain, evaluator practice.
The case reports that the AI-assisted exercise took approximately two days, whereas a traditional MELF design process in this context was estimated at between five and ten days.
3. The Initial AI Prototype
The evaluator began by asking GPT-4 to create a baseline MELF with minimal prompting.
The resulting framework provided a generic structure but revealed weaknesses in evaluation criteria and contextual sensitivity.
- Vague evaluation criteria.
- Insufficient sensitivity to organizational context.
- Generic rather than organization-specific logic.
The initial draft was therefore treated as diagnostic rather than final. Its weaknesses became the basis for subsequent refinement.
Important lesson
A coherent and professionally written AI-generated framework is not necessarily a methodologically strong framework.
4. The Five-Stage Design Process
Stage 1 — Initial Prototype Generation
GPT-4 generated a baseline MELF with minimal prompting. The draft exposed weaknesses in criteria and contextual sensitivity.
Stage 2 — Content Modeling of Past Outputs
More than 100 outputs from the think tank were analyzed, including website content, policy briefs, and analytical reports.
GenAI was used to identify dominant themes and patterns of engagement and to cluster them into preliminary evaluative domains.
Stage 3 — Literature Synthesis
GenAI supported a structured review of evaluation theory and think tank performance literature. Draft criteria and standards were developed drawing on Scriven’s logic of evaluation and realist principles.
Stage 4 — Realist-Informed Theory of Change
Prompts were used to structure potential Context–Mechanism–Outcome (CMO) pathways.
However, the initial AI-generated pathways tended toward linear and mechanistic models. One important assumption was that producing high-quality, evidence-based analysis would itself trigger policy uptake.
The evaluator subsequently challenged these assumptions using organizational feedback and documentation concerning policy impact, reflecting the contingent nature of policy contribution.
Stage 5 — Staff Feedback and Redrafting
Feedback was collected from nine staff members across analytical, engagement, and leadership functions through short written reflections and optional voice notes.
GPT-4 was used to identify common concerns and suggestions. Staff highlighted:
- The importance of shaping policy discourse alongside formal policy uptake.
- The importance of relationships and networks as conditions for impact.
- The need to account for informal and relational forms of influence.
Staff judgment was then used to moderate and reinterpret AI-generated themes where they risked over-generalization or insufficient contextualization.
5. A Key Problem: Frequency Is Not Significance
One important weakness was GenAI’s tendency to associate frequency with significance.
For example, “regional partnerships” appeared repeatedly and was therefore treated as a central theme. Less frequently mentioned issues, including “budgetary trade-offs” and “policy coordination failures,” received less emphasis even though they could be strategically significant.
The evaluator therefore had to reassert interpretive authority and determine whether frequently repeated themes represented meaningful contributions or simply reflected organizational habits.
Student reflection: Why might an AI system identify frequently mentioned topics more readily than strategically important but less frequently mentioned issues?
6. The Theory of Change Challenge
The initial AI-generated logic tended toward a simplified causal relationship:
The case identified this as an overly linear assumption.
Policy contribution may depend on context, relationships, networks, engagement, timing, and other conditions. The evaluator therefore used a realist-informed approach to interrogate potential Context–Mechanism–Outcome pathways.
Student Task
Context: What conditions might affect whether evidence contributes to change?
Mechanism: What might cause decision-makers, institutions, or policy communities to respond?
Outcome: What changes could occur beyond formal policy adoption?
Consider changes in policy discourse, awareness, relationships, policy options considered, evidence use, organizational learning, or eventual policy change.
7. What Human Participants Added
The staff feedback helped identify dimensions of impact that could be under-emphasized by a more mechanical analysis.
| Potential AI Limitation | Human Contribution |
|---|---|
| Emphasis on formal policy uptake | Recognition of influence on policy discourse |
| Over-generalization | Context-specific interpretation |
| Under-emphasis on relationships | Recognition of networks and relational influence |
8. The Resulting MELF
The resulting MELF was approximately 3,500 words and contained six sections:
- Context and scope
- Theory of change
- Evaluation questions and criteria
- Indicators and data sources
- Methods and analysis
- Learning and use
The framework was coherent and included explicit criteria, evaluative domains, and a basic realist-informed theory of change.
However, the case concluded that the framework was not particularly fit for purpose. It lacked the contextual nuance and contextual fit required for a high-quality operational product.
Overall finding
The outcome of using GenAI was adequate but not fit for purpose.
9. What GenAI Did Well
- Accelerated synthesis.
- Generated initial drafts and prototypes.
- Processed large volumes of text.
- Surfaced patterns and themes.
- Supported alternative design options.
- Helped process staff feedback.
- Supported iterative interrogation of the framework.
10. What GenAI Did Poorly
- Produced vague outputs.
- Generated overly generic frameworks.
- Tended toward linear causal assumptions.
- Could give disproportionate weight to frequently occurring themes.
- Could over-generalize organizational patterns.
- Could under-emphasize informal and relational forms of impact.
- Did not independently produce a framework sufficiently fit for operational use.
11. Evaluator Competencies
The case identifies evaluator competencies as a key safeguard when GenAI is incorporated into MELF development.
Evaluative Reasoning
The ability to distinguish surface-level adequacy from substantive quality and to interrogate assumptions in AI-generated outputs.
Contextual Literacy
The ability to understand and incorporate the realities of the organization, policy environment, relationships, and networks.
Responsible AI Literacy
The ability to understand AI limitations, verify outputs, document decisions, and maintain transparency and accountability.
The case also highlights critical experimentation, including testing prompting approaches, approaches for eliciting human input, and the effectiveness of different GenAI tools in quality-assurance functions.
12. Student Exercise — Critique the AI
Imagine that you are the evaluator reviewing the initial AI-generated MELF.
Identify five weaknesses you would look for.
- Are the evaluation criteria sufficiently specific?
- Does the framework reflect the organization’s context?
- Are the causal assumptions realistic?
- Are important issues being identified, or simply frequently mentioned issues?
- Are indicators and evidence requirements sufficiently clear for operational use?
13. Student Exercise — Human or AI?
For each activity, decide whether it should be primarily AI-supported, evaluator-led, or shared between AI and the evaluator.
| Activity | Your Decision |
|---|---|
| Summarizing large volumes of documents | AI / Human / Shared |
| Identifying recurring themes | AI / Human / Shared |
| Developing evaluation questions | AI / Human / Shared |
| Judging significance | AI / Human / Shared |
| Assessing causal assumptions | AI / Human / Shared |
| Understanding organizational context | AI / Human / Shared |
| Making final methodological decisions | AI / Human / Shared |
Discussion: What should never be delegated entirely to generative AI in evaluation design?
14. Final Student Challenge
The think tank asks:
“Can we use GPT to develop our MEL Framework?”
Prepare a one-page recommendation to the organization addressing:
- What AI should be used for.
- What AI should not be trusted to do.
- How human review should be incorporated.
- How stakeholder input should be incorporated.
- How AI-generated assumptions should be tested.
- How the final framework should be validated.
- What safeguards should be established.
End your recommendation with one of these decisions:
Justify your decision using evidence and reasoning from the case.
15. Responsible AI and Future Evaluation Practice
The case considers the possibility that increasingly capable AI systems will become more agentic, with capabilities involving planning, memory, multimodal reasoning, and real-world action execution.
Such systems could automate substantial components of data gathering, coding, and synthesis. At the same time, they introduce concerns around verification, attribution, transparency, accountability, and opaque decision-making.
The case therefore emphasizes the importance of contestability: stakeholders should be able to question and challenge evaluation results rather than simply accept automated outputs.
As AI capabilities increase, evaluators may need stronger skills in interrogating AI outputs, identifying problematic causal reasoning, monitoring audit trails, and validating AI-supported processes.
16. Key Takeaway
AI generates. The evaluator interrogates. Evidence constrains. Context decides.
The evaluator retains ownership of methodological decisions and treats AI as a critical interlocutor rather than an authoritative designer.
The central lesson of the case is that the value of GenAI in evaluation design lies less in producing a ready-to-use framework and more in accelerating the process of surfacing weaknesses, testing alternative design logics, and interrogating assumptions.
The approach is particularly useful for early-stage tasks such as mapping evaluative domains, stress-testing theories of change, and identifying implicit assumptions. It is less suitable when evaluation purposes are poorly specified, institutional context is weakly articulated, or AI outputs are treated as authoritative rather than provisional.
Recommended Course
Continue Your Learning
Custom GPTs for Evaluators
Learn how to build and use Custom GPTs for evaluation work, structure AI-assisted workflows, develop reliable instructions, work with evaluation knowledge, test outputs, and apply responsible AI practices.
Key principle: AI drafts — humans decide. Final methodological judgment should remain with the evaluator.
Source & Further Reading
Primary Case Source
Fung, Geordie. “Designing with generative AI: Using GPT-4 to build a Monitoring, Evaluation, and Learning Framework.” Chapter 5 in From Algorithms to Evidence: Using Generative AI in Evaluation Practice.
Pertinent Sources from the Chapter
The following references are selected from the bibliography provided at the end of Fung’s chapter. They are particularly relevant to the themes explored in this case: generative AI in evaluation, evaluation design, AI ethics, realist evaluation, and the changing role of evaluators.
Head, C. B., Jasper, P., McConnachie, M., Raftree, L., & Higdon, G. (2023).
“Large language model applications for evaluation: Opportunities and ethical implications.”
New Directions for Evaluation, 2023(178–179), 33–46.
DOI: 10.1002/ev.20556
Ferretti, F. (2023).
“Hacking by the prompt: Innovative ways to utilize ChatGPT for evaluators.”
New Directions for Evaluation, 2023(178–179), 47–61.
DOI: 10.1002/ev.20557
Christou, P. A. (2023).
“The use of artificial intelligence (AI) in qualitative research for theory development.”
The Qualitative Report, 28(9), 2739–2755.
DOI: 10.46743/2160-3715/2023.6536
Jobin, A., Ienca, M., & Vayena, E. (2019).
“The global landscape of AI ethics guidelines.”
Nature Machine Intelligence, 1(9), 389–399.
DOI: 10.1038/s42256-019-0088-2
Nielsen, S. B., Rinaldi, F. M., & Petersson, G. J. (2025).
Artificial Intelligence and Evaluation: Emerging Technologies and Their Implications.
Routledge.
Rompczyk, K. (2025).
“Technological revolution in evaluation: Artificial intelligence and the adherence to evaluation standards.”
Evaluation, 31(3), 331–351.
DOI: 10.1177/13563890251331066
Reid, A. M. (2023).
“Vision for an equitable AI world: The role of evaluation and evaluators to incite change.”
New Directions for Evaluation, 2023(178–179), 111–121.
DOI: 10.1002/ev.20559
Pawson, R., & Tilley, N. (1997).
Realistic Evaluation.
Sage Publications.
Scriven, M. (1991).
Evaluation Thesaurus (4th ed.).
Sage Publications.
Patton, M. Q. (2010).
Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use.
Guilford Press.
Why these sources matter for this case
Generative AI and evaluation: Head et al., Ferretti, Christou, Nielsen et al., and Rompczyk provide relevant foundations for understanding opportunities, limitations, and implications of AI use in evaluation practice.
Evaluation theory and causal reasoning: Pawson and Tilley’s work on realist evaluation is relevant to the case’s use of Context–Mechanism–Outcome pathways and its critique of overly linear causal assumptions.
Evaluation judgment: Scriven’s work provides one of the evaluative foundations used in assessing criteria and standards.
Responsible and ethical AI: Jobin et al., Reid, and Rompczyk provide relevant perspectives for considering ethics, accountability, transparency, and the changing responsibilities of evaluators working with AI.
Source note: The further-reading references above are selected from the bibliography included in Fung’s chapter. They are presented for students as additional reading and are not additional sources used independently to construct this case study.
The full bibliography, including additional references, appears in the original source chapter.
Disclaimer
This case study is provided by EvalCommunity Academy for educational and learning purposes. It is an adapted teaching resource based primarily on the case presented by Geordie Fung, “Designing with generative AI: Using GPT-4 to build a Monitoring, Evaluation, and Learning Framework,” as identified in the source material provided with this lesson.
The case study has been reorganized and formatted for students to support learning, reflection, discussion, and practical application in Monitoring, Evaluation, and Learning (MEL). It should not be interpreted as a reproduction of the original chapter or as an independent empirical study conducted by EvalCommunity.
The student exercises, questions, explanatory transitions, and learning prompts have been developed for educational use. They are intended to encourage critical analysis and should not be interpreted as findings reported by the original author unless explicitly identified as such.
Students should consult the original source chapter for the complete case, methodology, discussion, limitations, declarations, and bibliography. Interpretations and recommendations developed while completing the exercises are the responsibility of the student and should be supported by appropriate evidence.
The inclusion of generative AI in this case study does not constitute an endorsement of using AI without appropriate safeguards. AI-generated content can be inaccurate, incomplete, overly generalized, or insufficiently sensitive to context. Evaluation professionals remain responsible for methodological decisions, verification, interpretation, confidentiality, ethical practice, and the quality of their work.
Important: Do not enter confidential, proprietary, personal, or otherwise sensitive information into a generative AI system unless its use has been appropriately authorized and the relevant privacy, security, ethical, and organizational requirements have been addressed.
