
Generative AI for Evaluation Planning
EvalCommunity Academy · Case Study
Generative AI for Evaluation Planning: Strengthening Participation and Ownership
How an AI-assisted planning workflow supported efficiency, participation, accessibility, and ownership in a grassroots nonprofit.
Case Overview
This case examines the use of generative AI during the evaluation planning phase of an engagement with My Very Own Bed (MVOB), a grassroots nonprofit in Minneapolis. The approach combined a custom AI assistant with lightweight, AI-generated multimodal tools to make planning more feasible for a resource-constrained organization while creating accessible entry points for participation.
The case is particularly useful for understanding an important distinction: AI was used to support the work of evaluation planning, not to replace evaluative judgment. The evaluator remained responsible for contextualization, methodological decisions, quality assurance, communication, ethics, and final decisions.
Primary evaluation phase: Evaluation planning
Cross-cutting purposes: Participation, accessibility, ownership, capacity strengthening
Main AI contribution: First-pass drafting, continuity across planning documents, formatting, and support for participatory tools
Human responsibility: Contextualization, methodological judgment, review, ethics, and final decisions
What You Will Learn
- How generative AI can reduce labor demands during evaluation planning.
- How a custom AI assistant can maintain continuity across multiple planning products.
- How AI-generated multimodal tools can create additional entry points for client participation.
- Why participation and ownership require more than simply producing technically correct documents.
- How evaluators can retain methodological control while using AI to support repetitive work.
- Why human review, traceability, consent, and data protection remain essential.
Study Context
Grassroots nonprofits often operate with intermittent funding, lean staffing, limited formal MERL infrastructure, and limited evaluation capacity. In such environments, evaluation can become focused on activity tracking and reporting rather than systematic inquiry connected to decisions and learning.
The organization in this case, My Very Own Bed (MVOB), is a Minneapolis nonprofit providing beds to children moving into stable housing. MVOB wanted to strengthen its internal MERL capacity and better understand and strengthen its contribution to family outcomes.
The evaluation approach drew on utilization-focused evaluation and empowerment evaluation. The emphasis was therefore not simply on producing an evaluation product, but on ensuring that the work was useful, participatory, understandable, and connected to organizational decision-making.
Visit My Very Own Bed
·
Learn about its work
·
Referral partners
Why Use Generative AI?
The case identifies several practical pressures that made AI-assisted planning attractive:
- Limited organizational capacity.
- Tight budgets.
- Compressed decision timelines.
- High labor requirements for drafting and formatting.
- The need to maintain consistency across multiple planning documents.
- The need to redirect evaluator time toward substantive participation and engagement.
The logic was straightforward: if AI could take on some repetitive first-pass work, the evaluator could spend more time on the parts of the process where professional judgment and human interaction were most important.
The Four Actors in the Workflow
| Actor | Role |
|---|---|
| Evaluator | Coordinated the planning process, managed quality and coherence, communicated with the client, shaped AI use, and retained methodological and ethical responsibility. |
| MVOB | Provided organizational priorities, operational constraints, data and capacity information, and staff input. |
| Custom AI assistant | Supported first-pass drafting, formatting, continuity, and repetitive planning tasks. |
| AI-generated multimodal tools | Provided interactive and accessible ways for staff and leadership to explore decisions, questions, methods, and causal pathways. |
What the AI Assistant Supported
The evaluator loaded relevant context documents and iteratively refined instructions. The assistant produced first-pass versions of planning materials that the evaluator then reviewed, revised, and finalized.
- Evaluation questions
- Evaluation matrix
- Evaluation plan
- Terms of Reference (TOR)
- Gantt chart
- Data collection instruments
- Technical and plain-language versions of planning content
The assistant did not directly interact with the client. The evaluator remained the intermediary and quality-control point.
Selecting the AI Environment
The evaluator tested several platforms and considered their usefulness for maintaining continuity across the planning process.
| Platform | Observation reported in the case |
|---|---|
| ChatGPT | Used as the main environment during the engagement; the case describes model changes during the period and adjustments to prompting and review. |
| Claude | Produced strong drafts but was considered too expensive for sustained use and had token limitations. |
| Gemini | The case reports a ten-document ingestion restriction that was impractical for the larger corpus. |
| ChatGPT Custom GPTs | Supported document upload but did not retain context consistently enough across time for the workflow described. |
| ChatGPT Projects | Selected because larger document ingestion and continuity/shared memory across threads better supported the planning workflow. |
AI-Generated Tools for Participation
The case goes beyond using AI to draft documents. It describes several lightweight tools designed to help non-MERL specialists engage with evaluation decisions.
1. Evaluability Assessment App
An interactive readiness tool addressing program maturity, leadership stability, data availability, and intended use. The case describes it as adapted from Wholey’s evaluability assessment approach.
2. TOR and Evaluation Plan Explainers
NotebookLM-based multimedia walkthroughs translated technical planning documents into audio-visual explanations so non-MERL staff could review and respond more easily.
3. Evaluation Questions Mini-App
A Google AI Studio application presented evaluation questions as interactive digital cards. Users could explore questions, subquestions, and rationales asynchronously.
4. Methods Options Tool
An interactive slide-based tool presented qualitative data-collection options together with strengths and limitations, helping leadership engage with methodological choices.
5. Impact Wizard
The team used the Impact Wizard to map causal pathways, identify assumptions, and develop a shared visual model of how the program was expected to contribute to outcomes.
What Changed?
The case reports several process-level results:
- Initial drafting was reduced from days to hours.
- Centralized drafting reduced contradictions and drift between planning documents.
- The workflow reduced cognitive and administrative burden.
- Interactive tools created additional ways for staff and leadership to participate.
- Multimodal explanations made technical planning content more accessible to non-MERL staff.
- The Impact Wizard supported visualization of causal pathways and assumptions.
- The overall process was made more feasible for a resource-constrained organization.
What the Case Does Not Demonstrate
- It is based on a single evaluation engagement.
- It does not establish causal effects of generative AI on evaluation quality.
- It does not demonstrate that the workflow performs equally well in other organizational contexts.
- It does not establish performance with highly sensitive data or substantially more complex programs.
- It does not demonstrate downstream effects on learning behavior or program outcomes.
Methodological Adaptation
The case represents an adaptation of participatory and utilization-focused evaluation rather than a replacement of those approaches with AI.
GenAI was assigned high-labor tasks such as first-pass drafting, formatting, and carrying definitions and decisions across documents. The evaluator retained responsibility for contextualization, design judgment, quality control, and ethical oversight.
Participation was supported through accessible prompts, question cards, visual explainers, and interactive tools. The intended benefit was to reduce the time and cognitive burden required for stakeholders to enter the evaluation conversation.
This is consistent with the case’s discussion of collaborative intelligence: humans and AI perform different parts of the workflow, with professional judgment remaining essential.
The Changing Role of the Evaluator
The evaluator’s role shifted away from repetitive production toward orchestration and stewardship.
| Capability | What it involved |
|---|---|
| Workflow design | Selecting tools, sequencing tasks, curating context, and shaping prompts. |
| Technical development | Working with lightweight tools and prototypes using platforms such as Lovable.dev, Google AI Studio, Codex, and Claude in terminal environments. |
| Quality assurance | Checking coherence, definitions, traceability, bias, and consistency—not merely grammatical accuracy. |
| Ethical stewardship | Managing privacy, transparency, consent, automation risks, and appropriate boundaries for AI use. |
Quality Assurance Beyond Accuracy
A central lesson from the case is that AI quality assurance cannot be reduced to asking whether an output is factually correct.
The evaluator also needed to consider:
- Coherence: Do the different planning documents tell the same methodological story?
- Definition propagation: Are key terms and decisions used consistently?
- Traceability: Can important claims and decisions be connected to their sources?
- Context: Does the output actually reflect the organization and evaluation purpose?
- Bias: Could AI-generated suggestions introduce inappropriate assumptions?
- Judgment: Is a professional evaluator still making the decision where judgment is required?
Ethical Risks
The case identifies risks including automation bias and algorithmic inequity. To manage these risks, AI outputs were treated as provisional and subjected to structured human review before being shared.
The case also highlights an important transparency issue: the client knew that GenAI was being used for the mini-apps, but explicit consent for AI assistance during TOR drafting was not initially sought.
The case subsequently describes an iterative consent model involving standardized disclosure and explicit opt-in for individual deliverables. AI inputs were limited to non-sensitive information during planning, with closed/local models used later for private materials.
Student Exercise: Reconstruct the Workflow
Imagine you are the evaluator working with another small grassroots nonprofit.
- Identify three repetitive planning tasks that could potentially be supported by AI.
- Identify three tasks where professional judgment must remain human.
- Design a human-review checkpoint for every AI-supported task.
- Identify what organizational information the AI system would need as context.
- Separate information that could safely be used from information that should not be entered into the AI system.
- Design one simple participatory tool that could help non-MERL staff understand an evaluation decision.
- Define how you would obtain and document consent for AI use.
Final Student Challenge
Design an AI-assisted evaluation planning workflow for a resource-constrained organization.
Your workflow should specify:
- The evaluation purpose and intended users.
- The planning tasks AI will support.
- The information AI will receive.
- The outputs AI will generate.
- The human review and validation process.
- The participation mechanism.
- The consent and transparency process.
- The criteria for deciding when AI should not be used.
Key Takeaway
Generative AI can make evaluation planning more feasible and potentially more participatory—but only when the evaluator remains responsible for context, judgment, quality, ethics, and final decisions.
The strongest lesson is therefore not simply “use AI to save time.” It is to redesign the allocation of evaluator effort: let AI support repetitive and labor-intensive work while protecting the human activities that create methodological quality, contextual relevance, participation, accountability, and ownership.
Recommended EvalCommunity Academy Courses
This case connects directly with three EvalCommunity Academy learning pathways:
1. AI in Monitoring & Evaluation Certificate
Build the broader foundation for responsible AI use across evaluation design, evidence analysis, data work, reporting, ethics, bias, transparency, and human oversight. The course includes a dedicated module on evaluation design, frameworks, and decision logic, including evaluation questions, theories of change, indicators, TORs, and evaluation matrices.
2. Custom GPTs for Evaluators
Learn how to build a tailored AI assistant around evaluation workflows, templates, organizational context, instructions, knowledge, testing, governance, and human review.
3. AI Agents for Evaluators Certificate
Move from individual AI assistance toward reusable, validated AI workflows and no-code agents for evaluation design, reporting, coding, indicator tracking, learning, and follow-up.
AI in Monitoring & Evaluation → Custom GPTs for Evaluators → AI Agents for Evaluators.This progression moves from responsible AI foundations, to contextual AI assistants, to more structured and reusable AI workflows.
Source & Further Reading
Organization and Tools Referenced
- My Very Own Bed — official website
- My Very Own Bed — Our Work
- My Very Own Bed — Referral Partners
- Impact Wizard — Theory of Change / Logic Model tool
Pertinent Methodological Sources
- Cousins, J. B., & Whitmore, E. (1998). “Framing participatory evaluation.” New Directions for Evaluation, 80, 5–23. DOI: 10.1002/ev.1114.
- Fetterman, D. M. (1996). Empowerment Evaluation: Knowledge and Tools for Self-Assessment and Accountability. Sage.
- Fetterman, D., & St. Roseman, P. (2025). “AI-informed empowerment evaluation.” AEA365.
- Parasuraman, R., & Riley, V. (1997). “Humans and automation: Use, misuse, disuse, and abuse.” Human Factors, 39(2), 230–253. DOI: 10.1518/001872097778543886.
- Patton, M. Q. (2008). Utilization-Focused Evaluation (4th ed.). Sage.
- Wholey, J. S. (1987). “Evaluability assessment: Developing program theory.” New Directions for Program Evaluation, 33, 77–92. DOI: 10.1002/ev.1457.
- Wilson, J., & Daugherty, P. R. (2018). “Collaborative intelligence: Humans and AI are joining forces.” Harvard Business Review.
Additional AI and Evaluation Sources Identified in the Chapter
- Anuj, H., Raimondo, E., & Den Boer, H. (2027). “Balancing innovation and rigor at the World Bank Independent Evaluation Group: Thoughtful integration of artificial intelligence for evaluation.”
- Catsambas, T. (2027). “Artificial intelligence in evaluation practice: Uses, good practices, and emerging competencies.”
- Dell’Acqua, F., McFowland, E., Mollick, E. R., et al. (2023). “Navigating the jagged technological frontier.” Harvard Business School Working Paper 24-013. DOI: 10.2139/ssrn.4573321.
- Fung, G. (2027). “Designing with generative AI: Using GPT-4 to build a monitoring, evaluation and learning framework.”
- Greenstein, N., & Cho, S.-W. (2025). “Ethics & equity in data science for evaluators.”
- Jacob, S. (2027). “Beyond the radar: Addressing undeclared AI use in evaluation practice.”
- Stępień, M. (2027). “Scoping evaluation with GenAI: A case study of the partnership principle in Poland.”
Source note
This case study is an educational synthesis based primarily on Paschke’s Chapter 2. Descriptions of the workflow, tools, results, limitations, ethical issues, and methodological adaptation are derived from that source. External organization and tool links are provided for learners who want to explore the resources referenced in the case.
Primary source:
Sadie Paschke, “Threading the needle: Making evaluation affordable and participatory for grassroots nonprofits,” in From Algorithms to Evidence: Using Generative AI in Evaluation Practice, Chapter 2, pp. 19–29.
Read the full chapter PDF
· DOI: 10.4324/9781003799139-2
Disclaimer
This case study is provided for educational purposes. It does not constitute an independent evaluation of My Very Own Bed, nor does it independently validate the effectiveness, accuracy, security, or suitability of the AI systems and tools described by the original author. Claims about results, workflow performance, platform limitations, consent, and AI use reflect the cited source and should not be interpreted as independently verified empirical findings. AI tools and platform capabilities can change over time. Evaluators should apply their own professional judgment, organizational policies, data-protection requirements, and applicable ethical standards before adopting similar workflows.
Practical AI for Monitoring & Evaluation
