
From Algorithms to Evidence
From Stakeholder Voices to Evaluation Evidence: Using GenAI to Design an Evaluation
A practice-based case from Poland showing how GenAI was used to structure more than 500 stakeholder questions, support the development of evaluation questions and shape a Terms of Reference — while keeping human judgment at the centre.
Why this case matters for evaluators
This is not a case about asking AI to “design an evaluation”. It is a case about using GenAI as a working assistant during evaluation scoping: organizing stakeholder input, identifying overlaps and gaps, exploring options, and accelerating drafting — with the evaluation coordinator retaining responsibility for decisions, verification and methodological quality.
Case at a glance
Evaluation phase: Design / scoping
Policy context: European Union Cohesion Policy
Evaluation focus: The partnership principle in Poland
Stakeholder input: More than 500 submitted questions and issues
AI tools: NotebookLM and ChatGPT
Human role: Continuous supervision, correction, selection, verification and stakeholder validation
Reported process time: Three working days for the coordinator’s GenAI-supported design work
Question reduction: Approximately 500 → 79 → 40 → 20 priority questions
1. The evaluation challenge
The Polish National Evaluation Unit (NEU) was preparing an evaluation of the partnership principle within European Union Cohesion Policy for the 2021–2027 period. Rather than beginning with a predetermined set of research questions, the NEU used a “blank page” process and invited stakeholders to propose the issues and questions they believed the evaluation should address.
The response produced more than 500 questions from different stakeholder organizations, including Cohesion Policy institutions, socio-economic partners and NGOs. The submissions covered governance, multilevel coordination, the quality of participation, barriers, regional differences, and examples of good and poor practice.
The evaluation therefore faced a familiar problem for evaluators: how do you preserve broad stakeholder input while turning a very large body of contributions into a coherent and feasible evaluation design?
The design problem
The task was not simply to summarize 500 questions. The evaluation also needed a prioritized set of research questions, evaluation criteria, methods, data sources, deliverables and a realistic timetable.
2. Why GenAI was introduced
The case identifies three connected needs behind the use of GenAI:
- Scale: rapidly synthesize more than 500 diverse stakeholder submissions.
- Structure: support clustering, deduplication and gap-spotting so that prioritization could be more transparent and defensible.
- Coherence: support the development of a consistent and usable Terms of Reference incorporating stakeholder review.
The tools were explicitly used to augment rather than replace expert judgment. The chapter describes the process as a close collaboration between technology and the human evaluator, with the coordinator deciding the direction of the work and correcting AI outputs when errors, lost threads or inaccuracies appeared.
3. The human-in-the-loop workflow
The process described by Stępień can be read as a sequence of controlled stages rather than a single AI task.
500+ stakeholder submissions
↓
Preparation & intake
↓
First-pass structuring with NotebookLM
↓
Iterative synthesis with ChatGPT
↓
Draft Terms of Reference
↓
Stakeholder validation
↓
Final Terms of Reference
Human supervision remained across the entire process.
4. Step 1 — Prepare the AI-assisted work
The coordinator first supplied GenAI with information about the partnership principle. This included the relevant regulations, guidelines and the evaluation conducted during the previous programming period.
The stakeholder submissions were lightly cleaned. ChatGPT was used to support formatting, remove obvious duplicates and help create two complete structured lists: one of research issues and one of research questions.
This mattered because some stakeholders had submitted only an issue or only a question. The coordinator used ChatGPT to formulate missing questions for submitted issues and to connect questions with corresponding issues, helping prevent contributions from being lost at the next stages.
EvalCommunity practice point:
Before asking AI to synthesize a consultation, establish a complete and auditable representation of the original inputs. Do not allow incomplete records to disappear simply because they are inconvenient for the model to process.
5. Step 2 — Learn from the first AI failure
ChatGPT was initially asked to analyze the entire dataset. The result was not considered adequate: salient points were omitted and some material was reduced to overly simplistic syntheses.
The coordinator therefore paused the ChatGPT analysis and changed the workflow. NotebookLM was used for the first-pass structuring of the stakeholder material.
This is one of the most useful lessons in the case. The evaluator did not treat the first AI output as the answer. The workflow was changed because the output was not good enough.
6. Step 3 — Structure before consolidating
NotebookLM was given two files containing the stakeholder material: one with research issues and one with research questions. According to the case, it worked only from those uploaded documents.
NotebookLM supported:
- an initial evaluation context and draft objectives;
- synthesis of research issues;
- preliminary grouping of questions by theme;
- grouping by level — national, regional and local; and
- grouping according to stages of the policy cycle.
A particularly important feature was that NotebookLM did not delete the submitted issues or research questions. It created categories and assigned questions to them, providing an initial ordering of a large volume of material.
7. Step 4 — Consolidate iteratively, not all at once
After first-pass structuring, ChatGPT was used again. This time, the coordinator did not ask it to process the complete corpus in one operation. Instead, the work proceeded sequentially, focusing on one domain and its questions at a time.
For each domain, ChatGPT was used to identify:
- redundancies;
- overlaps;
- opportunities for consolidation;
- gaps; and
- potential priority questions.
This change in workflow directly responded to the earlier problem of over-compression.
8. A concrete example: when AI missed the point
The case provides a useful example involving partnership and conflict of interest in the selection of Monitoring Committee members.
ChatGPT initially proposed a question about whether applying the partnership principle affected the occurrence or avoidance of conflicts of interest in the selection process.
The evaluator considered that formulation incomplete. The intended issue was more specific: how Monitoring Committee members were selected and whether the selection process enabled conflicts of interest to be avoided.
After the evaluator clarified the problem, ChatGPT produced a revised formulation focused on the selection process and the mechanisms for preventing conflicts of interest.
The pattern to learn
AI proposes → evaluator detects conceptual loss → evaluator clarifies the intended meaning → AI reformulates → evaluator decides whether the result is suitable.
This is a practical example of why evaluation expertise cannot be separated from AI-assisted evaluation design.
9. Step 5 — Preserve traceability during consolidation
The initial corpus of approximately 500 questions was reduced to 79, then to 40, and finally to 20 critical questions for inclusion in the detailed Terms of Reference.
During consolidation, the human coordinator used colour-coding in Microsoft Word to show which original questions had been consolidated into the final questions. This created a way to trace how the final research questions were formed and which original proposals informed them.
Why traceability matters
An AI-generated final question may look reasonable while hiding what was removed, merged or reframed. The colour-coding approach created a visible connection between the original stakeholder proposals and the final evaluation design.
10. Step 6 — Use GenAI to explore the wider evaluation design
Once the research questions had been consolidated, GenAI was used to support the next elements of the Terms of Reference.
- improving the research context, objectives and scope;
- developing initial evaluation criteria;
- exploring a methodology mix;
- considering potential data sources;
- developing sampling logic;
- outlining deliverables; and
- developing a draft timetable.
The methodology options identified in the case included desk review, surveys, interviews, case studies and participatory workshops. Proposed deliverables included a methodology report, interim outputs, a final report and briefs aimed at different audiences.
These were proposals for consideration. The coordinator assessed which were suitable and how they should be developed or adapted.
11. Step 7 — Return the design to stakeholders
The draft Terms of Reference were sent to the stakeholders who had participated in the initial collection of questions and issues.
The consultation involved a first round of comments, amendments and a second round of comments. The draft was assessed positively. Stakeholders described some of the methodological treatment as “fresh” and “interesting”, while raising mainly minor clarifying comments.
Most importantly, the vast majority of stakeholders recognized the issues they had originally proposed within the Terms of Reference.
The final Terms of Reference incorporated nearly all stakeholder comments and refined the scope, research questions and evaluation methodology.
Core lesson for EvalCommunity evaluators:
AI can help process participation. It does not replace participation. Stakeholders still need an opportunity to review whether the resulting evaluation design reflects their information needs.
12. What GenAI did well
- Rapid synthesis at scale: the case reports that work that might otherwise have taken several weeks was completed in three working days of the coordinator’s GenAI-supported work.
- Structured thinking: clustering, deduplication and gap-spotting helped turn a large volume of inputs into clearer thematic areas.
- Simplified design exploration: GenAI generated alternative options for methods, criteria, data sources and scheduling, supporting deliberation.
- Stakeholder resonance: the process allowed stakeholder proposals to be taken into account, and stakeholders recognized their original issues in the final Terms of Reference.
- Resource efficiency: the case reports that the process could be managed by a single coordinator rather than requiring a larger group for the same initial design work.
13. Where GenAI struggled
Over-compression and omissions: the early “all-at-once” analysis omitted salient points. The response was chunked into one theme at a time, with document-based structuring before returning to ChatGPT.
Inconsistent granularity and overconfident generalizations: the coordinator harmonized levels of detail and verified claims against source documents.
Prompt sensitivity: output quality depended heavily on how questions were framed, so prompts were refined iteratively with explicit scope and constraints.
Risk of bias or crowding out niche topics: under-represented issues were manually reinserted during review and transparent decision logs were maintained.
14. What this case says about evaluator competencies
The case makes clear that successful GenAI-supported scoping required more than knowing how to write prompts.
The coordinator needed substantive evaluation knowledge, critical reading skills, knowledge of evaluation logic, prompt-design skills and documentation skills for traceability.
The chapter describes an evolving evaluator role in terms of a knowledge broker: someone able to connect stakeholders’ information needs with the processes that generate evaluation evidence, while also having enough technological competence to use AI-supported tools across evaluation work.
The coordinator’s competencies also evolved during the project. Prompt design improved through repeated cycles, and the discipline of questioning AI outputs strengthened critical thinking about constructs, assumptions and trade-offs in the research design.
15. A practical workflow for EvalCommunity evaluators
- Prepare the evidence. Define the evaluation purpose, scope, concepts and source materials before asking AI to synthesize them.
- Preserve the original inputs. Keep a complete record of stakeholder issues and questions before consolidation.
- Structure first. Use AI to organize themes, levels and relationships before asking it to prioritize.
- Work in manageable units. If whole-dataset analysis produces omissions or over-compression, change the workflow.
- Challenge proposed questions. Check whether the AI has preserved the actual meaning of the stakeholder concern.
- Keep a traceable decision trail. Record which original inputs contributed to consolidated questions and why decisions were made.
- Use AI to explore options. Let it propose methods, criteria, data sources and schedules, but retain methodological judgment.
- Return to stakeholders. Validate whether the AI-assisted synthesis still reflects their information needs.
- Document AI use. Keep the prompts, important outputs, corrections and decisions needed to understand how AI contributed to the work.
16. The EvalCommunity AI output check
Before accepting an AI-assisted evaluation question or design element, ask five questions:
- Fidelity: Does this accurately represent the original issue or evidence?
- Relevance: Does it contribute to the evaluation purpose?
- Specificity: Is it precise enough to investigate?
- Feasibility: Can the evaluation realistically answer or implement it?
- Traceability: Can you show where it came from and why it was retained?
17. Questions for your own evaluation practice
- What part of the scoping problem actually requires GenAI?
- What could be lost if the model compresses the stakeholder material?
- What decisions must remain with the evaluator?
- How will you identify omissions?
- How will you protect minority or less frequently mentioned issues?
- Can you trace final evaluation questions back to the original stakeholder inputs?
- How will stakeholders validate the AI-assisted synthesis?
18. The main lesson
The case shows a practical model of human–AI collaboration. AI accelerated synthesis and generated options, while the evaluator retained control over interpretation, prioritization, methodological choices, verification and stakeholder engagement.
For evaluators, the important question is therefore not simply “Can AI do this task?” but:
“Where can AI increase our capacity without transferring away the judgment that makes evaluation credible?”
Source
Stępień, M. (2027). “Scoping evaluation with GenAI: A case study of the partnership principle in Poland.” In K. Bruce, V. J. Gandhi, & S. B. Nielsen (Eds.), From Algorithms to Evidence: Using GenAI in Evaluation Practice, pp. 30–41. Routledge. Chapter DOI
EvalCommunity Academy
Practical learning for evaluators using AI in real monitoring, evaluation and learning workflows — with methodological rigor, verification and responsible use at the centre.
