Episode 10: From Raw Data to Draft Findings
- Categories AI Series
- Date February 28, 2026
EVALCOMMUNITY AI SERIES — EPISODE 10
From Raw Data to Draft Findings
Turning days of manual synthesis into hours — without losing rigor
One of the biggest bottlenecks in evaluation is moving from raw data to a structured findings section.
You’ve spent weeks in the field. Now you’re staring at:
- 40 pages of interview transcripts
- Survey tables with hundreds of rows
- Pages of field notes
- A donor deadline looming
The traditional approach—printing transcripts, highlighting with markers, creating endless coding spreadsheets—can take days or even weeks. But what if you could turn that into hours, while actually improving your analytical rigor?
In this episode, we share a practical AI workflow that transforms how evaluators move from raw data to draft findings. It’s not about cutting corners—it’s about clearing the path so you can focus on what really matters: interpretation, judgment, and insight.
The Problem: Synthesis at Scale
Qualitative analysis has always been time-intensive. Reading and re-reading transcripts, identifying themes, pulling quotes, triangulating with survey data—it’s essential work, but it’s also where evaluation timelines get stretched thin.
The challenge isn’t just speed. It’s consistency.
When you’re working through 40 interviews over two weeks, your own interpretation can drift. What felt like a theme on Monday might get coded differently by Friday.
We needed a way to:
- Analyze the full dataset holistically, not piece by piece
- Maintain consistency across hundreds of pages
- Integrate qualitative and quantitative evidence seamlessly
- Generate a structured foundation that still leaves room for evaluative judgment
The Solution: A Three-Step AI Workflow
After testing multiple approaches across real-world evaluations, we’ve landed on a workflow that consistently delivers. It uses large-context AI platforms—specifically, Projects in Anthropic’s Claude.ai or the file upload features in OpenAI tools—to analyze entire datasets at once.
Upload and Analyze the Full Dataset
Instead of fragmenting your analysis into small chunks, upload everything at once. Claude Projects (or ChatGPT’s file upload) can handle dozens of transcripts, survey tables, and field notes in a single context window.
“Based on ALL uploaded transcripts and survey data, identify the 5–7 main themes related to program effectiveness. For each theme, provide representative quotes and relevant statistics.”
The AI reads across all your materials, identifying patterns that might take a human coder days to spot. It doesn’t just list themes—it pulls supporting evidence and connects qualitative quotes to quantitative survey results.
What you get:
- Cross-cutting themes that appear consistently across data sources
- Integrated qualitative and quantitative evidence
- Clear, representative quotations for each theme
- Early pattern recognition that might challenge your initial assumptions
Time investment: 30-45 minutes
Organise by Evaluation Framework
Raw themes are useful, but they don’t yet form a coherent findings section. The next step is to structure them in a way that aligns with professional evaluation standards.
“Organise these themes under the OECD-DAC criteria: relevance, effectiveness, efficiency, impact, sustainability. Create an outline for the findings section.”
Why OECD-DAC? Because it’s the international standard for evaluation. Using this framework ensures:
- An internationally recognised structure that donors expect
- Clear evaluative logic that separates different dimensions of performance
- Strong alignment with reporting requirements across most development organizations
Within minutes, the AI transforms your list of themes into a structured outline. “Effectiveness” gets its own section. “Sustainability” another. Each theme finds its logical home.
What you get:
- A clean, professionally structured findings outline
- OECD-DAC aligned headings and subheadings
- Clear mapping between your evidence and evaluation criteria
Time investment: 15-20 minutes
Draft the Narrative
Now for the heart of the findings section. With your outline in place, you can ask the AI to draft complete narrative paragraphs.
“Draft 3–4 paragraphs for the ‘effectiveness’ section. Begin with an overall finding, provide supporting evidence, and note key challenges. Use professional evaluation language.”
The AI draws on the themes, quotes, and statistics from Step 1, weaving them into coherent prose. It starts with an overarching finding, layers in evidence, and acknowledges limitations or challenges—just as a good evaluator would.
What you get:
- Draft narrative that’s 70-80% ready for final editing
- Consistent professional tone throughout
- Evidence properly integrated, not just tacked on
Time investment: 45-60 minutes
The Result: From 5 Days to 5 Hours
In approximately 2 hours, evaluators can generate a comprehensive draft findings section that would traditionally take 4-5 days.
But here’s the critical point: that’s just the draft. The remaining time is dedicated to what truly matters:
- Verifying interpretations against the original data
- Strengthening causal reasoning and logical connections
- Adding contextual insight that no AI could know
- Ensuring ethical and methodological soundness
The workflow doesn’t replace evaluative judgment—it accelerates the synthesis so evaluators can focus on insight.
What This Workflow Is (and Isn’t)
✅ It is:
- A tool for managing large volumes of qualitative data
- A way to ensure consistent theme identification
- A method for quickly structuring findings against international standards
- A time-saver that frees evaluators for higher-order thinking
❌ It is not:
- An excuse to skip reading your transcripts
- A replacement for field experience and contextual knowledge
- A way to generate findings without validation
- A substitute for ethical oversight
Try It Yourself
This workflow works best when you:
- Use platforms with large context windows – Claude Projects (Anthropic) or ChatGPT with file upload
- Upload all relevant materials at once – transcripts, survey tables, field notes
- Iterate on the prompts – refine based on what the AI produces
- Always review against original sources – the AI draft is a starting point, not a final product
We’ve used this approach across health, education, and livelihood evaluations. It consistently cuts synthesis time by 70-80% while maintaining—and often improving—analytical rigor.
Want to Go Deeper?
This workflow is part of our practical training on responsible AI use in Monitoring & Evaluation. We cover prompt engineering, ethical safeguards, quality assurance, and integration with traditional methods.
The EvalCommunity AI Series is designed for practicing evaluators who want to use AI responsibly and effectively. Each episode focuses on a practical workflow you can implement immediately.
Episode 10
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
