Processing Key Informant Interviews with CoLoop – Case Study
- Categories AI, Case Studies
- Date March 13, 2026
Processing Key Informant Interviews with CoLoop AI: An ITAD Case Study
What is the role of AI in processing key informant interviews for evaluations?
In monitoring and evaluation, AI tools like CoLoop.AI are used to automate the transcription, translation, and preliminary analysis of interviews. These tools can process large volumes of qualitative data faster than human transcribers, and some offer analytical features to identify themes. However, as this case shows, the technology's analytical depth often requires careful calibration and human verification to maintain evaluative rigour.
Which program did ITAD evaluate using CoLoop AI, and why?
ITAD conducted the final evaluation of the MacArthur Foundation’s “Big Bet on Nigeria” initiative—a 10‑year, $154 million program aimed at reducing corruption through transparency, participation, and accountability. The evaluation employed a mixed‑methods design, including case studies and key informant interviews (KIIs).
Given the program’s longevity, geographic scope, and the need to capture diverse stakeholder voices, the team faced a large volume of interview data. They introduced CoLoop.AI to record, transcribe, and translate over 50 interviews (many conducted in local languages), hypothesising that the tool would enable them to collect and analyse more data than traditional methods would permit.
How did the team initially intend to use CoLoop for analysis?
However, the tool’s outputs were disappointing: themes were shallow, oversimplified, and failed to capture the complexity of the qualitative data. The results simply did not address the evaluation questions adequately.
What iterations did the team try to improve AI analysis?
🔄 Prompt iteration
The team refined prompts multiple times, attempting to guide CoLoop toward more nuanced responses.
💬 Chat interface
They switched from analysis grids to the chat interface, hoping for more flexible interrogation of the data.
⚠️ Result
Neither approach produced usable analytical outputs. The tool continued to flatten context and miss critical distinctions.
What was the final working approach that succeeded?
After internal reflection and trial‑and‑error sessions, the team redesigned the workflow entirely:
- Download transcripts – Raw, verbatim transcripts from CoLoop.
- Manually code the data – Evaluators applied human‑driven coding to capture nuance and context.
- Upload coded excerpts back into CoLoop – Only the coded segments (not full transcripts) were re‑ingested.
- Analyse individual coded sets – CoLoop was asked to synthesise themes within each bounded code category.
This approach limited the context window, allowing CoLoop to focus on specific, human‑defined categories rather than the entire dataset. The result: themes became accurate, sufficiently nuanced, and evidence remained fully traceable.
How well did CoLoop perform on transcription and translation?
| Transcription accuracy | Strong – only minor mistakes after manual review of every transcript |
| Translation quality | Effective – enabled inclusion of non‑English interviews |
| Speed | Much faster than manual transcription – a clear efficiency gain |
| Analysis (initial) | Poor – shallow themes, failed to answer evaluation questions |
| Analysis (final workflow) | Good – after human pre‑coding and focused prompting |
What limitations and risks emerged?
🧩 Nuanced classification challenges
CoLoop struggled to understand case context and flattened complex stakeholder perspectives.
🌫️ Conceptual fuzziness
The tool had difficulty linking abstract concepts (e.g., “accountability”) to the specific program context.
⚙️ Over‑reliance on AI analysis
Initial expectation that AI could replace manual coding proved incorrect; thematic analysis required human framing.
How did the team mitigate these limitations?
- Quality assurance: Every transcript and all AI‑generated content were manually reviewed.
- Trial and error: Internal reflection sessions and workflow experiments led to the final approach.
- Task simplification: Breaking analysis into smaller, code‑based tasks dramatically improved output quality.
- Manual pre‑coding: Human coding before AI analysis ensured context was preserved.
- Clear definitions: Prompts embedded precise definitions of abstract concepts to reduce fuzziness.
Lessons for evaluators: when does CoLoop add value?
✅ Transcription & translation
Highly effective—saves time, expands linguistic reach.
✅ Assistive synthesis
Useful after human coding, on bounded data subsets.
⚠️ Not a replacement
Cannot replace manual coding for nuanced, context‑rich analysis.
🧠 Human first
Human framing, verification, and coding remain essential.
Frequently asked questions
What is CoLoop AI used for in evaluation?
How did ITAD use CoLoop for the MacArthur evaluation?
Did CoLoop successfully replace manual qualitative analysis?
What is the key takeaway for evaluators using AI transcription tools?
Main reference & original source
📘 This case study is based on the official ITAD guide on artificial intelligence in evaluation, published in March 2026.
Source: ITAD guide on AI (PDF) – EvalCommunity repository
Resources for further learning
Advance your skills in AI‑assisted evaluation
Explore practical guides, expert courses, and a global network of M&E professionals using AI responsibly.
The courses and articles have been developed by an experienced team of evaluators and software developers under the guidance of Fation Luli. The EvalCommunity Academy combines practical expertise in Monitoring & Evaluation with cutting-edge AI technologies to provide high-quality, accessible learning experiences for professionals around the world.
