Grant Proposal Analysis and Classification with Copilot – Case Study
- Categories AI, Case Studies
- Date March 13, 2026
Grant Proposal Analysis and Classification with Copilot: ITAD Case Study
What is the role of AI in grant proposal analysis for evaluations?
In monitoring and evaluation (M&E), AI tools like Microsoft Copilot 365 process large volumes of unstructured text—such as grant proposals—to extract evidence, identify themes, and classify documents against evaluation frameworks. The technology performs structured data extraction and initial synthesis, which human evaluators then verify and interpret. According to the OECD Development Assistance Committee, such augmentation can increase both efficiency and transparency in evaluation processes.
Which organization used Copilot for proposal analysis, and why?
ITAD, a global monitoring and evaluation consultancy, employed Microsoft Copilot 365 during its evaluation of the Wellcome Climate Impacts Awards. These awards fund short, high‑impact projects combining evidence generation, stakeholder engagement, and policy influence. The evaluation aimed to assess the awards’ effectiveness and implementation across multiple proposal cohorts.
The team introduced Copilot to manage the analytical challenge of processing dozens of dense proposals. By automating extraction and initial synthesis, evaluators could dedicate more time to sense‑making, triangulation with interview data, and developing actionable recommendations—tasks that require human judgment and contextual knowledge.
How did the ITAD team integrate Copilot into the evaluation workflow?
📋 1. Framework‑led prompting
Evaluators built a detailed analytical framework aligned with three evaluation questions. They crafted prompts that included clear components, defined criteria, and example answers to guide Copilot.
📤 2. Data extraction
Copilot extracted structured evidence from year 1 and 2 proposals, populating the framework with relevant text.
⚙️ 3. Task decomposition
For year 3 proposals, the team separated extraction from classification—first extracting data into Excel, then prompting Copilot to classify based on that structured data.
🔍 4. Synthesis & pattern ID
Copilot summarised extracted evidence to highlight recurring themes and patterns across proposals.
✅ 5. Human verification
A post‑hoc check of 15% of proposals confirmed high accuracy. AI outputs were then integrated with key informant interviews and committee scoring.
What specific tasks did Copilot perform?
- Evidence extraction: Pulling relevant information from proposals into a structured format based on the evaluation framework.
- Synthesis and thematic analysis: Identifying and summarising themes, patterns, and recurring elements across multiple documents.
- Classification: Categorising year 3 proposals against predefined criteria, using data previously extracted into spreadsheets.
Performance & verification results
| Evidence extraction verified 15% | Strong – complete, accurate, no hallucinations |
| Synthesis (relevance) rated by rubric | Strong – comprehensive alignment with framework |
| Synthesis (insightfulness) | Adequate – lacked nuanced differentiation across applicant groups |
| Classification (structured) | Good – but needed manual coding for context‑dependent cases |
| Quantification of trends | Fragile – replaced by Excel + human QA, not used in final analysis |
What limitations were identified?
📉 Quantitative fragility
Inconsistent aggregated counts → not used.
🧩 Nuanced classification
Misclassifications when criteria depended on context (e.g., longlisted vs. awarded).
⚖️ “Meets criteria” bias
Syntheses flattened distinctions between applicant groups.
💡 Insight ceiling
AI alone insufficient for deep, differentiated insights.
How did the team mitigate these risks?
- Human‑in‑the‑loop: spot checks + rubric scoring.
- Task decomposition: “extract then classify” improved transparency.
- Process fragmentation: spreadsheet‑based counting + human QA replaced AI quantification.
- Manual coding: for categories needing deep context.
- Triangulation: anchored in KIIs and committee scoring.
These measures follow guidance from World Bank IEG and UNICEF Innocenti on responsible AI use in evaluation.
Frequently asked questions
What is grant proposal analysis and classification with Copilot?
How did ITAD use Copilot in the Wellcome evaluation?
What were the main limitations of using Copilot?
Can this approach be replicated by other evaluation teams?
Main reference & original source
📘 This case study is based on the official ITAD guide on artificial intelligence in evaluation, published in March 2026.
Source: ITAD guide on AI (PDF) – EvalCommunity repository
Resources for further learning
Deepen your expertise in AI for M&E
Access practical tools, courses, and a global network of evaluators using AI responsibly.
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
