World Bank – AI for Complex Portfolio Evaluation – Case Study
- Categories AI, Case Studies
- Date March 13, 2026
World Bank – AI for Complex Portfolio Evaluation: A Case Study
What is AI-powered complex portfolio evaluation at the World Bank?
AI-powered complex portfolio evaluation refers to the use of natural language processing, machine learning, and knowledge graph technologies to analyze large collections of project documents. The World Bank Independent Evaluation Group (IEG) piloted these methods to assess whether AI could support theory-driven evaluations of programs implemented over 5-10 years, by automating content analysis, generating emergent insights, and structuring evidence according to a theory of change (ToC).
Why did the World Bank IEG explore AI for evaluation?
IEG thematic evaluations assess the performance of World Bank projects and activities over extended periods. These evaluations traditionally rely on manual qualitative synthesis of project reports—a time-consuming and resource-intensive process that limits the scale of analysis. The increasing volume of evaluation evidence presents both an opportunity and a challenge: comprehensive synthesis could generate powerful learning, but manual methods cannot systematically examine such large evidence bases. The IEG pilot sought to determine whether AI could offer faster, more comprehensive, and systematic synthesis while maintaining methodological rigor and enabling theory-driven analysis.
Which AI approaches were tested?
📑 Supervised ML
Deductive coding: train models to label text according to a predefined conceptual framework (74 labels across 3 main categories)
🔍 Unsupervised ML
Inductive topic modeling: identify emergent themes from text labeled "factors affecting success or failure"
🧠 Knowledge Graphs
Structure machine learning outputs according to a theory of change schema; explore relationships among ToC components
📊 t-SNE Visualization
Project high-dimensional text data onto 2D planes to identify patterns and distinct country program characteristics
What was the step-by-step AI workflow?
1. Conceptual framework development
IEG experts developed a theory of change for child undernutrition, identifying categories: nutrition challenges, interventions, outcome achievement, and factors affecting success/failure.
2. Training data preparation
Experts manually extracted text from project reports into correspondence tables and labeled content with 74 hierarchical codes across 3 top-level categories.
3. Supervised ML model training
Multiple classification algorithms were trained on labeled data to predict codes for unseen text. Performance assessed on held-out pilot country data.
4. Unsupervised ML topic modeling
Topic modeling applied to "factors affecting success/failure" text. Emergent topics interpreted by domain experts and validated against project performance data.
5. Knowledge graph construction
ML outputs mapped onto a knowledge graph schema representing the ToC. Rule-based reasoning used to query relationships (e.g., intervention X → outcome Y).
6. Visualization and interpretation
t-SNE visualizations created to identify patterns, distinct country clusters, and relationships among topics.
How well did the AI approaches perform?
Supervised Machine Learning (Deductive Coding)
| Exact label prediction (74 labels) | Weighted F1: 0.44 – modest performance given high class imbalance and 74-way classification |
| Top-level category prediction | 90–95% accuracy – model reliably identified the correct high-level category (nutrition challenges, interventions, outcomes) |
| Training set size | 274 projects |
| Test set size | 118 projects |
Unsupervised Machine Learning (Inductive Topic Modeling)
| Topics identified | 10 coherent, domain-relevant topics (e.g., Program Design and Setup, Adaptive Management, Risk Mitigation, Sustainability Factors) |
| Novel insights | Topics aligned with widely recognized good practices; subsequent IEG analysis confirmed they were key predictors of project performance |
| Pattern discovery | t-SNE visualization revealed gradual transitions: sociopolitical → technical themes, project-specific → country/system-specific themes; topics mapped onto project implementation cycle |
| Success vs. failure factors | Failure factors were essentially the inverse of success factors; 10 topics represented prerequisites for project success |
Knowledge Graphs
| Rule-based reasoning | Successfully performed simple statistical analyses (e.g., intervention X achieved outcome Y in 78% of cases) |
| Limitations | Multilabel complexity, incomplete evidence trails, and need for more granular ToC schema prevented full theory-based evaluation |
| Potential | Could evolve into a "smart ToC" for streamlined portfolio review and evidence interrogation |
What key insights did AI reveal?
🌍 Country program distinctiveness
t-SNE visualization showed only 5 of 64 countries had statistically distinct program characteristics. The pilot country's unique focus on water, sanitation, and hygiene was visually confirmed.
🧩 10 prerequisites for project success
UML identified 10 factors that were key predictors of project performance—providing empirical evidence for good practices in international development.
⚙️ Multidimensionality matters
Projects with low intervention variety should pay particular attention to design, implementation, and risk mitigation; M&E and sustainability are universally important.
🔗 Success/factor symmetry
Failure factors were the inverse of success factors—confirming that the 10 topics represent conditions necessary for success.
What limitations were identified?
📉 Class imbalance
Some labels had >200 training examples; others had <10. This hindered SML performance on rare labels.
✍️ Manual extraction required
Training data required experts to manually extract text of interest—a time-consuming step. (IEG has since developed an automated section extractor.)
🧠 Knowledge graph complexity
Multilabel relationships, incomplete evidence, and the need for a more granular ToC schema limited theory-based evaluation.
⚖️ Front-end investment
Significant effort required to set up frameworks and prepare data—only worthwhile for large or regularly updated portfolios.
📄 Document quality
PDF quality, inconsistent formatting, and lack of standardized outcome statements affect AI performance.
How did the team mitigate these limitations?
- Iterative model optimization: Multiple algorithms and preprocessing methods were tested to improve SML performance.
- Expert validation: All UML topics were reviewed and interpreted by domain experts; findings were statistically validated against project performance data.
- Visualization for pattern detection: t-SNE provided intuitive, powerful insights that complemented quantitative metrics.
- Modular approach: SML, UML, and knowledge graphs were used for distinct purposes, allowing each to add unique value.
- Future automation: IEG's new document section extraction routine could remove the manual text extraction step in future applications.
- Focus on scalability: Methods were designed to be applied to large portfolios where the front-end investment is justified.
What are the key lessons for evaluators?
✅ SML: top-down coding works for broad categories
Predicting exact sublabels is challenging, but high accuracy on top-level categories is achievable and useful.
✅ UML: emergent insights at scale
Topic modeling can generate novel, validated findings from large datasets—scaling inductive analysis beyond manual feasibility.
✅ Visualization: powerful pattern recognition
t-SNE and similar methods reveal patterns invisible to human analysts (e.g., country distinctiveness, topic gradients).
⚠️ Knowledge graphs: promising but developmental
Useful for structured queries, but full theory-based evaluation requires more granular schemas and complete data.
⚖️ Front-end investment matters
AI for evaluation synthesis is currently most viable for large, ongoing, or regularly updated portfolios.
🧠 Human-AI collaboration is essential
Domain expertise is needed for framework design, training data preparation, validation, and interpretation. AI augments, not replaces.
Frequently asked questions
What AI methods did the World Bank IEG test?
Can AI replace human evaluators in complex portfolio evaluation?
What were the most promising AI results?
What are the main challenges of using AI for portfolio evaluation?
Where can I find the full World Bank IEG report?
Main reference & original source
📘 This case study is based on the World Bank Independent Evaluation Group working paper: "Advanced Content Analysis: Can Artificial Intelligence Accelerate Theory-Driven Complex Program Evaluation?" by Samuel Franzen, Cuong Quang, Lukas Schweizer, Alexander Budzier, Jenny Gold, Mercedes Vellez, Santiago Ramirez, and Estelle Raimondo (January 2022).
Source: World Bank Documents (PDF)
Resources for further learning
Advance your skills in AI for complex evaluation
Learn how to apply supervised and unsupervised machine learning, knowledge graphs, and visualization techniques in your own evaluation work. Access courses, templates, and expert guidance.
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
