World Bank – Unsupervised Machine Learning for Project Portfolio Analysis: A Case Study
- Categories Case Studies
- Date March 13, 2026
World Bank – Unsupervised Machine Learning for Project Portfolio Analysis: A Case Study
What is unsupervised machine learning for project portfolio analysis?
Unsupervised machine learning (UML) refers to AI techniques—such as topic modeling—that identify patterns, themes, and relationships in data without requiring pre-labelled examples. Unlike supervised learning, which learns from human-coded training data, UML explores the data inductively, allowing novel insights to emerge. In evaluation contexts, UML can analyze large volumes of project documents to uncover hidden factors influencing success or failure, complementing traditional deductive methods.
Why did the World Bank IEG explore AI for portfolio analysis?
The IEG thematic evaluation on child undernutrition assessed World Bank projects across 64 countries, involving 392 unique project reports. Traditional qualitative synthesis methods—reading and manually coding documents—are time-consuming and limit the scale of analysis. The evaluation team sought to determine whether unsupervised machine learning could:
- Generate new, emergent insights from a large dataset without prior hypotheses
- Identify factors affecting project success and failure at scale
- Complement deductive supervised learning approaches with inductive discovery
- Provide visual and statistical evidence to support evaluation findings
Which AI methods were used?
📄 Topic Modeling
Unsupervised machine learning technique that identifies clusters of related words and concepts, grouping them into coherent topics based on statistical similarity
📊 t-SNE Visualization
t-distributed stochastic neighbor embedding projects high-dimensional text data onto two dimensions, revealing semantic relationships and patterns among topics
🔍 NLP & Vectorization
Natural language processing converts text into numerical representations (vectors) that preserve semantic meaning, enabling machine learning analysis
What was the step-by-step AI workflow?
1. Data preparation
392 project reports from 64 countries were collected. Text from sections labeled "factors affecting success or failure" was extracted for unsupervised analysis.
2. Topic modeling
UML algorithms clustered the text data, extracting statistically significant guide words for each cluster. Texts with the highest scores against these guide words were selected as prime examples.
3. Expert interpretation
Domain experts reviewed the prime example texts and guide words, using their knowledge to develop coherent topic descriptors (e.g., "adaptive management," "risk mitigation").
4. t-SNE visualization
Topics were mapped in a continuous two-dimensional space to visualize relationships, gradients, and patterns—revealing transitions from project-specific to country/system-specific themes.
5. Validation
IEG performed statistical analysis confirming that the UML-identified topics were key predictors of project performance, as measured by nutrition indicator achievement.
What key insights did unsupervised ML reveal?
10 prerequisites for project success
| Topic 1: Program Design and Setup | Multisectoral project design, objectives appropriate to scope, budgeting cycles |
| Topic 2: Adaptive Management | Dynamically adjusting objectives, continuity of program and staff, use of third parties |
| Topic 3: Performance Improvement Strategies | Flexibility to adapt, government commitment, operational capacity |
| Topic 4: Implementation | Technical implementation, donor-institution alignment, needs-based targeting |
| Topic 5: Risks | Community-level interventions, political instability, natural disasters |
| Topic 6: Risk Mitigation | Aligning program to context, implementation planning, capacity assessment |
| Topic 7: Leadership and Management | Working with local leaders, project management, stakeholder management |
| Topic 8: Operations | Planning and coordination, inclusivity, empowerment, collaboration |
| Topic 9: Sustainability Factors | Government-led implementation, stakeholder coordination, financial sustainability |
| Topic 10: Evaluation and Performance Review | Specific KPIs, civil unrest affecting delivery, accountability |
Patterns discovered
| Success/factor symmetry | Failure factors were the inverse of success factors—confirming the 10 topics represent necessary conditions for success |
| Implementation cycle mapping | Topics progressed clockwise through stages: design → adaptive management → implementation → sustainability → evaluation |
| Context gradients | Topics transitioned from project-specific to country/system-specific themes along the east-west axis; from sociopolitical to technical themes along the north-south axis |
| Multidimensionality matters | Projects with low intervention variety should prioritize design, implementation, and risk mitigation; M&E and sustainability are universally important |
How did the team validate these findings?
The validation approach demonstrated that unsupervised machine learning could generate robust, evidence-based hypotheses about good practices in international development—topics that were later shown to predict project success.
What were the main benefits of this approach?
🔍 Novel discovery
Identified patterns and relationships that would not be obvious to human analysts, including the mapping of topics onto project implementation cycles.
📈 Scalability
Analyzed 392 reports across 64 countries—a scale infeasible for manual inductive coding—demonstrating UML's power for large evaluation portfolios.
📊 Visual insights
t-SNE visualization revealed gradients and relationships among topics, providing intuitive understanding of complex semantic patterns.
✅ Hypothesis generation
Generated empirically validated hypotheses about prerequisites for project success, supporting evidence-based programmatic decision-making.
What limitations and risks were identified?
🧠 Expert interpretation required
Topic models produce clusters and guide words, but domain expertise is essential to interpret them into coherent, meaningful topic descriptors.
⚙️ Front-end investment
Text extraction and preparation required significant effort—though IEG has since developed automated section extraction to reduce this burden.
📄 Document quality dependence
Model performance depends on the quality, consistency, and completeness of source documents and the text extracted from them.
🔍 Validation necessity
UML findings require rigorous validation (e.g., statistical testing against performance data) to confirm they are meaningful, not spurious.
How did the team mitigate these limitations?
- Expert-in-the-loop: Domain experts interpreted and validated topic model outputs, ensuring coherence and relevance.
- Statistical validation: IEG tested whether UML-identified topics predicted actual project performance, confirming empirical robustness.
- Iterative refinement: The approach was piloted and refined, with topic descriptors reviewed and adjusted based on expert feedback.
- Automated extraction development: IEG subsequently developed automated document section extraction to reduce manual preparation.
- Triangulation with deductive methods: UML complemented supervised learning and traditional qualitative analysis, providing a multi-method perspective.
What are the key lessons for evaluators?
✅ UML enables inductive discovery at scale
Topic modeling can surface novel, validated insights from large evaluation datasets—far beyond manual feasibility.
✅ Visualization reveals hidden patterns
t-SNE and similar methods illuminate relationships (e.g., topic gradients, implementation cycles) invisible to traditional analysis.
⚠️ Domain expertise remains essential
AI identifies patterns; humans must interpret them meaningfully. UML augments, not replaces, evaluator judgment.
⚖️ Validation is non-negotiable
UML findings must be tested against outcomes to ensure they are predictive, not coincidental.
🔍 Front-end investment pays off for large portfolios
UML is most valuable for large, cross-country, or regularly updated evaluation datasets where manual analysis is impractical.
🧠 Success factors may be universal
The 10 identified prerequisites (design, adaptation, risk mitigation, sustainability) align with widely recognized good practices—suggesting transferable lessons.
Frequently asked questions
What is unsupervised machine learning in evaluation?
What did the World Bank IEG discover using unsupervised ML?
Can unsupervised ML replace human evaluators?
What is t-SNE visualization and why is it useful?
How can other organizations replicate this approach?
Main reference & original source
📘 This case study is based on the World Bank Independent Evaluation Group working paper: "Advanced Content Analysis: Can Artificial Intelligence Accelerate Theory-Driven Complex Program Evaluation?" by Samuel Franzen, Cuong Quang, Lukas Schweizer, Alexander Budzier, Jenny Gold, Mercedes Vellez, Santiago Ramirez, and Estelle Raimondo (January 2022). The case is also featured in the OECD publication "Governing with Artificial Intelligence: The State of Play and Way Forward in Core Government Functions" (2025).
Sources: World Bank PDF · OECD section on AI in policy evaluation
Resources for further learning
Advance your skills in AI for evaluation
Learn how to apply unsupervised machine learning, topic modeling, and visualization techniques in your own evaluation work. Access courses, templates, and expert guidance.
The courses and articles have been developed by an experienced team of evaluators and software developers under the guidance of Fation Luli. The EvalCommunity Academy combines practical expertise in Monitoring & Evaluation with cutting-edge AI technologies to provide high-quality, accessible learning experiences for professionals around the world.
