How Do Organizations Learn?
EvalCommunity Case Study
How Do Organizations Learn? The Diffusion of Scientific Evidence on Generative AI
What a World Bank field experiment tells us about evidence diffusion, organizational hierarchy, and learning in the age of generative AI
Based on research by Mahvish Shaukat, Andreas Stegmann, and Mattie Toma
Research note: This case study is an educational adaptation of World Bank Policy Research Working Paper 11305, How Do Organizations Learn? The Diffusion of Scientific Evidence on Generative AI, published in February 2026. The findings and interpretations reported in the original paper are those of the authors and do not necessarily represent the views of the World Bank.
The question behind the study
Most organizations produce more information than their people can realistically absorb. Evaluation reports, research papers, monitoring data, learning products, guidance notes, dashboards, and policy briefs may all be available. Yet availability does not mean that the evidence travels.
Someone has to notice it. Someone has to decide that it matters. Someone has to pass it on. And, eventually, someone else has to remember enough of it for the evidence to influence a conversation or a decision.
This is the starting point for an important 2026 study by Mahvish Shaukat, Andreas Stegmann, and Mattie Toma. The researchers wanted to understand how scientific evidence moves through a large organization and, in particular, whether the person who receives the evidence first affects how far it travels.
The practical question
If you have important evidence that you want an organization to use, who should receive it first?
Why this is relevant to M&E
For monitoring and evaluation professionals, this is more than a question about internal communication. Evidence use is often the point at which an otherwise strong evaluation either becomes useful or disappears into a document repository.
An evaluation can be methodologically sound and still have little influence if its findings do not reach the people who make decisions. A learning product can be well designed and still fail to change practice. A donor report can accurately describe results without creating much organizational learning.
The study therefore offers a useful reminder: evidence production and evidence use are different parts of the evidence-to-decision process.
The authors place the question within a broader literature on organizational learning and information diffusion. They note that hierarchy can slow, filter, or distort communication, while individual beliefs about the relevance or credibility of information can also affect whether it is transmitted.
The setting: the World Bank
The field experiment was conducted at the World Bank headquarters in Washington, D.C. The study population consisted of 6,254 employees working across 511 divisions.
The researchers were interested in the early stages of organizational learning around generative AI. Rather than asking employees abstract questions about whether AI was important, they gave selected employees actual scientific evidence about the workplace effects of ChatGPT and then followed what happened next.
This made the study unusual. It was not simply a survey of attitudes. It was a randomized field experiment embedded in a real organization.
What evidence was shared?
The experiment was conducted in two rounds between December 2023 and April 2024.
Round 1: productivity
Seeds received evidence from Noy and Zhang (2023), an experimental study examining the effects of ChatGPT on skilled professionals performing writing tasks. The study reported improvements in task completion time and output quality.
Round 2: creativity
Seeds received evidence from Hauser and Doshi (2024) concerning ChatGPT’s effects on creativity. Participants completed a short-story task, and the study examined third-party assessments of the novelty and usefulness of the resulting work.
This distinction was useful because the researchers could examine evidence about different dimensions of generative AI rather than treating “AI impact” as one undifferentiated concept.
The experiment in simple terms
One employee was randomly selected from each division to participate in the initial survey. The researchers called these employees seeds.
In 75% of divisions, the seed received the scientific evidence. The remaining 25% formed the control group and completed the survey without receiving evidence about ChatGPT’s effects.
Within the treated divisions, the researchers randomly varied whether the evidence was given to a senior or junior employee.
Senior seed
World Bank grades GG or above, including positions such as Senior Specialist, Senior Economist, or Director.
Junior seed
World Bank grades GF or below, including positions such as Economist, Analyst, or Assistant.
The experiment also used a saturation design across the World Bank’s 42 Program Management Units. PMUs were randomly assigned to different proportions of treated divisions, allowing the researchers to investigate potential spillover effects.
The researchers tested more than seniority
The study did not assume that hierarchy was the only reason people might share evidence. Two additional treatments were introduced.
1. What did employees think their colleagues were doing?
The researchers experimentally changed seeds’ beliefs about colleagues’ use of ChatGPT and their views about whether World Bank employees should use it frequently.
In the positive-information condition, seeds were told that 75% of colleagues in a relevant subset used ChatGPT and 80% agreed or strongly agreed that it should be used frequently. In the negative-information condition, the corresponding figures were 20% and 18%.
The treatment successfully changed participants’ beliefs. But that did not translate into greater evidence sharing.
2. Did stronger credibility signals make evidence travel further?
In the first round, some seeds received additional information about the credibility of the evidence. They were told that the study had been published in Science, that the authors were affiliated with MIT, and that the findings had been independently replicated by researchers from Harvard Business School and Wharton.
Again, the intervention changed perceptions of credibility to some extent. But the researchers did not find a significant increase in sharing intentions, actual engagement, or colleagues’ recall.
How was diffusion measured?
This is one of the most useful aspects of the study for M&E professionals. The researchers did not define “evidence use” simply as whether someone had received an email.
| Stage | Measure | Why it matters |
|---|---|---|
| Intention | Seed’s stated intention to share | Captures willingness to transmit evidence |
| Engagement | Clicks on the shared materials | Provides behavioural evidence of engagement |
| Recall | Colleagues’ ability to identify study details | Tests whether information actually travelled and was retained |
For the contact survey, colleagues were asked open-ended questions. Their responses were manually coded for whether they correctly identified details such as the study authors, institutional affiliations, journal, or noteworthy features of the research.
That is a much harder test than simply asking whether someone had seen a link.
The main finding: hierarchy mattered
The clearest result concerned the organizational rank of the initial recipient.
14.2 percentage points
Senior seeds were 14.2 percentage points more likely than junior seeds to report that they were likely or very likely to share the evidence with colleagues.
15.8 percentage points
The likelihood of multiple clicks on the infographic link was 15.8 percentage points higher when the evidence was seeded with a senior employee.
5.0 percentage points
Colleagues of senior seeds were 5.0 percentage points more likely to recall at least one notable feature of the study than colleagues of junior seeds.
These results were statistically significant in the study’s main analyses. The paper reports p = 0.01 for the sharing-intention result, p < 0.01 for the multiple-click outcome, and p = 0.03 for colleagues’ recall.
But there is an important qualification
It would be too simple to conclude that seniority itself explains everything.
Senior and junior employees differed in several observable characteristics. Senior employees were, on average, more educated, more likely to hold managerial responsibilities, and more likely to occupy policy-oriented positions. The authors therefore discuss the possibility of both selection into senior positions and an independent effect of organizational status.
The results nevertheless persisted when employees with managerial responsibilities were excluded. The authors suggest that senior employees may have greater status, more central network positions, or fewer interpersonal barriers when sharing information.
This distinction matters. The paper provides evidence that organizational rank is associated with stronger diffusion under the experimental design, but it does not justify the simplistic rule that every organization should always route evidence through senior management.
The result that may matter most for M&E
Perhaps the most striking finding was not the size of the seniority effect. It was how little evidence travelled overall.
In Round 1, only about 3% of colleagues in treated divisions recalled at least one relevant feature of the evidence two to three weeks after it had been shared with the seed. In Round 2, the researchers measured no recall.
This happened even though a majority of seeds expressed an intention to share and there was substantial engagement with the materials.
In other words:
Wanting to share is not the same as sharing. Clicking is not the same as learning. Receiving evidence is not the same as using it.
That distinction is highly relevant to evaluation practice.
What the researchers expected — and what happened instead
The study included an additional and particularly interesting element: experts were asked to predict which treatments would have the strongest effects.
The experts expected changing beliefs about colleagues’ adoption and attitudes toward ChatGPT to be important. They expected hierarchy to be the least relevant predictor.
The experimental results told a different story.
This is a useful lesson for M&E professionals because it illustrates why plausible explanations should not be treated as established mechanisms without testing them. People inside and outside an organization may have strong intuitions about why evidence does or does not spread. Those intuitions can be wrong.
What this means for evidence use in M&E
The study does not test an M&E dissemination model directly. The following implications are therefore an EvalCommunity interpretation of the findings for M&E practice, rather than conclusions claimed by the authors.
1. Identify the people who can move evidence
When an evaluation produces an important finding, do not think only about the target audience. Think about the people through whom information actually travels inside the organization.
Formal seniority may be one factor. Technical authority, cross-team relationships, institutional trust, or network centrality may be others.
2. Measure what happens after dissemination
An evaluation communication plan should ideally go beyond counting downloads, emails, workshops, or presentations.
Where feasible, ask whether users engaged with the evidence, whether the findings reached secondary audiences, whether people remember the main message, and whether the evidence affected a decision or adaptation.
3. Treat evidence diffusion as part of the learning system
If an organization says that it is learning from monitoring and evaluation, it should have some way of examining whether learning is actually taking place.
This can include simple measures of reach and recall, structured reflection with decision-makers, evidence-use logs, management-response tracking, or follow-up interviews.
4. Do not assume that better evidence automatically creates better uptake
Credibility remains important to good research practice. But the experiment shows that adding credibility signals alone did not produce a measurable increase in evidence diffusion in this setting.
The quality of the evidence and the organizational pathway through which it travels are related but different problems.
5. Be careful with “AI adoption” metrics
The study is about evidence concerning generative AI, but the researchers did not find a clear overall increase in actual adoption among contacts. The strongest evidence concerns information sharing, engagement, and recall.
That distinction is important. Evidence that travels through an organization is not necessarily evidence that changes behaviour.
A useful evidence-to-use chain for M&E teams
The study suggests a practical sequence that M&E teams can use when designing evidence-dissemination strategies:
1. Produce — What did the evaluation or monitoring system actually find?
2. Seed — Who is best positioned to start the conversation?
3. Engage — Did the first recipients actually engage with the evidence?
4. Diffuse — Did the evidence reach people beyond the initial audience?
5. Recall — Can people accurately explain the important findings?
6. Use — Did the evidence contribute to a decision, adaptation, or action?
This is not a framework tested by Shaukat, Stegmann, and Toma. It is a practical M&E interpretation of the stages their experiment helps us distinguish.
Possible indicators for an M&E evidence-use system
| Dimension | Example indicator |
|---|---|
| Reach | Percentage of intended evidence users who received the finding |
| Engagement | Percentage of users who opened, read, discussed, or interacted with the evidence product |
| Diffusion | Percentage of secondary audiences reached beyond the original recipients |
| Recall | Percentage of users able to accurately identify key findings |
| Use | Examples of decisions, adaptations, or actions informed by the evidence |
What does this mean for AI in M&E?
The study has an interesting second layer for today’s M&E professionals. The evidence being disseminated was about generative AI, but the organizational problem is broader: how does an organization learn about a new technology well enough to use it responsibly?
That question is increasingly relevant inside M&E teams. Organizations are experimenting with AI for evaluation design, qualitative analysis, data processing, reporting, evidence synthesis, dashboards, and workflow automation. But acquiring AI tools is not the same thing as developing organizational capability.
The study suggests that the social and organizational side of AI adoption deserves as much attention as the technology itself. Who introduces the evidence? Who validates it? Who explains it to colleagues? Who can challenge an AI-generated conclusion? Who decides when an AI-assisted workflow is appropriate?
For M&E leaders, these are organizational learning questions as much as technology questions.
The bigger lesson
AI capability is not created simply by giving staff access to AI tools. Organizations need people who can evaluate evidence, test outputs, understand limitations, share learning, and build reliable workflows around the technology.
A methodological lesson for evaluators
There is also something here for the way we conduct M&E research.
The researchers did not simply ask employees which factor they thought mattered. They experimentally varied who received the evidence and what information they received, then observed what happened.
That design allowed them to distinguish between explanations that sounded plausible and explanations that were supported by the observed data.
For evaluators, this is a useful reminder to ask:
- What is our theory about how evidence will lead to learning?
- Which part of that theory can actually be observed?
- Which assumptions are we taking for granted?
- Can we test whether our dissemination strategy is working?
- Are we measuring outputs, engagement, learning, or actual use?
Questions for reflection
For your own organization:
- Who are the people through whom evaluation evidence actually travels?
- Are they necessarily the most senior people?
- How do you currently measure whether evidence reaches secondary audiences?
- Can your stakeholders recall the most important findings from your last evaluation?
- What happens between publication of an evaluation and the decision that follows it?
- Where could AI help strengthen that evidence-to-learning pathway?
What should M&E professionals take away?
The World Bank study does not offer a simple formula for organizational learning. Its value is that it makes a familiar problem measurable.
Evidence can be produced and still fail to travel. People can intend to share it without doing so. People can engage with a document without remembering it. And an organization can be surrounded by high-quality research without necessarily learning from it.
The experiment provides evidence that organizational hierarchy can shape the movement of scientific information. At the same time, the limited overall diffusion is a reminder that evidence use requires more than a good report, a credible source, or an email with a link.
For M&E professionals, the practical challenge is therefore not only to produce better evidence. It is to understand the organizational pathway through which evidence becomes knowledge and, eventually, action.
Evidence has to travel before it can be used.
The quality of an evaluation matters. So does what happens to its findings after the report is finished.
Continue learning with EvalCommunity Academy
Build the skills to work with AI in real M&E workflows
If this case raises questions about how AI can support evidence analysis, evaluation practice, and organizational learning, the EvalCommunity Academy offers two practical certificate pathways designed specifically for M&E professionals.
AI in Monitoring & Evaluation Certificate
Learn how to apply AI across evaluation design, data collection, qualitative and quantitative analysis, evidence synthesis, reporting, governance, ethics, and human-in-the-loop quality assurance.
AI Agents for Evaluators Certificate
Move from using individual AI tools to designing no-code AI workflows and agents for reporting, qualitative coding, data quality review, indicator tracking, evidence synthesis, learning, and follow-up.
Both courses are designed around practical M&E work rather than generic AI training. The Academy also offers the two courses together as the AI for M&E Professional Bundle.
Main reference
Shaukat, M., Stegmann, A., & Toma, M. (2026). How Do Organizations Learn? The Diffusion of Scientific Evidence on Generative AI. World Bank Policy Research Working Paper 11305.
Read the full World Bank working paper
The paper reports that the study was preregistered and that a verified reproducibility package is available through the World Bank’s reproducibility platform.
