AI Most Significant Change for Monitoring and Evaluation
Last updated: June 2026 | Version 1.0
Estimated reading time: 45–50 minutes
EvalCommunity Academy Practical Tutorial
AI Most Significant Change for Monitoring and Evaluation: A Practical Step-by-Step Tutorial
Learn how to use AI to collect, score, and select transformative beneficiary stories for monitoring, evaluation, accountability, and learning (MEAL).
Quick Answer
AI Most Significant Change (MSC) uses artificial intelligence to collect, score, and select the most transformative beneficiary stories. This tutorial teaches M&E professionals how to scale MSC analysis from hundreds to thousands of stories, identifying the most powerful examples of program impact for learning and accountability.
⚠️ What Is Most Significant Change (MSC)?
Most Significant Change is a participatory monitoring and evaluation method that collects stories of change from beneficiaries and then selects the most significant stories through group discussion. AI helps process large numbers of stories, but human selection of the final “most significant” stories remains essential.
This tutorial focuses on using AI to scale the collection, initial scoring, and analysis phases of MSC.
Before You Start
To complete this tutorial you will need:
- Beneficiary stories (from interviews, written submissions, focus groups)
- Spreadsheet software (Excel, Google Sheets, or similar)
- Access to an AI tool (examples include ChatGPT, Claude, Gemini, or equivalent AI systems)
- Understanding of MSC methodology (see resources section)
Note: The specific AI tools mentioned are examples only. Equivalent AI systems can also be used.
What You Will Learn
Who This Tutorial Is For
MEAL officers
Program managers
Evaluators
NGO staff
Humanitarian practitioners
Most Significant Change Domains AI Can Identify
| Change Domain | Description | Example Story Excerpt |
|---|---|---|
| Livelihoods乳 | Changes in income, employment, assets | “I now earn enough to send my children to school.” |
| Health乳 | Changes in health status, access, behaviors | “I learned about nutrition and my family is healthier.” |
| Education乳 | Changes in learning, skills, enrollment | “My daughter now attends school regularly.” |
| Empowerment乳 | Changes in confidence, decision-making, agency | “I now speak at community meetings without fear.” |
| Relationships乳 | Changes in social connections, cooperation | “Our community now works together to solve problems.” |
| Resilience乳 | Changes in ability to cope with shocks | “We saved money during the good season to prepare for drought.” |
📊 MSC Story Volume Guidelines
Small program: 50-200 stories — Manual review possible, AI can help with scoring
Medium program: 200-1,000 stories — AI recommended for efficient processing
Large program: 1,000-10,000+ stories — AI essential for scaling MSC
Real Example: AI Scoring of MSC Stories
Original Story (Beneficiary)
“Before the program, I had no income. I couldn’t feed my children. After the training, I started a small business selling vegetables. Now I earn enough to pay school fees and have savings. My children are healthier. I feel like a different person. This changed everything for my family.”
AI MSC Scoring Output
- Significance Score: 9.5/10
- Change Domains: Livelihoods, Health, Empowerment
- Depth of Change: High (transformative)
- Reasoning: Story shows complete life transformation across multiple domains with clear before/after contrast.
Table of Contents
Why Most Significant Change Matters for Monitoring and Evaluation
Most Significant Change (MSC) captures the human stories behind the numbers. But traditional MSC is time-consuming — collecting stories is easy, but reading and selecting the most significant stories from hundreds or thousands of submissions is difficult.
The Challenge: Too Many Stories, Too Little Time
Organizations collect hundreds of stories but lack the capacity to read and select the most significant ones. AI can help score and prioritize stories, enabling participatory selection at scale.
Source: This tutorial is adapted from MSC methodology and real AI applications in international development. For foundational guidance, see BetterEvaluation: Most Significant Change.
16-Step Workflow for AI Most Significant Change
STEP 1
Define MSC Domains and Questions
What to do: Define what types of change you want to capture.
How to do it: Create 3-8 domains of change (e.g., Livelihoods, Health, Empowerment) and a simple story collection question.
Example question: “Since participating in our program, what is the most significant change that has happened in your life?”
Common mistake: Too many domains (more than 8) makes analysis difficult.
STEP 2
Design Story Collection for AI Compatibility
What to do: Structure story collection to maximize AI usability.
How to do it: Collect stories in plain text. Include metadata: storyteller ID, location, gender, date.
Example format: Each story as a separate row in a spreadsheet with columns: ID, Date, Location, Gender, Story Text.
Common mistake: Stories stored in multiple formats (PDFs, images, handwritten).
STEP 3
Review a Sample of Stories Manually
What to do: Read 20-50 stories manually to understand the range of changes.
How to do it: As you read, note patterns in change types, language, and depth of impact.
Example observations: Many stories mention income increase. Empowerment stories are often from women. Some stories describe unexpected community-level changes.
Common mistake: Feeding stories to AI without understanding the range of changes first.
STEP 4
Create a Gold-Standard MSC Scoring Sample
What to do: Manually score a representative sample of 50-100 stories.
How to do it: Have two team members independently score each story on significance (1-10). Resolve disagreements through discussion.
Example scoring criteria: 1-3 = Minor change, 4-6 = Moderate change, 7-8 = Significant change, 9-10 = Transformative change.
Common mistake: Using different scoring criteria across coders.
STEP 5
Define MSC Scoring Criteria for AI
What to do: Write clear criteria that AI can use to score stories.
How to do it: Define what makes a story more significant: depth of change, breadth of change (multiple domains), unexpectedness, number of people affected.
Example criteria: Significance score based on: (1) Depth: surface vs. life-changing, (2) Breadth: 1 domain vs. multiple domains, (3) Unexpectedness: expected vs. surprising.
Common mistake: Vague criteria that AI cannot apply consistently.
STEP 6
Design the AI MSC Scoring Prompt
What to do: Write structured instructions for AI to score stories.
How to do it: Include scoring criteria, domain identification, and output format.
You are an MSC assistant for M&E.
Score this story on a scale of 1-10, where:
1-3 = Minor change (small, surface-level)
4-6 = Moderate change (noticeable improvement)
7-8 = Significant change (major improvement)
9-10 = Transformative (life-changing)
Also identify:
- Change domain(s): Livelihoods/Health/Education/Empowerment/Relationships/Resilience
- Depth of change: Low/Medium/High
- Unexpectedness: Expected/Unexpected
Return as JSON with: story_id, significance_score, domains, depth, unexpectedness, reasoning.
Story: {story}
STEP 7
Test AI on Your Gold-Standard Sample
What to do: Process your manually scored sample through AI.
How to do it: Compare AI significance scores to human scores.
Example comparison: Human score: 8, AI score: 7 (close). Human score: 9, AI score: 4 (disagreement to investigate).
Common mistake: Expecting perfect agreement — look for patterns, not exact matches.
STEP 8
Calculate MSC Scoring Agreement
What to do: Measure how well AI matches human scoring.
How to do it: Calculate correlation between human and AI scores. Identify systematic biases (e.g., AI consistently scores lower on empowerment stories).
Interpretation: Aim for correlation > 0.8. If correlation is low, review disagreements and refine criteria.
Common mistake: Accepting low correlation without investigation.
STEP 9
Refine Scoring Criteria Based on Errors
What to do: Review stories where AI and human scores diverge significantly.
How to do it: Identify what the human noticed that AI missed (or vice versa).
Example refinement: If AI misses “unexpected” changes, add “unexpectedly,” “surprisingly,” “never thought” to the criteria.
Common mistake: Making only one attempt. Expect to iterate 2-3 times.
STEP 10
Scale to All Stories
What to do: Process all collected stories through AI.
How to do it: Use batch processing for hundreds or thousands of stories.
Example: Process 2,000 stories in batches of 50.
Common mistake: Processing without saving original story text and metadata.
STEP 11
Generate MSC Story Rankings
What to do: Sort stories by AI significance score.
How to do it: Create a spreadsheet with stories sorted from highest to lowest score.
| Rank | Significance | Domain | Story Preview |
|---|---|---|---|
| 1 | 9.8 | Livelihoods, Empowerment | “Before the program, I had nothing…” |
| 2 | 9.5 | Health, Resilience | “My children are healthy for the first time…” |
Common mistake: Only looking at top scores — also review mid-range stories for diversity.
STEP 12
Extract Change Domains by Group
What to do: Analyze which change domains are most common across stories.
How to do it: Group stories by domain and count frequencies.
| Domain | Count | % of Stories | Avg Significance |
|---|---|---|---|
| Livelihoods乳 | 420乳 | 84%乳 | 7.8 |
| Empowerment乳 | 310乳 | 62%乳 | 8.2 |
Common mistake: Not calculating average significance per domain.
STEP 13
Identify Stories for Participatory Selection
What to do: Shortlist top stories for stakeholder selection.
How to do it: Select top 20-50 stories by AI score. Also include stories that represent each domain and diverse storytellers.
Example shortlist: Top 15 by score + 5 from each domain + 5 from underrepresented groups.
Common mistake: Only using top scores — may miss important but lower-scored stories from certain groups.
STEP 14
Generate MSC Summaries for Stakeholders
What to do: Create summaries of each shortlisted story.
How to do it: Use AI to generate concise summaries (2-3 sentences) of each story.
Example summary: “A woman who started a vegetable business after training now earns enough to send her children to school. She reports increased confidence and community respect.”
Common mistake: Summaries that lose the emotional impact of the original story.
STEP 15
Facilitate Participatory Selection
What to do: Present shortlisted stories to stakeholders for final selection.
How to do it: Share summaries (not AI scores) with stakeholders. Ask them to select the 5-10 most significant stories. Discuss selections as a group.
Documentation: Record which stories were selected and why.
Common mistake: Sharing AI scores with stakeholders (which may bias their selection).
STEP 16
Report MSC Findings
What to do: Document and share the most significant changes.
How to do it: Create an MSC report with the selected stories, analysis of change domains, and aggregate statistics.
Example statement: “AI-assisted MSC analysis of 2,000 stories identified that 84% of beneficiaries reported livelihood improvements. Stakeholder selection of the top 20 stories highlighted empowerment and health as the most valued changes.”
Common mistake: Only reporting numbers without the stories themselves.
Participatory Selection: AI + Human Judgment
The Most Significant Change method is fundamentally participatory. AI supports but does not replace human selection.
AI Role
- Score all stories on significance
- Extract change domains
- Shortlist top stories
- Generate summaries
- Aggregate statistics
Human Role
- Select final most significant stories
- Discuss why changes matter
- Interpret stories in context
- Make program decisions based on findings
- Share stories with stakeholders
“The most important part of MSC is the discussion about why certain changes are considered significant. This conversation reveals values, priorities, and learning that no AI can capture.” — Rick Davies (MSC originator)
AI Prompt Library for Most Significant Change
MSC Scoring Prompt
Story Summary Prompt
Domain Classification Prompt
Practical Code Examples for Implementation
Example: Batch MSC Scoring with OpenAI API
import openai
import pandas as pd
import json
client = openai.OpenAI(api_key="your-api-key")
prompt_template = """
Score this story on significance (1-10).
1-3: Minor change
4-6: Moderate change
7-8: Significant change
9-10: Transformative change
Also identify: domains (Livelihoods/Health/Education/Empowerment/Relationships/Resilience),
depth (low/medium/high), unexpectedness (expected/unexpected).
Return as JSON.
Story: {story}
"""
def score_story(story):
response = client.chat.completions.create(
model="your-selected-model",
messages=[{"role": "user", "content": prompt_template.format(story=story)}],
temperature=0.2
)
return json.loads(response.choices[0].message.content)
df = pd.read_csv("stories.csv")
results = []
for _, row in df.iterrows():
result = score_story(row["story_text"])
results.append({
"id": row["story_id"],
"significance": result["significance"],
"domains": ", ".join(result["domains"]),
"depth": result["depth"],
"unexpectedness": result["unexpectedness"]
})
output_df = pd.DataFrame(results)
output_df.to_csv("msc_scores.csv", index=False)
print(f"Scored {len(results)} stories")
Essential Quality Checks for MSC
- Human-AI correlation: Compare AI scores to human scores on gold-standard sample (aim for correlation > 0.8)
- Domain accuracy: Check AI domain classification against human judgment
- Story representativeness: Ensure shortlist includes diverse storytellers (gender, location, age)
- Summary fidelity: Verify AI summaries accurately represent original stories
- Participatory validation: Confirm stakeholders agree with the significance of selected stories
- Bias check: Test whether AI systematically scores certain storytellers higher or lower
Frequently Asked Questions About AI Most Significant Change
What is Most Significant Change (MSC)?
Most Significant Change is a participatory M&E method that collects stories of change from beneficiaries and then selects the most significant stories through group discussion. It captures what matters most to participants, not just what the program expected.
Can AI replace the participatory selection process?
No. The participatory discussion is the heart of MSC. AI helps by scoring and shortlisting stories, but the final selection of most significant stories must be made by stakeholders discussing what change means to them.
How many stories can AI process?
AI can process thousands of stories cost-effectively. For example, scoring 10,000 stories with GPT-4 typically costs $5-15, depending on story length.
What makes a story “most significant”?
Significance is defined by stakeholders, not by AI. Common criteria include: depth of change (surface vs. transformative), breadth of change (affects one area or many), unexpectedness (surprising outcomes), and alignment with values. For related guidance on outcome identification, see AI Outcome Harvesting.
How do I ensure stories are authentic and not fabricated?
Collect stories through trusted facilitators. Use verification questions (e.g., “When did this happen?” “Who else was involved?”). For sensitive programs, see Designing Informed Consent for AI Tools for ethical guidance.
EvalCommunity Final Checklist
Key Takeaway for M&E Professionals
AI Most Significant Change transforms how we collect, analyze, and select beneficiary stories. When AI handles the heavy lifting of scoring thousands of stories, evaluators can focus on what matters most: facilitating participatory discussions about why certain changes are significant and using those insights to improve programs.
Sources and Further Reading
AI for MSC in Development
Related EvalCommunity Tutorials
Suggested Citation
EvalCommunity Academy (2026). “AI Most Significant Change for Monitoring and Evaluation: A Practical Step-by-Step Tutorial.” Retrieved from https://academy.evalcommunity.com/ai-most-significant-change-for-monitoring-and-evaluation/
