AI Future Isn’t About Learning AI
Why this matters for evaluation work: when any funder, partner, or reviewer can generate a competent-sounding evaluation summary in seconds, the evaluator’s value stops being the summary. It becomes the judgment behind which findings matter, the courage to name an uncomfortable result, and the originality to ask the question a template wouldn’t. This tool is a self-assessment, not a scorecard — there’s no passing grade, just a clearer picture of where to focus.
1. Where Do You Stand?
Rate yourself on the eight qualities
For each quality, pick the option closest to how your M&E practice actually runs right now, not how you’d like it to run. Be honest — the point is to find where to focus, not to score well.
2. The Eight Qualities in Practice
What each one looks like in evaluation work
None of these are about rejecting AI tools. They’re about what stays valuable once everyone has access to the same tools — grounded here in the evaluation tasks and OECD-DAC criteria most evaluators work against daily.
Why tie this to OECD-DAC criteria specifically
Most evaluation reports are already structured around relevance, coherence, effectiveness, efficiency, impact, and sustainability. AI tools can draft competent text against every one of those headings. What they can’t do is decide which coherence gap actually matters, or whether an effectiveness claim survives a skeptical read — that judgment call sits with the evaluator, criterion by criterion.
3. Building the Habit, Not the Skill
These aren’t learned in a course
Prompting skills will keep getting easier to pick up as the tools improve. These qualities work the other way — they’re built through repeated small choices, not a single training session.
Notice when you’re confirming instead of discovering
Before reviewing data or an AI-generated summary, write down what you expect to find. Afterward, check whether you actually found that, or just accepted it because it matched.
Practice disagreeing with the AI in writing
Next time a tool gives you a recommendation, write your own reasoning first, before you read its justification. Compare the two afterward.
Say the uncomfortable thing while it’s still small
Findings get harder to raise the longer they sit unspoken. Naming a concern early, while it’s still one data point, is easier than naming it once it’s undeniable.
Close the loop on visibility, every time
A finding that never reaches a decision-maker didn’t happen, professionally speaking. Treat sending it to the right person as part of the deliverable, not an optional extra.
Pressure-test one causal link before you finalize it
Before signing off on a theory of change or logframe, argue against your own weakest causal link the way an external reviewer would. If you can’t defend it, that’s the assumption worth testing first, not the one to quietly leave in.
4. This Week’s Practice
Pick a few, not all eight
Trying to work on all eight qualities at once usually means working on none of them. Check off what you’ll actually attempt this week.
5. FAQ
Frequently asked questions
Isn’t learning to use AI tools still important for M&E work?
Yes, and it’s worth doing. The point isn’t to skip that skill, it’s to recognize that as the tools get easier, knowing how to use them stops being what sets one evaluator apart from another. Judgment, originality, and courage don’t get easier with a better interface.
What if I score low across most of the eight qualities?
That’s common and not a diagnosis of anything. Pick one or two to focus on this week using the practice tab, rather than trying to fix all eight at once.
Is this self-assessment scientifically validated?
No. It’s a structured reflection exercise, not a psychometric instrument, and it makes no clinical or diagnostic claims. Treat the result as a prompt for thinking, not a score to optimize.
How does this connect to the OECD-DAC evaluation criteria?
Each quality tends to matter most against a specific criterion: judgment when separating attribution from contribution in an impact section, courage when naming a sustainability risk early, curiosity when a relevance assessment risks only reaching people the program already serves. Tab 2 maps all eight this way.
Does this app send my answers anywhere?
No. Everything runs in your browser to build your reflection and downloadable file. Nothing is transmitted to or stored on EvalCommunity’s servers.
Where can I read the full article this app is based on?
See the full piece on EvalCommunity Academy: Preparing for an AI Future Isn’t About Learning AI.
Want to build the practical AI side too?
This reflection is one half of the picture. EvalCommunity Academy’s course covers the practical half — applying AI across the evaluation cycle without losing the judgment that makes the work yours.
