
From Chatbot to Autonomous M&E Agent
EvalCommunity Academy
From Chatbot to Autonomous M&E Agent: 10 Levels of AI Automation
How evaluators can move from asking AI questions to building useful, repeatable and human-supervised workflows.
Most people are still using AI at Level 1
For many M&E professionals, working with AI still means opening a chatbot, typing a question and copying the answer into a document.
There is nothing wrong with that. It can save time. But it is only one way of using AI.
Newer agentic systems can work with documents and data, use external tools, perform several steps in sequence, delegate tasks to other agents and run scheduled workflows.
That changes the question from:
to something more like:
For M&E, that distinction matters. A good agent should reduce repetitive work without taking responsibility away from the people who need to make the final judgement.
The 10 levels of AI automation for M&E
Think of these as a progression rather than a checklist. You may only need Level 4 for one task and Level 8 for another.
Level 1 — AI as an assistant
You ask. AI responds.
This is the familiar starting point. You might ask AI to suggest evaluation questions, draft an indicator or help improve an interview guide.
Level 2 — AI with context
Give the AI something real to work with.
Instead of asking a generic question, provide the programme documents, Theory of Change, MEL plan, evaluation framework or other relevant material. The quality of the answer often improves because the AI is working from your actual context.
Level 3 — AI with instructions
Tell the system how you want it to work.
You can establish rules for evidence, sources, uncertainty, terminology, formatting and quality assurance. For example, you might instruct the system to separate evidence from interpretation and flag claims that cannot be verified.
Level 4 — AI workflow
Turn a task into a sequence.
Rather than asking AI to “check this indicator”, define the process: review the definition, check measurability, examine the data source, identify weaknesses and suggest improvements. The same process can then be reused.
Level 5 — AI with tools
Let AI work with the systems you already use.
This could mean working with spreadsheets, documents, databases, search tools or other connected applications. Instead of describing a dataset to AI, for example, the system can inspect the actual file and help analyse it.
Level 6 — AI team
Give different jobs to different agents.
One agent might review qualitative evidence while another examines quantitative results. A third can act as a reviewer. A coordinating agent can then bring the outputs together.
Level 7 — Background work
The agent keeps working after you leave.
Some agent platforms can continue a task in the background. That could be useful when a research task takes longer than a normal conversation or when you want the system to prepare material while you work on something else.
Level 8 — Scheduled AI
Make useful tasks recurring.
For example, an agent could check for new evaluation publications every Friday and prepare a short evidence update. You no longer need to remember to start the task yourself.
Level 9 — Coordinated M&E agents
Connect specialised jobs into one workflow.
An evidence scout can find material, a screening agent can decide whether it is relevant, an analysis agent can extract findings and a reviewer can check the result before a briefing agent prepares the final summary.
Level 10 — Autonomous M&E workflow
The system runs a defined process with limited intervention.
At this point, the system can monitor a workflow, perform several steps, identify issues, prepare outputs and bring important decisions back to a person. This is where an AI tool starts becoming an AI system.
A practical example: an AI evidence monitor
Consider a simple problem that many evaluators know well: keeping up with new evidence.
Instead of manually checking the same sources every week, you could design a workflow with several specialised steps:
Evidence Scout
Looks for new studies, evaluations and reports in selected sources.
Screening Agent
Checks whether the material meets your relevance criteria.
Evidence Analyst
Extracts the methodology, population, intervention, outcomes and key findings.
Quality Reviewer
Checks important claims against the underlying source.
Synthesis Agent
Organises the relevant findings and adds them to the evidence record.
Briefing Agent
Produces a short update for the evaluator.
The evaluator still decides what the evidence means for the evaluation. The system simply removes much of the repetitive searching, sorting and first-pass analysis.
More autonomy also means more responsibility
It is tempting to think that the more autonomous an AI system becomes, the better it is. In M&E, that is not necessarily true.
An agent that can access files, search the web, analyse information and take actions needs appropriate limits. This becomes particularly important when working with confidential programme information, beneficiary data or evidence that could influence important decisions.
Autonomous execution does not mean autonomous accountability.
Before giving an agent more freedom, ask:
- What information can it access?
- What can it change or send?
- Which information should never be shared with the model?
- Which decisions require human approval?
- How will sources and outputs be checked?
- What happens when the system is uncertain?
- Can you see what the agent did and why?
Where Hermes Agent fits
Hermes Agent is one useful platform for exploring this model of agentic work. Its capabilities provide practical examples of several of the levels described above, including persistent instructions, specialised skills, connected tools, background work, sub-agents, scheduled tasks, profiles, browser automation and API access.
For example, a persistent instruction file can establish how an agent should operate. Skills can give it specialised capabilities, while connected tools allow it to work beyond the conversation itself. Sub-agents can divide a larger task into smaller pieces.
The important point is not to learn Hermes for its own sake. The bigger lesson is understanding how to design an M&E workflow that AI can perform reliably. The same thinking can then be applied to other agent platforms.
Start with one task
You do not need to build a Level 10 system tomorrow.
In fact, starting small is usually the better approach.
Try this with your own M&E work
Think about a task you perform repeatedly. It might be checking indicators, screening documents, monitoring new evidence, reviewing survey responses or preparing routine reports.
Then ask yourself:
- Which part of this task is repetitive?
- Could AI handle that part?
- Could I describe the process as a series of steps?
- Would the AI need access to a file, dataset or external tool?
- Where should a human check or approve the result?
The real opportunity
The next step in AI for M&E is not simply getting better answers from a chatbot. It is designing workflows where AI can take care of repetitive work, use the right tools, check its outputs and bring important decisions back to people.
Key takeaway
Start with one workflow that is genuinely worth improving.
Automate one step. Test it. Add quality checks. Give the system only the access it needs. Then decide whether more autonomy would actually help.
The goal is not maximum automation. The goal is better M&E work with less repetitive effort and enough human oversight to trust the result.
EvalCommunity Academy
Practical guidance for using AI in monitoring, evaluation, learning and evidence work.
