AI Detection in Monitoring and Evaluation: Accuracy, Risks and Responsible Use
A practical guide to the accuracy, limitations and responsible use of AI-detection tools in monitoring, evaluation, research and international development.
Last reviewed: 23 July 2026 | For evaluators, M&E/MEAL teams, researchers, universities, NGOs, donors and consultants.
Can AI detectors reliably identify AI-written evaluation reports?
Not with certainty. AI detectors estimate whether text resembles patterns found in AI-generated examples. A score can support further review, but it cannot prove who wrote a report, which tool was used or how much human editing occurred.
AI-detection tools are increasingly used to screen student assignments, reports, proposals, articles and professional documents. This creates an important question for evaluators and organisations: can these tools reliably determine whether an evaluation report was written with AI?
The answer requires caution. Detection tools can sometimes identify statistical patterns commonly associated with AI-generated text, but their results can be affected by document length, language, editing, translation, formal writing conventions and the type of material being analysed. These limitations matter in M&E because evaluation documents are often highly structured, technical, repetitive and written by multilingual teams.
This tutorial explains how AI detectors work, why false results occur, how to test them responsibly and what organisations should do before making decisions that affect students, staff, consultants or partners.
What you will learn
- What an AI detector actually measures
- Why detector scores can be wrong or misleading
- Why M&E documents are especially vulnerable to false positives
- How to run a small detector reliability experiment
- How to review suspected AI use fairly and consistently
- What to assess instead of relying only on authorship detection
- How to disclose responsible AI assistance in an evaluation report
Tutorial contents
- Why AI detection matters in M&E
- How AI detectors work
- What a detector score means
- Main limitations
- Why M&E writing is at risk
- Practical detector experiment
- Responsible decision process
- Organisational policy checklist
- What matters more than detection
- AI-use disclosure templates
- Frequently asked questions
- Next tutorial: edit AI-assisted M&E writing
- Resources and further learning
1. Why AI detection has become an M&E issue
Generative AI can help professionals organise interview notes, improve language, draft headings, summarise documents and prepare first versions of reports. At the same time, organisations need to protect research integrity, confidentiality, authorship, professional judgement and accountability.
This has encouraged some institutions to use AI detectors. A detector may be introduced in situations such as:
Education and training
Checking an evaluation assignment, thesis, course submission or professional certificate assessment.
Consultancy management
Reviewing an inception report, evaluation report, proposal, technical offer or knowledge product.
Research and publication
Screening an article, evidence synthesis, evaluation brief, case study or conference paper.
Why the stakes are high
A wrong classification may affect a learner’s grade, a consultant’s reputation, a staff member’s employment, a supplier’s contract or the credibility of an evaluation team. Detector results therefore require due process, documented review and human judgement.
2. How AI detectors work
An AI detector normally does not find a hidden label showing which tool wrote a document. Instead, it analyses characteristics of the submitted text and compares them with patterns learned from examples of human-written and AI-generated material.
Depending on the tool, the analysis may consider:
- How predictable the next word appears to be
- Variation in sentence length and structure
- Repetition of common expressions
- Paragraph rhythm and grammatical consistency
- Similarity to text generated by models included in the detector’s training data
- Patterns that the detector associates with human or machine writing
A detector therefore provides an inference. It does not directly observe who typed the text, which tool was used or how much human editing occurred.
3. What does an “80% AI” score actually mean?
Detector interfaces use different labels. One may display “likely AI-generated,” another may highlight selected sentences and another may show a percentage. These outputs are not always defined in the same way.
| Possible output | What it may indicate | What it does not prove |
|---|---|---|
| Likely AI-generated | The detector found patterns associated with its AI examples. | That a named person used a specific AI tool. |
| 80% AI | A score calculated according to that tool’s own method. | That 80% of the document was necessarily written by AI. |
| Human-written | The text did not cross the detector’s AI threshold. | That no AI assistance was used. |
Do not convert a detector score into a factual statement. Write “the tool produced a high AI-likelihood score,” not “the tool proved that the author used AI.”
4. Main limitations of AI-detection tools
OpenAI withdrew its AI-written text classifier on 20 July 2023 because of low accuracy. Its published guidance warned that the classifier was not fully reliable, performed poorly on short text, could incorrectly label human writing and should not be used as the main decision-making tool. These cautions remain useful when assessing any detector, although products use different methods.
False positive
Human-written material is classified as AI-generated. This is particularly serious when the result is used to accuse or penalise someone.
False negative
AI-generated or AI-assisted material is classified as human-written. A low score cannot confirm that AI was not used.
Short-text problem
A short paragraph gives the detector fewer signals. Results for abstracts, recommendations, emails or brief findings may be especially unstable.
Language limitation
Performance may vary across languages. Translated documents and writing by multilingual professionals may not resemble the detector’s training material.
Editing changes the signal
Human editing, copy-editing, translation, paraphrasing and formatting can change the patterns used for classification.
Training-data mismatch
A detector may be less reliable when the submitted material differs from the types of documents, sectors, languages or models used in its development.
5. Why M&E writing can be wrongly classified
Evaluation reports often use standard formats required by commissioners and donors. They may contain repeated indicator names, formal terminology, parallel findings, technical definitions and consistent recommendation structures. Those characteristics can look statistically predictable even when the text was written by a person.
| Normal M&E characteristic | Why it may look machine-like | Example |
|---|---|---|
| Standard report headings | The structure is repeated across many reports. | Methodology, findings, conclusions and recommendations |
| Repeated indicator language | Terms must remain consistent for accuracy. | “Percentage of supported facilities submitting reports on time” |
| Formal donor language | Institutional style can reduce variation. | “The programme contributed to the following outcomes…” |
| Recommendation templates | Recommendations may follow the same sentence pattern. | Action, responsible party, deadline and priority |
| Translated or edited text | Translation tools and editors can regularise style. | A French field report translated into formal English |
| Multi-author production | One editor may harmonise several writing styles. | Country inputs consolidated by a lead evaluator |
Equity consideration
A review process should consider whether a detector may disadvantage people writing in an additional language, professionals required to follow strict templates or teams whose documents have been translated and centrally edited.
6. Practical exercise: test detector reliability yourself
Before adopting a detector, test it with materials whose origin and editing history are already known. The goal is not to learn how to evade detection. The goal is to understand whether the tool produces stable and useful results for your organisation’s real documents.
Prepare five controlled samples
- Human original: a report section written before generative AI was used in your organisation.
- AI first draft: a paragraph generated from a non-sensitive fictional M&E prompt.
- Human-edited AI draft: the same paragraph reviewed by an evaluator for accuracy and clarity.
- Translated sample: a known human-written paragraph translated into another language and back into English.
- Short extract: a recommendation or executive-summary paragraph from each sample.
Record the results
| Sample | Known origin | Length | Language | Detector result | Correct? | Notes |
|---|---|---|---|---|---|---|
| 1 |
Duplicate the blank row for each sample, or copy these columns into Excel or Google Sheets.
Questions to ask after the test
- Did the tool falsely classify any known human writing?
- Did the result change when the passage became shorter?
- Did translation or editing affect the score?
- Did two detectors disagree?
- Is the tool suitable for your languages and document types?
- Would you be comfortable defending a serious decision based on this result?
7. A responsible decision process for suspected AI use
A detector may provide one signal, but it should not replace an investigation. Use the following process when AI use may violate a stated policy or contract.
Neutral questions for the author
- Please describe how you planned, drafted and revised this section.
- Which source documents support the main findings?
- Which tools, including language or AI tools, were used?
- How were generated suggestions checked against the evidence?
- Can you explain how you reached this conclusion and recommendation?
8. Organisational policy checklist
The most effective response to AI is not simply purchasing a detector. Organisations need clear rules defining acceptable assistance, prohibited practices, disclosure, privacy and accountability.
For a broader approach to responsible AI in evaluation, use EvalCommunity Academy’s free Principles for AI Use in Evaluation and the practical AI Governance Toolkit for M&E professionals.
9. What matters more than detecting authorship?
In evaluation, the central question should not be limited to “Was AI involved?” The stronger questions concern evidence quality, ethics, transparency and accountability.
Review the work against these ten questions
- Are the findings supported by traceable evidence?
- Are quotations and statistics accurate?
- Can the analysis be reproduced or explained?
- Are uncertainty and limitations reported?
- Are alternative explanations considered?
- Were affected groups represented fairly?
- Was confidential information protected?
- Are recommendations connected to findings?
- Was AI assistance disclosed when required?
- Has an accountable evaluator approved the final product?
| Weak review question | Stronger evaluation-quality question |
|---|---|
| Does this sound like AI? | Which statements are unsupported, vague or inconsistent with the source data? |
| What percentage was generated? | What assistance was used, for which tasks and under whose review? |
| Can we ban all AI? | Which uses create unacceptable risk, and which can be governed responsibly? |
10. AI-use disclosure templates for evaluators
Disclosure should match the actual use. Do not claim that a document was fully human-written when AI materially supported drafting or analysis. Do not imply that AI made professional decisions that remained the evaluator’s responsibility.
Template A: language and structure support
Generative AI was used to support language editing and the initial organisation of selected sections. The evaluation team reviewed and approved the final wording. All findings, interpretations, references and recommendations remain the responsibility of the authors.
Template B: qualitative analysis support
An approved generative AI tool supported the preliminary organisation of de-identified qualitative material. Evaluators reviewed the proposed categories against the source records, revised the coding and completed the interpretation and triangulation. No directly identifiable participant data were entered into the tool.
Template C: no generative AI used
The authors did not use generative AI to draft, analyse or interpret the material in this report. Standard spelling, grammar, reference-management and document-formatting tools were used.
Important: Adapt disclosures to the organisation’s policy, the tool used, the sensitivity of the data and the actual level of human review.
11. Frequently asked questions
Can an AI detector prove that a report was written by ChatGPT?
No. A detector can indicate that text resembles patterns associated with AI-generated examples, but it cannot normally establish the complete authorship history of a document.
Should we run only the executive summary through a detector?
Treat the result with particular caution. Shorter samples provide less information and executive summaries are often highly edited, formal and structurally predictable.
Can a human-written report receive a high AI score?
Yes. False positives are possible, especially when the writing is highly standardised, translated, technical or different from the detector’s training data.
Does a low score prove that no AI was used?
No. AI-assisted text may be classified as human-written, particularly after substantive human revision.
What evidence is stronger than a detector score?
Draft history, source notes, tracked changes, interview records, calculation files, citation checks and the author’s ability to explain the analytical reasoning are usually more informative.
Should evaluators avoid AI completely?
The decision depends on the task, data sensitivity, contract and organisational policy. Responsible use requires appropriate tools, data protection, verification, transparency and accountable human oversight.
NEXT EVALCOMMUNITY TUTORIAL
Go beyond detector scores: improve AI-assisted M&E writing
AI detectors cannot tell you whether a report is accurate, specific or useful. Continue with the practical guide to removing vague claims, generic language, unsupported conclusions and repetitive AI-style wording from evaluation reports, proposals and donor communications.
12. Resources and further reading
EvalCommunity resources
- Principles for AI Use in Evaluation — a practical framework for ethical and accountable AI use in M&E.
- AI Governance Toolkit for Monitoring & Evaluation — tools for risk, fairness, governance and oversight.
- AI Tools in Monitoring and Evaluation — an overview of practical tools and use cases.
- EvalCommunity Academy — courses, tutorials and practical resources for evaluators and development professionals.
External guidance
- OpenAI: limitations of an AI-written text classifier — explains why the tool was withdrawn and includes cautions about false results, short passages and primary decision-making.
- UNESCO guidance for generative AI in education and research — a human-centred approach to policy, capacity and responsible use.
- NIST AI Risk Management Framework — a voluntary framework for managing AI risks and improving trustworthiness.
Final takeaway
Do not ask only, “Was AI involved?” Ask whether the work is accurate, ethical, transparent, traceable and accountable. A detector can support inquiry, but professional judgement must remain in charge.
CONTINUE YOUR PROFESSIONAL LEARNING
Learn to use AI responsibly across the full M&E cycle
The AI in Monitoring & Evaluation Certificate helps evaluators apply AI to evaluation design, evidence analysis, qualitative and quantitative work, reporting, privacy, bias checks and human review. The course is self-paced, requires no coding and includes a certificate of completion.
Working in humanitarian action or international development? Explore the AI for International Development and Humanitarian Practitioners Certificate.
