Avoid AI in Sensitive Assessment
Catch Me If You Can Series
AI can support administration — but humans must judge the work.
EvalCommunity Tutorial
Avoid AI in Sensitive Assessment
A practical guide for avoiding substantial AI use in peer review, proposal assessment, scoring, ranking, evaluation panels, and review of unpublished work.
Quick Answer
Substantial AI use should be avoided in peer review, proposal assessment, scoring, ranking, or evaluation of unpublished work.
AI may support limited administrative tasks, such as formatting templates, organising public criteria, or preparing blank checklists. It should not assess, score, rank, compare, recommend, or replace human expert judgment.
What You Will Learn
- What sensitive assessment means in research, evaluation, and proposal review.
- Why substantial AI use should be avoided in peer review and proposal assessment.
- Which AI uses are high risk in scoring, ranking, and evaluation panels.
- Which limited administrative AI uses may be lower risk.
- How to keep assessment human-led, fair, confidential, and accountable.
Main Sources and Related Tutorials
This tutorial is based on the ERA Living Guidelines on the responsible use of generative AI in research and related EvalCommunity guidance. The key principle is that AI may support process administration, but sensitive assessment should remain human-led.
Sensitive Assessment Safety Workflow
Use this workflow before using AI in peer review, proposal review, scoring, ranking, selection, or evaluation of unpublished work.
Step 1
Identify Task
Admin support or judgment?
Step 2
Check Sensitivity
Unpublished or confidential?
Step 3
Avoid AI Judgment
Do not score or rank.
Step 4
Protect Material
Do not expose submissions.
Step 5
Document Use
Record permitted support.
Step 6
Keep Human Decision
Humans make final calls.
1. Why Sensitive Assessment Matters
AI can support many administrative and editorial tasks. But not every task should be delegated to AI.
Peer review, proposal assessment, scoring, ranking, and evaluation of unpublished work involve judgment, fairness, confidentiality, and responsibility. Substantial AI use in these activities can create risks for applicants, researchers, evaluators, reviewers, organisations, and funders.
Core Rule
Substantial AI use should be avoided in peer review, proposal assessment, scoring, ranking, or evaluation of unpublished work.
2. What Is Sensitive Assessment?
Sensitive assessment is any process where a person, project, paper, proposal, application, organisation, or unpublished work is judged, scored, ranked, selected, rejected, approved, funded, or evaluated.
These decisions can affect funding, publication, careers, partnerships, reputations, access to opportunities, and organisational credibility.
Practical Test
Ask: is AI helping administer the process, or is AI judging the work? If AI is judging, scoring, ranking, comparing, or recommending, stop and keep the task human-led.
3. What Counts as Substantial AI Use?
Substantial AI use means AI is not only helping with formatting or administration. It is influencing the core judgment.
Assessment Judgment
- Assessing proposal quality
- Judging scientific merit
- Evaluating methodology
- Identifying strengths and weaknesses
Scoring and Ranking
- Scoring proposals
- Ranking applications
- Comparing applicants
- Recommending winners
Review Outputs
- Drafting peer review comments
- Writing panel conclusions
- Generating final recommendations
- Assessing unpublished work
4. What AI May Support Safely
AI may support limited administrative or non-judgmental tasks if confidentiality, data protection, and organisational rules allow it.
- Formatting a blank review template
- Creating a checklist from public criteria
- Explaining public assessment criteria
- Summarising public reviewer guidance
- Preparing a panel meeting agenda
- Drafting neutral administrative reminders
- Checking whether required sections are present
- Organising human-written notes without adding judgment
Safe Boundary
AI may help organise assessment materials. It should not assess the merit of the work.
5. What AI Should Not Do
- Do not ask AI to score, rank, compare, or recommend proposals.
- Do not ask AI to evaluate unpublished research or draft peer review judgments.
- Do not ask AI to judge innovation, feasibility, methodology, or scientific merit.
- Do not ask AI to produce final panel comments or funding recommendations.
- Do not upload confidential submissions into external AI tools without safeguards.
- Do not let AI become a hidden reviewer, scorer, judge, or panel member.
6. Why AI in Sensitive Assessment Is Risky
| Risk | Why It Matters |
|---|---|
| Unfair treatment | AI may favour certain writing styles, institutions, sectors, languages, or regions. |
| Hallucination | AI may invent strengths, weaknesses, claims, or assessment points. |
| Confidentiality breach | Unpublished proposals, papers, reports, or review materials may be exposed. |
| False objectivity | AI-generated scores may look neutral even when they reflect model limitations. |
| Loss of accountability | It may become unclear whether the reviewer or AI made the judgment. |
7. Decision Table: Can AI Be Used?
| AI Use | Classification | Safer Practice |
|---|---|---|
| Format a public assessment checklist | Usually acceptable | Verify criteria against the official source. |
| Summarise public reviewer guidance | Usually acceptable | Check accuracy manually. |
| Score a confidential proposal | Avoid | Human reviewers should score using official criteria. |
| Rank grant applications | Avoid | Use human panel review and documented reasoning. |
| Organise human-written reviewer notes | Use caution | Use approved tools and confirm no judgment is added. |
8. High-Risk Assessment Areas
Proposal Assessment
Do not use AI to score proposals, judge fundability, compare applicants, rank consortia, or draft final panel comments.
Peer Review
Do not use AI to review unpublished manuscripts, judge novelty, assess methods, or draft acceptance or rejection comments.
Evaluation Panels
Do not use AI to rank applications, score technical proposals, compare candidates, recommend bidders, or replace panel discussion.
9. Safer Prompt Formula
Formula
Administrative Task + Public Criteria + No Judgment + Human Review
Create a blank proposal review checklist based only on the public criteria below. Do not assess, score, rank, compare, or recommend any proposal. Do not add criteria that are not provided. The checklist will be reviewed by a human before use.
10. Safer Prompt Examples
Public Criteria Checklist
Turn public assessment criteria into a checklist. Do not evaluate any proposal or suggest scores.
Blank Matrix
Create a blank scoring matrix using only the criteria provided. Do not populate scores or comments.
Completeness Check
Check whether the form includes required sections. Do not assess quality, merit, or fundability.
11. Risky Prompts to Avoid
- Score this proposal.
- Rank these applications.
- Which proposal should be funded?
- Draft a peer review report.
- Identify weaknesses in this unpublished manuscript.
- Compare these applicants and select the best one.
- Evaluate the scientific merit of this project.
- Write final panel comments.
- Assess the quality of this unpublished report.
12. Good Practice for Reviewers, Panels, and Funders
| Reviewers and Panels | Funders and Organisations |
|---|---|
| Follow official assessment criteria. | Set clear rules on AI use in assessment. |
| Read original submissions directly. | Define prohibited and permitted AI uses. |
| Protect confidential and unpublished material. | Protect confidential submissions and review materials. |
| Write human reasoning and comments. | Avoid AI scoring or ranking of scientific content. |
13. Simple AI Use Notes
For Administrative Support
AI was used only to support administrative preparation, such as formatting a checklist or organising public criteria. It was not used to assess, score, rank, compare, or recommend any proposal, paper, applicant, or project.
For Review Panels
AI tools were not used to assess the scientific, technical, or evaluative content of submissions. All scoring, ranking, comments, and recommendations were made by human reviewers and panel members.
For Confidential Material
No confidential proposals, unpublished research, peer review materials, or protected assessment documents were uploaded into external AI tools.
14. Practical Exercise: Is This Sensitive Assessment Safe?
| Example | Classification |
|---|---|
| AI turns public criteria into a blank checklist. | Lower risk |
| AI scores a confidential research proposal. | Avoid |
| AI drafts peer review comments for an unpublished paper. | Avoid |
| AI formats an empty scoring matrix. | Lower risk |
| AI ranks grant applications. | Avoid |
15. Frequently Asked Questions
What is sensitive assessment?
Sensitive assessment is any process where people, proposals, projects, papers, applications, or organisations are judged, scored, ranked, selected, approved, rejected, or funded.
Can AI help with peer review?
AI should not substantially shape peer review judgments, comments, recommendations, or decisions. Limited administrative support may be possible if allowed by policy and if confidentiality is protected.
Can AI score proposals?
Substantial AI use for scoring proposals should be avoided. Scoring should be carried out by human reviewers using official criteria and documented reasoning.
Can AI create a blank scoring template?
Yes. This is usually lower risk if it uses public criteria and does not assess, score, rank, compare, or recommend any submission.
What is the safest rule?
AI may support administration, but it should not judge the work.
16. Final Takeaway
AI Can Support the Process. Humans Must Make the Assessment.
AI can help organise review processes, format checklists, and summarise public guidance. But it should not become a hidden reviewer, scorer, judge, or panel member.
Keep assessment human-led. Protect unpublished work. Protect fairness. Protect trust.
