How to Use the Gender-Responsive AI Safety Evaluation Framework
What is this tool?
The Gender-Responsive AI Safety Evaluation Framework is a free, browser-based interactive tool built by EvalCommunity for M&E professionals, researchers, and AI auditors. It provides a structured, participatory methodology to evaluate AI systems for gendered risks — moving beyond generic technical benchmarks.
Core premise: an AI system cannot be considered “safe” if it produces gendered harms, regardless of benchmark performance. The tool helps you identify, test, document, and escalate those harms — with community co-evaluators at the centre.
Who this tutorial is for
The 7 tabs at a glance
Work through them in order for your first evaluation, then return to individual tabs as needed.
Step-by-step guide
Each step corresponds to one or more tabs in the tool. Follow in order for your first evaluation.
Orient yourself to the framework
Start here if you are new to gender-responsive evaluation. This tab sets the criteria your evaluation will be held to.
What to do on this tab:
aRead the 4 foundational principles. These are the lens through which all evaluation steps are interpreted. Principle 1 (intersectionality) shapes your community mapping in Phase 1.
bUse the interactive phase timeline. Click any of the 4 coloured circles to jump to that phase in the Eval Cycle tab — fastest navigation during an active evaluation.
cNote the core premise banner. Screenshot or print it for briefing colleagues or presenting evaluation rationale to programme leadership.
Work through the 4-phase evaluation cycle
The operational core of the tool. Maps the full evaluation from pre-scoping to governance, with specific actions at each stage.
Phase 1 · Pre-Evaluation Scoping
Four action cards: community mapping (intersectional identity matrix), contextual harm taxonomy, community-led safety goals, and trauma-informed protocols.
Phase 2 · Evaluation Execution
Four methods: Participatory Test-a-thons, Contextual Red-Teaming, Community-Based Audits, Safety Sandbox Testing — with application guidance for each.
Phase 3 · Documentation & Remediation
Gender-specific transparency requirements, remediation pathways for detected harms, and how to feed findings into incident databases and practitioner networks.
Phase 4 · Capacity & Governance
Build institutional capacity for sustained evaluation — mandatory gender screenings for high-risk sectors and contribution to public leaderboards.
Classify gendered harms before testing
Understanding harm typology before testing determines which prompts you design, what outputs you flag, and which remediation you recommend.
aReview the intersectional harm risk matrix. The table maps risk levels across 5 identity-context combinations. Use it to prioritise which harms to test first.
bClassify harms before filling in the Audit tab. Each checklist item maps to one of these three typologies — knowing which type helps you assess it accurately.
Select your evaluation methods and test cases
The bridge between principles and the checklist — method framework plus specific test scenarios to run.
How to use the Methods table:
Complete the 15-item audit checklist
The scoring step. Rate each criterion, watch your score update live, and export your full results.
How scoring works:
aUse the dropdown for each of the 15 criteria. Dropdown colour updates instantly (green / amber / red). Score ring and count badges update in real time.
bRead the score threshold label. Below 40% = “Critical gaps.” 40–70% = “Partial compliance.” Above 70% = “Strong gender safety posture.”
cClick Export to copy the full audit to clipboard. Includes all 15 criteria with scores, final percentage, and source citation. Paste directly into reports or procurement dossiers.
Generate your Gender Safety Disclosure Card
A structured, publishable evaluation record modelled on AI Model Card standards — for transparency, procurement, and regulatory purposes.
The 10 fields explained:
1Fill in all 10 fields, then click Generate Card. A formatted disclosure table appears below the form.
2Click Copy to copy to clipboard. Use Clear to reset for a new evaluation.
Submit results and contribute to the public dataset
Your submission contributes to a growing, community-generated dataset of AI gender safety performance.
1Enter your system name, sector, and both scores. Click Add to Leaderboard — ranked automatically against benchmark data.
2Average ≥75 = Compliant · 50–74 = Partial · below 50 = At Risk. Use your ranking to benchmark against comparable deployments.
Frequently asked questions
Open the Gender-Responsive
AI Safety Evaluation Tool
Free, browser-based, no login required. Desktop and mobile. For M&E practitioners across 90+ countries.
