Signpost AI Information Assistant
- Categories AI, Case Studies
- Date April 22, 2026
Signpost AI Information Assistant: A Pilot Case Study on Responsible AI in Humanitarian Response
Signpost AI — a consortium of the International Rescue Committee (IRC), Mercy Corps, Internews, and local partners — developed and piloted a Retrieval-Augmented Generation (RAG) based AI Information Assistant to help frontline responders scale one-to-one information delivery. Over six months across Greece, Italy, Kenya, and El Salvador, the tool reduced response time by up to 80%, improved moderator efficiency, and demonstrated safe usage only with a human-in-the-loop (HITL).
The Humanitarian Information Gap: Demand Outstrips Capacity
Signpost — the world's first scalable community-driven information program — is jointly implemented by the International Rescue Committee (IRC), Mercy Corps, Internews, and a network of local partners. Since 2015, Signpost has launched programs in over 30 countries, reaching approximately 20 million users and delivering 500,000+ one-to-one information counselling sessions. Yet a fundamental challenge persists: individualized information support remains tied to human resources, limiting scalability.
The critical bottleneck:
In late 2023, a single social media campaign in Afghanistan generated 48,000 inquiries within weeks, overwhelming a team of three staff and forcing them to halt two-way counseling. Demand for empathetic, localized, accurate information far exceeds humanitarian capacity — a gap that AI, if used responsibly, could help bridge.
Signpost AI (SPAI) — a dedicated initiative within the consortium — launched in early 2024 with a mandate to responsibly build, derisk, pilot, and scale AI tools that augment frontline responders, never replace human judgment.
The AI Solution: A Retrieval-Augmented Generation (RAG) Assistant
The consortium developed an LLM‑based Information Assistant designed to help moderators respond to user queries using vetted, localized knowledge. The RAG architecture connects a Large Language Model to a vector database containing ~30,000 fact-checked Signpost articles, service maps, and country-specific content. Workflow:
1. User query (FB, WhatsApp, SMS) → 2. AI classifies intent → 3. Retrieves relevant articles from vector DB → 4. Concatenates with country prompts → 5. LLM generates response → 6. Human moderator reviews and edits → 7. Safe output delivered.
Each pilot country (Greece, Italy, Kenya, El Salvador) had a localized knowledge base and custom system prompts. The tool integrated into Zendesk (CRM), enabling moderators to generate, score, and refine AI-drafted replies seamlessly.
Grounding Humanitarian & AI Principles: The Consortium's Ethical Compass
Through cross-consortium consultations (IRC, Mercy Corps, Internews, local partners), the team developed four core principles:
Safety, do-no-harm, human-centered at every stage.
Full documentation, AI Impact Assessments, open case studies.
Rigorous metrics, sandbox testing, and red teaming.
Partnerships with local organizations, academia, and tech.
Design, Sandbox Testing & Red Teaming
Signpost AI assembled an interdisciplinary team across consortium partners: product leads, developers, M&E officers, protection officers (from IRC and Mercy Corps), and a dedicated Red Team. The Red Team adversarially tested the assistant for bias, hallucinations, security leaks, and unsafe outputs using metrics such as Harmful Information Flag, Bias Flag, Hallucination Rate, and Response Time.
📌 Quality framework for outputs (scored by moderators):
- Trauma-Informed (1-3): Psychological first aid language, tone matching
- Client-Centered (1-3): Individualized, clear, accessible, prioritizes user concern
- Safety/Do No Harm (Yes/No): No stereotypes, political statements, confidentiality
After months of sandbox iteration (testing Claude, GPT-4o, Gemini), the assistant reached acceptable safety levels for a 6‑month live pilot with human moderators in four countries.
Pilot Results: Performance, Trust & Efficiency Gains
Country highlights: Greece (885 inquiries, 62% good), Italy (323, 67% good), Kenya (719, 311 rated good/usable), El Salvador (241, 70% good). All moderators across consortium partners requested to keep the tool post-pilot. Trust increased significantly over time, though the team noted occasional automation complacency — reinforcing the need for ongoing AI literacy.
⭐ Critical finding: The assistant performed well on Safety (85% safe) and Trauma-Informed (79%), but Client-Centeredness (75%) was the main driver of failures. Context-rich knowledge bases are the single most critical factor for quality — a lesson for all consortium partners scaling AI.
Implications for Monitoring & Evaluation (M&E) in Humanitarian Settings
Human-in-the-Loop (HITL) is non-negotiable
With 23% of outputs requiring correction to avoid potential harm, the tool is unsafe for autonomous use in high-stakes contexts. M&E systems must integrate continuous human review metrics.
Knowledge base completeness = performance driver
M&E frameworks should measure knowledge base relevance and coverage. Greece and Italy (mature KBs) outperformed contexts with thinner content — directly informing resource allocation.
AI literacy and trust require longitudinal tracking
Pre/post surveys showed increased AI knowledge and productivity, but also unrealistic expectations. M&E should track both performance gains and cognitive biases (over-trust, automation complacency).
Key Takeaways for Humanitarian & M&E Professionals
Build, test, and de-risk through prototyping and red teaming.
No fully autonomous client-facing AI in humanitarian information delivery – yet.
Invest in high-quality, localized, regularly updated knowledge bases.
IRC, Mercy Corps, Internews and local partners shared M&E insights, reducing duplication and risk.
Frequently Asked Questions
Who comprises the Signpost AI consortium?
Signpost AI is a consortium led by the International Rescue Committee (IRC), in strategic partnership with Mercy Corps, Internews, and a network of local implementing partners across 30+ countries. This collaborative structure ensures humanitarian expertise, local contextual knowledge, and responsible AI governance.
Is the Signpost AI assistant safe to use autonomously?
No. The pilot demonstrated that 23% of outputs required human correction to avoid potential harm. The assistant is safe and effective only with a human-in-the-loop (HITL) — moderators review, edit, and approve every response before it reaches a client.
What LLMs were tested during the pilot?
The team tested Claude (Anthropic), GPT-4o (OpenAI), and Gemini (Google). Claude was noted for its "emotional intelligence" and trauma-informed language, and was selected for the final pilot based on safety and performance metrics.
How can other humanitarian organizations adopt this approach?
Signpost AI has open-sourced its case study, quality framework, red team metrics, and RAG architecture documentation. Visit signpostai.org/research for toolkits, training materials, and implementation guidance.
A Blueprint for Responsible AI in Humanitarian Information Delivery
The Signpost AI pilot demonstrates that generative AI, when grounded in humanitarian principles, tested through red teaming, and deployed with human-in-the-loop oversight, can safely enhance the scale and quality of information services. The consortium model — bringing together IRC, Mercy Corps, Internews, and local partners — proved essential for contextualizing AI, sharing M&E insights, and building trust.
For Monitoring & Evaluation professionals, this case provides actionable metrics: pass rates, hallucination tracking, client-centered scores, and trust indices. The future of AI in humanitarian action is not autonomous systems — it's augmented human responders, equipped with tools that respect dignity, ensure safety, and expand access to life-saving information.
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
