Safe Generative AI Chatbots: Mercy Corps
- Categories AI, Case Studies
- Date May 8, 2026
Safe Generative AI Chatbots: Mercy Corps
Technology: Generative AI · RAG · LLMs
Focus: MEL · Knowledge Management · Data Protection
1. Introduction: Balancing AI Opportunities and Risks
Generative AI tools such as ChatGPT are freely available and easy to use. Their use creates opportunities but also significant risks. Most organizations have put policies in place to provide guardrails and guidance for individual staff members using generative AI tools, but in reality, compliance is difficult to track and enforce. At the same time, generative AI is becoming more prevalent both as standalone tools and built into existing software applications. To balance governance efforts with the need for staff to become familiar with generative AI, Mercy Corps decided to build two in-house chatbots — one for the Monitoring, Evaluation, and Learning (MEL) team and one for organization-wide knowledge access.
2. Background: A Safe Alternative to Public Chatbots
Mercy Corps approached the project as an opportunity for the organization to learn about generative AI — how to use it safely and responsibly, how to build and manage generative AI-based solutions, and how to assess its benefits. Realizing there was no way to avoid or prevent the use of generative AI tools, the organization wanted to provide a safe alternative to using freely available tools, while providing a commensurate user experience. “Safe” in this context meant:
- Alignment with Mercy Corps’ responsible data policy
- Reduced risk of data leakage from accidentally incorporating personally identifiable information (PII) into public chatbots
- Parameters set to limit or mitigate bias and hallucinations by restricting source material primarily to the Mercy Corps Digital Library
3. AI Development: Two In-House Chatbots
MEL Chatbot
Built at the request of Mercy Corps’ Monitoring, Evaluation and Learning (MEL) team. Serves a small team wanting to interact with a custom set of documents specific to their area. Synthesizing large volumes of documents to identify relationships and generate new insights had been a long-standing challenge for MEL teams — this chatbot helps address that.
Digital Library Chatbot
Aimed at the entire organization, serving as a “front-end” to Mercy Corps’ Digital Library, which contains over 20,000 documents. Users can ask questions in natural language covering practical topics (e.g., how to file expenses) and strategic topics (program design). The chatbot can respond in the user’s own language, enhancing accessibility.
Technical Architecture
For both chatbots, Mercy Corps leverages off-the-shelf large language models (LLMs) in their own Microsoft Azure environment. Retrieval Augmented Generation (RAG) has been activated so the chatbots can, when needed, also access public resources on the internet — ensuring responses are grounded in authoritative sources.
4. Operationalizing AI: From Pilot to Organization-Wide Tool
User feedback on the MEL chatbot was very positive. Users reported between 40% and 60% increase in efficiency — meaning time saved reading and processing multiple documents and generating text summaries. Additionally, users reported that they were discovering new insights in the summaries provided by the chatbot, going beyond simple time savings to enhanced analytical capability.
The Digital Library chatbot is currently being tested by forty staff members from across Mercy Corps’ countries and functions, evaluating whether responses are accurate and usable. The project team is conducting adversarial testing or “red teaming” — trying to see to what extent they can make the chatbot provide wrong or inappropriate responses — enabling them to correct and fine-tune the model. Users can also see which documents the chatbot uses to generate output, helping to validate results and make corrections where needed.
As part of the overall project, Mercy Corps’ Data Protection and Privacy team has launched an “Ethical AI” workstream, ensuring integration of personal data protection principles into any AI use. The team has conducted research on emerging AI regulations and distilled recommendations, merging them with Mercy Corps’ humanitarian principles to create processes such as ethical AI assessments.
5. Impact on International Development, Humanitarian Action, and M&E
Efficiency Gains for MEL Teams
40-60% time savings in document synthesis and summarization allows M&E professionals to focus on higher-value tasks: validation, analysis, and strategic recommendations rather than manual data extraction.
New Insights Discovery
Beyond efficiency, chatbots help identify relationships across documents that humans might miss — enhancing the quality of evaluations, literature reviews, and evidence synthesis.
Democratized Knowledge Access
Natural language querying and multilingual support make 20,000+ organizational documents accessible to field staff, breaking down barriers related to language, search literacy, and time constraints.
Responsible AI Governance
The Ethical AI workstream and standardized assessment processes provide a replicable model for other humanitarian organizations to adopt safe, accountable generative AI tools.
6. Key Learnings: Technical Surprises, Expertise, and Sustainability
Technical Performance
“I would say technically I’ve been surprised at how well these work,” says Shadrock Roberts. While fine-tuning and system prompting are still needed, the overall accuracy of LLMs trained on organizational documents is seen as good.
Cross-Functional Collaboration
Having domain experts (e.g., MEL specialists) work closely with technology experts is essential to get chatbot output right. Technical accuracy alone is insufficient without contextual relevance.
Policies and Tools Together
Developing policies and tools for safe and responsible AI use alongside the technology solution reassures users and creates a standardized evaluation approach for future AI requests.
Retaining AI Expertise
Organizations must think about how to retain AI expertise given high market demand — whether through in-house capacity building or contracting with service providers. A solid capacity foundation is needed to scale AI across an organization.
7. Future Plans: Broad Rollout and Programmatic Use
Mercy Corps is planning to roll out the Digital Library chatbot broadly across the organization as a Beta version so users can start utilizing the solution and become familiar with how they can employ generative AI to assist in their work. Initially, English, French, and Spanish language interaction will be available. As the chatbot is run in-house on Mercy Corps’ documents, scaling up to more users will increase infrastructure cost over time, but with up to five thousand potential users, it will still be more cost-efficient than acquiring licenses for publicly available solutions.
With the rollout, the organization expects to gain a better understanding of how users interact with generative AI, what other use cases might emerge, and eventually even programmatic areas where tools might interact with individuals and communities that Mercy Corps serves.
8. Frequently Asked Questions (FAQ)
Where to Learn More
For inquiries about Mercy Corps’ generative AI initiatives or ethical AI framework:
Download the full case study to explore technical architecture, ethical AI assessment processes, and implementation lessons.
The courses and articles are developed by a team of experienced evaluators, collaborators, authors, and software developers, guided by Fation Luli. EvalCommunity Academy combines practical expertise in Monitoring & Evaluation and International Development with the latest advances in AI to create high-quality, accessible, and practical learning experiences for professionals worldwide.
