Using NotebookLM and Claude for MERL
MERL professionals routinely handle large volumes of information—evaluation reports, interview transcripts, survey data, project documentation, and donor reports. Reviewing, synthesizing, and transforming these materials into actionable knowledge is often the most time‑consuming part of the job.
Artificial Intelligence tools like NotebookLM (for source‑grounded analysis) and Claude (for fluent drafting) can dramatically accelerate this process. Used together, they create a powerful analysis‑to‑output pipeline that maintains the rigor and professional judgment required in MERL.
Why Use NotebookLM and Claude Together?
NotebookLM specializes in document‑grounded question answering. It only responds based on uploaded sources, making it reliable for evidence extraction, cross‑document comparison, and theme identification without hallucinations.
Claude excels at structured writing and adaptation. It can take a bullet‑point evidence brief and turn it into a narrative report, executive summary, or donor update, while respecting tone, audience, and length constraints.
Why not use just one? NotebookLM lacks sophisticated writing capabilities; Claude lacks source‑grounded verification. Combining them gives you both accuracy and fluency—a combination essential for credible MERL work.
Theoretical Foundation: AI‑Augmented Sensemaking
This workflow is rooted in sensemaking theory (Weick, 1995) and the knowledge creation spiral (Nonaka & Takeuchi, 1995). In MERL, we move through a chain:
AI tools accelerate the middle stages:
- NotebookLM performs externalization – extracting and connecting implicit patterns across documents, turning raw data into structured information and evidence.
- Claude performs combination – reconfiguring extracted evidence into new knowledge products (reports, briefs, summaries).
- Human judgment ensures internalization – validating that the output is contextually sound, actionable, and aligned with organizational learning goals.
Key insight: This is not automation—it is augmentation. The evaluator remains the critical thinker, responsible for framing questions, interpreting nuance, and making final judgments. AI handles the heavy lifting of extraction and drafting, freeing time for higher‑order analysis.
The 6‑Step Practical Workflow
→
2. Upload & Explore
→
3. Extract Evidence
→
4. Structure Brief
→
5. Draft with Claude
→
6. Validate & Finalize
Iterative loop: each step may feed back to previous ones as new insights emerge.
1. Curate Your Source Documents
Quality in = quality out. Gather all materials relevant to the specific evaluation or research question. Organize them by type (quantitative reports, qualitative transcripts, secondary literature) and chronology (baseline, mid‑term, endline).
Pro tip: Create a simple inventory table with columns: Document Name, Type, Date, Key Topics. This helps NotebookLM produce more focused answers and allows you to quickly verify sources later.
2. Upload and Explore with NotebookLM
Upload your curated collection to NotebookLM. Resist the urge to ask generic questions. Instead, use iterative probing:
- Start broad: “What are the main themes across all documents?”
- Then narrow: “What evidence supports the effectiveness of intervention X?”
- Cross‑cut: “How do the endline findings differ from the mid‑term review?”
- Gap analysis: “What outcomes have limited or no evidence?”
As you get answers, follow up with clarifying questions. NotebookLM cites sources, so you can always verify. This is a dialogue, not a one‑off query.
3. Extract and Organize Evidence
Create a master list of findings, evidence statements, and contradictions. For each finding, note:
- Claim – what does the evidence say?
- Source – which document(s) support this?
- Strength – is it a single mention, or triangulated across sources?
- Limitations – any caveats or missing data?
This structured extraction is the critical bridge between analysis and writing. It transforms raw AI outputs into a curated evidence base.
4. Build a Comprehensive Evidence Brief
Now, structure your evidence into a brief that tells a story—even if it is still in bullet form. A good evidence brief includes:
- Context – evaluation questions, project background, theory of change.
- Key findings – organized by outcome or theme, with supporting data.
- Supporting evidence – specific quotes, numbers, or case examples from sources.
- Unexpected findings – surprises, contradictions, or negative results.
- Gaps – what we do not know and why it matters.
- Emerging lessons – patterns that might inform future programming.
Remember: This brief is the input to Claude. The more structured and clear it is, the better the output. Treat it as the blueprint for your final product.
5. Draft with Claude – Strategic Prompting
Copy your evidence brief into Claude and provide a detailed prompt that sets the stage:
- Role – “You are a senior MERL consultant with expertise in [sector].”
- Task – “Draft an evaluation findings section for a donor report.”
- Audience – “The reader is a program officer familiar with the project but not the details.”
- Tone – “Professional, evidence‑based, concise, and actionable.”
- Structure – “Organize by outcome, with sub‑headings for each finding. Include a brief interpretation after each finding.”
Claude can also rewrite for different audiences (e.g., a 2‑page executive summary vs. a full technical report) and suggest visualizations or key messages based on the evidence. You can also ask it to generate an outline first, then draft section by section.
6. Validate, Refine, and Finalize
This is the most critical step. AI is a drafting assistant, not a final authority. Your validation checklist:
- Does each claim trace back to a source in the evidence brief?
- Are there any factual errors, misinterpretations, or overstatements?
- Is the tone and framing appropriate for the intended audience?
- Are the conclusions and recommendations logically supported by the evidence?
- Have you added context, nuance, or caveats that Claude could not know?
Iterate: You can go back to NotebookLM to verify a specific point, then ask Claude to revise that section. This back‑and‑forth refines the output until it meets your standards.
Advanced Prompt Library for MERL
Beyond the basics, here are prompts for specific MERL tasks:
For NotebookLM (evidence extraction):
"Compare the recommendations in the endline report with those from the mid‑term review. Which recommendations were acted upon, and which were not? Provide evidence and cite sources."
For Claude (data interpretation and discussion):
"Here are the key findings from our evaluation. Draft a 'Discussion' section that interprets these findings in the context of the theory of change, identifies plausible causal pathways, acknowledges alternative explanations, and highlights implications for future programming."
For Claude (knowledge product adaptation):
"Using the evidence brief below, develop a 2‑page learning brief for program staff. Focus on actionable lessons, not just findings. Use a narrative style with one main message per section. Include a 'So What?' box for each lesson."
For Claude (executive summary):
"Based on the attached evidence brief, write a 1‑page executive summary for a donor. Highlight the most significant findings, the key recommendation, and the sustainability implications. Use clear, non‑technical language."
Scenario: Endline Evaluation Synthesis
Context: You have 15 documents: an endline survey report, 20 interview transcripts, 4 focus group summaries, a project logframe, and 5 quarterly reports. The evaluation covers a 3‑year health program in a rural region.
- Upload all to NotebookLM.
- Ask: “What is the overall change in Outcome 1 (maternal health access) since baseline?” → NotebookLM pulls data from the survey report and transcripts.
- Ask: “What contextual factors influenced implementation?” → NotebookLM finds patterns across quarterly reports and interview data (e.g., seasonal road access, community health worker turnover).
- Ask: “Are there any contradictions between the quantitative survey data and the qualitative interviews?” → NotebookLM highlights discrepancies that require deeper investigation.
- Create an evidence brief with findings organized by outcome, plus a section on barriers and enablers, and a separate section on data quality limitations.
- Prompt Claude: “Draft the ‘Findings’ and ‘Discussion’ sections of a final evaluation report. Use the evidence brief. The audience is a donor who needs to understand impact, sustainability, and value for money. Include a brief subsection on cost‑effectiveness if the data allows.”
- Review: Check every claim against the original sources. Add a paragraph on the program’s unique political and cultural context that Claude could not know. Revise the discussion to reflect your professional interpretation of the mixed findings.
Result: A high‑quality draft produced in hours instead of weeks, with full traceability, professional insight, and a clear audit trail from raw data to final narrative.
Ethical and Practical Guardrails
- Data privacy and security: Never upload personally identifiable information (PII), proprietary donor data, or sensitive community‑level data without explicit permission. Use anonymized transcripts and aggregated data. Follow your organization’s data protection policies.
- Bias awareness: AI can amplify biases present in the source material. Actively look for underrepresented perspectives, contradictory evidence, and alternative explanations. Use NotebookLM to specifically ask: “What evidence contradicts the dominant narrative?”
- Transparency and attribution: In final reports, consider a footnote or disclosure: “This report was drafted with the assistance of AI tools, with all content reviewed and validated by the evaluation team.” This builds trust and transparency.
- Iterative use across the evaluation cycle: Do not treat AI as a one‑shot tool. Use it throughout the evaluation lifecycle—from scoping (reviewing existing literature) to data collection (summarizing transcripts) to analysis and dissemination.
- Human‑in‑the‑loop: Always maintain a human reviewer who understands the program context, the evaluation questions, and the stakeholders. AI is a tool, not a replacement for professional expertise.
Common Pitfalls and How to Avoid Them
Pitfall 1: Treating AI outputs as final.
Solution: Always validate against source documents. Use NotebookLM’s citation feature to trace every claim. Build a validation step into your workflow, not as an afterthought.
Pitfall 2: Over‑reliance on a single query.
Solution: Use NotebookLM iteratively. Start broad, then drill down. Ask follow‑up questions. Treat it as a conversation with a research assistant, not a search engine.
Pitfall 3: Ignoring contradictory evidence.
Solution: Explicitly ask NotebookLM: “What findings contradict the main trends?” and “What evidence is weak or inconclusive?” Use Claude to draft sections that honestly discuss limitations and alternative interpretations.
Pitfall 4: Vague prompting.
Solution: Be specific with Claude. Provide role, task, audience, tone, and structure. The more context you give, the better the output. Use the prompt library above as a starting point and adapt it to your specific needs.
Conclusion
NotebookLM and Claude support different but complementary stages of the MERL workflow. NotebookLM assists with document review, evidence extraction, and synthesis, while Claude helps transform structured information into clear and organized written outputs. When combined with professional expertise, critical thinking, and quality assurance processes, these tools can support more efficient evidence review, reporting, and organizational learning.
The future of MERL is not AI replacing evaluators—it is evaluators using AI to work at a higher level of insight and impact. This workflow is a practical step toward that future. Start small, iterate, and always keep the human at the center.
AI‑augmented MERL is about working smarter—not replacing expertise, but amplifying it. Always validate, always contextualize, always apply professional judgment.
Tutorial for MERL Professions
