AI Research Methods for Nonprofits
Teaches nonprofit staff to evaluate AI tools and claims rigorously on a constrained budget, rather than to build models.
- Who it's for
- Program leads, M&E staff, development/fundraising teams, and operations managers at nonprofits and NGOs
- Format
- 2 days, instructor-led, in-person or remote
- Prerequisites
- No coding required. Comfort with spreadsheets and reading a basic chart is enough.

By the end
What your team walks out with.
- Read an AI vendor claim and identify precisely what evidence would need to exist for it to be true
- Design a small, honest pilot that produces a defensible yes/no decision inside one grant cycle
- Build an evaluation set from your own historical data instead of relying on public benchmarks
- Write an AI section for a grant report or board memo that neither overclaims nor hides real results
- Recognise where donor, beneficiary, or case data makes a given AI tool inappropriate regardless of performance
4 modules
How the programme runs.
01What "AI research" means when you are not a lab
Half dayA shared vocabulary and a clear split between questions you can answer in-house and questions you must outsource.
02Reading claims: benchmarks, demos, and what they hide
Half dayParticipants dissect three real vendor pitches and identify the missing evidence in each.
03Building an evaluation set from your own data
Half dayA working 50-example evaluation set drawn from the organisation's actual documents and cases.
04Running a pilot that survives a board review
Half dayA written pilot protocol with a baseline, a success threshold, a stop rule, and a named owner.
Nonprofits are being sold AI faster than almost any other sector, and with less capacity to check what they are buying. Vendors arrive with discounted nonprofit pricing, a compelling demo, and a claim about hours saved. The staff evaluating that claim are usually programme people with no budget for a technical consultant and no time to run a proper trial.
This curriculum is the response to that gap. It is not a course on building AI, and it is not a tools tour. It teaches a nonprofit team to do research — to ask what evidence would make a claim true, to generate that evidence cheaply from their own data, and to reach a decision they can defend to a funder.
Why this curriculum exists
Sector surveys have consistently found that a large minority of nonprofits are already experimenting with AI, concentrated in fundraising and communications, and that adoption is growing considerably faster than evaluation capacity. That mismatch is the actual risk. An organisation that adopts a tool which quietly performs badly on its specific caseload does not usually find out from a benchmark — it finds out from a mis-triaged case or an embarrassing donor communication, months later.
The core argument of this curriculum is that evaluation is cheaper than most nonprofits assume, and much cheaper than a failed rollout. A useful evaluation set is fifty examples, not fifty thousand. A useful pilot is one programme area for six weeks, not an organisation-wide transformation. The skills required are closer to monitoring and evaluation practice — which most nonprofits already have — than to machine learning.
Who this is for
This works best for a mixed group of six to twenty people spanning programme delivery, monitoring and evaluation, fundraising, and operations. Mixing functions matters: the fundraising team's enthusiasm for a donor-scoring tool and the safeguarding lead's concerns about case data need to meet in the same room, working from the same evidence, rather than negotiating by email after a purchase decision.
It is a poor fit for organisations looking for a hands-on prompt-writing workshop, or for teams that have already committed to a specific vendor and want validation rather than evaluation.
What the two days actually look like
Day one — reading claims
The morning establishes what can and cannot be known about an AI system from the outside. Participants learn the difference between a benchmark score, a demo, a case study, and a controlled comparison, and why the first three are frequently compatible with a tool that will not work for them.
The afternoon is adversarial and hands-on. Groups take three real vendor pitches — anonymised, drawn from actual nonprofit-sector sales material — and for each one write down the single piece of evidence that would most change their confidence. This is usually the moment the framing clicks: the question stops being "is this tool good" and becomes "good at what, measured how, on whose data."
Day two — generating your own evidence
The morning builds an evaluation set. Participants bring anonymised historical examples from their own work — past grant applications, case summaries, donor queries, whatever the candidate tool claims to help with — and assemble roughly fifty examples with known correct answers. This artefact is the durable output of the whole programme. It outlives any particular tool, and it can be reused every time a new vendor appears.
The afternoon turns that into a pilot protocol: a baseline measurement of current performance, an explicit success threshold agreed before the tool is run, a stop rule, a duration, and a named owner. Teams leave with this written down and signed off in the room, which is the single highest-value artefact of the two days.
What we deliberately do not cover
- Model building or fine-tuning. Almost no nonprofit should be training models, and a curriculum that implies otherwise is selling the wrong ambition.
- Tool-specific instruction. Tools change faster than a curriculum can track; evaluation methods do not.
- Procurement and contract law. Important, but a different specialism — we cover the technical evidence that should inform procurement, not the contracting itself.
Governance and data constraints
Data classification is treated as a gate, not an afterthought. Before any evaluation runs, teams work through which categories of their data can leave the organisation at all: donor personal data, beneficiary case notes, safeguarding records, and financial information each carry different constraints, and in some jurisdictions and some programme areas the honest answer is that a cloud-hosted model is not an appropriate destination regardless of how well it performs.
Teams that reach that conclusion have not wasted the exercise. Establishing clearly that a category of work is off-limits, and documenting why, is a legitimate and valuable outcome — and considerably cheaper to reach in a workshop than after a data incident.
Related curricula
AI safety and guardrails for healthcare teams covers a similar evaluation discipline in a higher-stakes setting, and ChatGPT for work is the better starting point for teams who need day-to-day usage skills first.
Related programmes and reading
This curriculum pairs naturally with broader evaluation and governance work. For background reading before a session, participants often find these useful:
- How to read AI benchmarks — the single most relevant prerequisite article for day one
- What is an AI benchmark, and what does it actually measure?
- Responsible AI and governance for teams
- AI interpretability and monitoring for teams
Sessions are delivered by explainx.ai and can be scoped to a single programme area or run across an entire organisation. Curricula are adapted to the sector, data constraints, and jurisdiction of the specific organisation rather than delivered from a fixed deck.
Common questions
- Do participants need a technical or data science background?
- No. This curriculum is built for programme and operations staff, not engineers. Everything is done in spreadsheets and plain documents. The goal is to make people confident evaluators of AI claims, not to teach them to train models.
- How is this different from a general "AI for nonprofits" workshop?
- Most introductory sessions teach tool usage — how to prompt a chatbot, how to draft a donor email. This curriculum assumes your team can already do that and addresses the harder question: how do you know whether a given AI tool actually works for your specific caseload, and how do you prove it to a funder or a board?
- Can this run for a small team with no budget for paid tools?
- Yes. The evaluation methods taught here work on free tiers and are specifically designed for constrained budgets — the entire point is to avoid spending scarce restricted funds on tools that will not survive contact with your real data.
- Does it cover donor data privacy and safeguarding?
- Yes, as a gating constraint rather than a separate compliance module. Module four covers how to determine which categories of data — beneficiary case notes, donor records, safeguarding reports — can legitimately be sent to a third-party model, and what to do when the answer is none of them.
- How long before a team sees a usable result?
- Teams typically leave with a completed evaluation set and a written pilot protocol on day two. A first defensible go/no-go decision on one tool usually takes three to six weeks after that, depending on how much historical data the organisation can assemble.
Make it fit your team
Shape this curriculum around your work.
Every session is adapted before delivery — to your tools, your data constraints, and the tasks your team actually does. Tell us the context and we will come back with a scoped outline.
Our practitioners’ training experience




Platforms our practitioners teach on
Udemy
Coursera
CodecademyOther curricula
Agent Harness Engineering
Teaches teams to evaluate and operate agent harnesses as production infrastructure, covering tool design, permissions, sandboxing, and observability.
AI Safety and Guardrails for Healthcare Teams
Teaches healthcare teams to find and contain AI failure modes before deployment, with escalation paths designed around clinical risk rather than model accuracy.
ChatGPT for Work
Teaches teams to convert ad-hoc ChatGPT use into shared, reviewable workflows using custom GPTs, projects, and data analysis.
Claude for Work
Teaches non-engineering teams to use Claude for repeatable work — projects, long documents, and shared workflows — rather than one-off chat prompts.