Case Study • Education.org
Education.org had already paid for its research, but reaching it took days, so most of it went unused. Their evidence decides where NGO funding and effort go, and it sat in scanned PDFs and languages the reader didn’t speak.
We changed the economics of an archive they already owned. We mapped their research workflow stage by stage and built for each one. The centrepiece made the whole of their research answerable in one question, for their team first and then for the public they serve.
Education.org came to us through Sallyann, founder of Gleac, whose community Ionio serves as its AI team. She brought us in and stayed close to the work.
“I try to make sure when I’m bringing in teams, when my internal team can’t handle it, that I bring in someone who can handle it from A to Z.”
“I had very little involvement. I think I attended three or four meetings, and you and your team handled that project.”
“I was comfortable already with you. I knew that you could manage that, and your team handled that project.”
“There was such depth in terms of the work. They’re an example of an NGO that really took the first step.”
Education.org had decided where it wanted to be: an AI-first organisation, one of the few NGOs in its field building AI into how it actually works. That was a positioning bet as much as a technical one. In a space where credibility and being early both matter, an NGO that can show it runs on AI stands out to the funders and partners it depends on, and it opens the door to more of the work it wants to do.
The trouble was the gap between wanting that and living it. The organisation ran on knowledge work, and while people had started reaching for AI here and there, none of it was regimented. There was no system, no shared tooling, nothing that turned scattered experiments into a way of working. Closing that gap is what the engagement was really about.
Everyone played it their own way, and nothing was keeping time.
Underneath the positioning sat the daily reality, and it is easiest to see through one researcher. Picture someone who needs to know what the organisation has already established about, say, girls’ schooling outcomes in a particular region. The answer almost certainly exists somewhere in the research. Getting to it is the hard part.
They would start searching, one repository at a time. Some of what they need is in languages they don’t read. Some is locked in scanned PDFs with no searchable text. Some was filed years ago by someone who has since moved on. So the search becomes a hunt, and the hunt takes days, and at the end of it they have a stack of documents they still have to read before they can answer the question they started with.
Days of searching outward, and the answer was at the centre.
That was one stage of a longer chain. Searching, then cataloguing, then screening each source against the inclusion criteria, then coding it, then pulling findings together, each step manual, each done in a different tool, each eating the time of skilled people. Every hour spent on it was an hour not spent on the judgment only a researcher can provide.
The organisation’s most valuable people were spending their days on work a machine could carry, and its research was going underused because reaching it cost too much.
We began in 2024, before the current generation of reasoning models existed, inside a specialist discipline neither of us could shortcut. Almost none of what made this difficult was code.
Evidence synthesis is a specialist craft with its own rules, and getting it wrong means the output can’t be trusted. We had to learn how their researchers actually work, what grey literature is and why it’s hard to source, how inclusion and exclusion criteria get applied, and what a librarian needs to see before they’ll rely on a tool for real work.
What this meantWe spent real time learning the discipline before we designed anything.
We began in 2024, before the current generation of reasoning models existed. One of the projects we took on, automating the judgment-heavy screening and coding work, ran into a wall the models of that moment couldn’t clear, so we held it at the pilot stage rather than ship something unreliable into a research organisation. When reasoning-focused models arrived not long after, that work became possible.
What this meantKnowing when to wait was as important as knowing what to build.
The source material was not clean text. Scanned PDFs with no readable layer, documents in languages the reader didn’t speak, formats that changed from one repository to the next. Before anything could be searched it had to be pulled in, read by OCR where needed, and translated into one common language, so that a single question could reach across the whole collection.
What this meantThe ingestion pipeline had to make messy documents behave like clean ones.
The people we were building for were deep domain experts who hadn’t yet had reason to work closely with AI. That changed the job. Part of the delivery was showing them what these tools could and couldn’t do, setting honest expectations, and designing something that felt trustworthy to someone encountering it for the first time.
What this meantWe were teaching as much as we were shipping.
A closed discipline, a technology that was not ready, and a room of experts who had to trust the thing before they would use it. The engineering was the tractable part.
Rolling out AI inside an organisation that is excellent at one thing is not mainly a technical problem. The people who have to trust the tool are often the ones with the least reason to have kept up with AI, because they were busy being expert at the actual work. Education.org was a room full of exactly that: sharp researchers who simply hadn’t had time to dig into what these models do or don’t do.
So the job was never just to automate a workflow. It was to bring a team of experts along at a pace they could trust, while building the thing that would earn that trust. In practice that meant three things:
The principle travels to any expert organisation adopting AI. Understand the work first, set honest expectations, and build for the people who have to adopt it.
The automation is the easy half. Getting a room of experts to rely on it is the half that decides whether the thing gets used at all.
That second half showed up in how we worked, not just what we built. We walked in assuming certain things were obvious, that everyone understood what these models could and couldn’t do. We were wrong. These were deeply smart researchers with real depth in their field who simply hadn’t had time to dig into AI, and teaching the technology’s limits was a prerequisite for the engineering to land, not a nice-to-have.
It shaped how we communicated too. We didn’t ask open-ended questions, we came with options, framing decisions as A, B or C rather than “what do you want,” explaining the tradeoffs, and making sure no one had to chase us for an update. That is how you keep a non-technical stakeholder in control of a deeply technical process, and over time it changes what they can do.
The same instinct applies once a product is live. When the tools underneath an AI system move on, the right response isn’t to wait for the client to notice. It is to come back and say plainly: here is what we built, here is what has changed, here is what it means for you. Framed as education rather than a sales pitch, that is what turns a delivered project into a lasting one.
“Here’s what we did before, here is what the new technology is right now. For a client that it is working, that’s actually better for them in the long run.”
From a plain-language question to a cited answer, then the library it draws from and the site where the public meets it. Here is the Ask Evidence platform in action.
The tool opens on a single prompt: ask a question about education and get a research-backed answer drawn from Education.org’s own verified resources. A “Popular Topics” row offers real starting points, from education policy for equity to the role of NGOs, so a first-time user isn’t staring at an empty box. One question, the whole evidence base behind it.
Ask something like what role NGOs play in education policy, and the answer comes back written out, with a Sources card naming the exact document it drew from and inline markers tying each claim to it. Underneath, a Related Questions list suggests where to go next. Every statement traces back to a real study, which is what makes it usable for research rather than just readable.
The Library holds every document the tool answers from, in one place. National education plans from Guinea-Bissau, Iraq, Malawi, Zimbabwe, Myanmar and more sit alongside research reports, each a real source the answers can cite. This is the corpus, kept in view and managed by the team rather than buried in a database.
The same tool lives on Education.org’s own site, reachable from an “Ask the Evidence” button on every page. A visitor reading about the organisation’s work can put a question to its research without leaving, which is how the evidence reaches the professionals it’s meant for.
We started by mapping Education.org’s whole research workflow, the chain that ran from searching for sources, through cataloguing and screening them, to coding and pulling findings together. Then we went stage by stage and built to replace the manual work at each one. That produced three projects: two focused tools that took on specific stages, and one larger platform that became the centrepiece. Ionio owned the full arc of each, product strategy, wireframes, UI and UX, the build, deployment and handover.
The workflow began where every piece of research does, with finding sources, and that was the first stage we automated. A researcher used to work through repositories one at a time, hunting across libraries, journal databases and open repositories, much of it grey literature that is awkward to find and awkward to read.
We built scrapers that went out to the sources that mattered, among them the World Bank and the African Education Database, and pulled the documents in automatically. Across the target sources this gathered more than a thousand documents into one place, turning days of manual desk searching into a process that ran on its own.
Days of desk searching, replaced by a process that runs on its own.
The next stage was the hardest, and the most revealing. Once sources are found, researchers screen each one for inclusion against a set of criteria, then code it, work that is slow, expert, and done entirely by hand. We built to automate it, and in doing so we found the real limit was not the tooling but the moment.
The coding framework leaned on subjective human judgment that the models of the day could not reproduce reliably, so we held this at the pilot stage rather than ship something a research team couldn’t trust. It became the clearest lesson of the engagement: sometimes the right build is the one you wait to finish.
A research team cannot use an answer that changes every time you ask.
The third project grew into the main one. Rather than speed up one stage, it changed the question a researcher had to ask at all: instead of hunting for documents, they could ask the research directly and get an answer back cited to the studies behind it. This is the platform the walkthrough above steps through, and it comes down to a few core systems.
Not hunting for documents. Asking the research directly.
The platform runs on a conventional, staffable stack, chosen so Education.org could maintain and extend it after handover.
Carries the query interface and the contribution and admin views. Questions, answers and document uploads happen without page reloads.
Runs the platform’s core modules: users, documents, the library, the processing workers and the chat layer. Mature, well-documented, and easy to hire into.
Stores the embedded representation of every document, so a plain-language question retrieves the passages that actually match its meaning, not just its keywords.
Holds accounts, document metadata, usage history and the structured data behind the dashboard.
Translates non-English documents during ingestion, so the whole library lives in one searchable language.
Generates the written answer from the retrieved evidence, grounded in the documents rather than in the model’s general knowledge.
Ingestion runs as background jobs that handle many documents at once, with OCR, language detection and translation happening in sequence before a document joins the library.
Every component containerised and deployed on AWS across RDS, EC2 and S3, so the platform runs consistently and scales without re-architecting.
Most of what Education.org owns is research it already paid to produce, and before this engagement much of it sat idle, out of reach because getting to it cost days. What we built changed what that asset is worth. Research that once took days to search now answers a question in seconds, so the same collection gets used far more often, in more decisions, and nothing new had to be commissioned for that value to appear. For an organisation whose whole output is evidence, the cheapest research it can produce is the research it already owns and can finally reach.
The first flame is no smaller for having lit the rest.
It also changed what kind of organisation they are. They used to publish reports and hope the right practitioner found the right one; now they run a public tool that answers questions straight from their library, cited to the studies behind it, in whatever language the question arrives in. The same shift hands their researchers back the hours that sourcing, translating and screening used to eat, and it turns every question the field asks into something they can finally see rather than guess at.
And it gave them a proof point. Education.org became one of the very few NGOs in their field with AI built into how they actually work, at a moment when almost none had. In a space where credibility and being early both count, that is something they can put in front of a funder.
Sallyann had spent a year telling NGO leaders they needed to become AI-first. When we finished, she could point at Education.org and show them one that was.
The engagement ran as a discovery phase, a hard pause, and then a focused build.
Everything we built afterwards, and the one thing we did not, traces back to that map.
It was the right call, and it is one we would make again.
Four lessons out of two years with a research organisation. Two of them are about knowing when not to build.
Working with experts new to AI, we learned not to assume what was obvious to us was obvious to them. Setting honest expectations about what the models could and couldn’t do wasn’t a side conversation, it was part of delivering the product.
The harder automation wasn’t reliable on the models we started with. Pausing rather than forcing it, then resuming when the technology caught up, produced a better product than pushing through would have. Timing is a design decision.
When the harder automation stalled, the real problem wasn’t the model, it was that the task leaned on subjective human judgment no model of the day could reliably apply. The lesson was to look hard at the process itself before blaming the tool.
Most of the real engineering went into the unglamorous work of getting messy, multilingual, scanned documents into clean, searchable shape. The quality of every answer traces back to that pipeline.
That is the whole point of putting a system over a body of knowledge. An organisation spends years producing research, and we build the thing that puts it in front of the people who have to act on it, in seconds, cited to the study behind it. Education.org did that with an archive it already owned and had already paid for.
We mapped how their researchers actually worked before proposing anything, said plainly what the models of the day could and could not do, held back the one project that wasn’t ready, and built the rest. What it left behind is an NGO that runs on AI and can point at the thing to prove it.
You hold a body of expertise or evidence that is expensive to reach, the way Education.org’s research was, and you can see that the next version of your organisation runs on it rather than around it.
You want a partner who understands the business model as well as the technology, not a dev shop that only builds what you specify.
You want a chatbot for your dashboard, AI for the press release, or features a foundation model will commoditize in six months.
We will tell you that directly.
30 minutes. No deck. We will talk through your challenge, share some relevant work, and see if it makes sense to work together. Or honestly, just grab coffee and chat, no agenda needed.
Book an Intro Call →Prefer email? contact@ionio.ai