Case Study • Education.org

How We Turned 1,000+
Unsearchable Documents
Into Research That
Answers in Seconds
For a Non-Profit

Education.org had already paid for its research, but reaching it took days, so most of it went unused. Their evidence decides where NGO funding and effort go, and it sat in scanned PDFs and languages the reader didn’t speak.

We changed the economics of an archive they already owned. We mapped their research workflow stage by stage and built for each one. The centrepiece made the whole of their research answerable in one question, for their team first and then for the public they serve.

Book a call
1,000+ documents
Ingested and made searchable
Multilingual
Auto-translated into one knowledge base
Internal Public
RAG platform, then a public Ask Evidence tool
AI-first NGO
Among the first in their space
Testimonial

What the Partner Said

Education.org came to us through Sallyann, founder of Gleac, whose community Ionio serves as its AI team. She brought us in and stayed close to the work.

On why she chose us

“I try to make sure when I’m bringing in teams, when my internal team can’t handle it, that I bring in someone who can handle it from A to Z.”

On staying out of the way

I had very little involvement. I think I attended three or four meetings, and you and your team handled that project.”

On the trust

“I was comfortable already with you. I knew that you could manage that, and your team handled that project.”

On the depth of the work

“There was such depth in terms of the work. They’re an example of an NGO that really took the first step.”

The Problem

The Expensive Problem

The Business Problem

Education.org had decided where it wanted to be: an AI-first organisation, one of the few NGOs in its field building AI into how it actually works. That was a positioning bet as much as a technical one. In a space where credibility and being early both matter, an NGO that can show it runs on AI stands out to the funders and partners it depends on, and it opens the door to more of the work it wants to do.

The trouble was the gap between wanting that and living it. The organisation ran on knowledge work, and while people had started reaching for AI here and there, none of it was regimented. There was no system, no shared tooling, nothing that turned scattered experiments into a way of working. Closing that gap is what the engagement was really about.

The Same Piece, Six Players

No Shared Tooling
THE SAME PIECE, SIX PLAYERS GOLD MARKS WHERE AI GOT USED 01 02 03 04 05 06 NO TWO PLAYING IT THE SAME WAY · NOTHING KEEPING TIME

Everyone played it their own way, and nothing was keeping time.

The Workflow Problem

Underneath the positioning sat the daily reality, and it is easiest to see through one researcher. Picture someone who needs to know what the organisation has already established about, say, girls’ schooling outcomes in a particular region. The answer almost certainly exists somewhere in the research. Getting to it is the hard part.

They would start searching, one repository at a time. Some of what they need is in languages they don’t read. Some is locked in scanned PDFs with no searchable text. Some was filed years ago by someone who has since moved on. So the search becomes a hunt, and the hunt takes days, and at the end of it they have a stack of documents they still have to read before they can answer the question they started with.

The Widening Search

One Question · Days of Hunting
THE ANSWER HERE ALL ALONG 1 2 3 STILL TO READ
  1. 1A language they don’t read
  2. 2Scanned, with no searchable text
  3. 3Filed by someone who has moved on

Days of searching outward, and the answer was at the centre.

That was one stage of a longer chain. Searching, then cataloguing, then screening each source against the inclusion criteria, then coding it, then pulling findings together, each step manual, each done in a different tool, each eating the time of skilled people. Every hour spent on it was an hour not spent on the judgment only a researcher can provide.

The organisation’s most valuable people were spending their days on work a machine could carry, and its research was going underused because reaching it cost too much.

Why This Was Hard

We began in 2024, before the current generation of reasoning models existed, inside a specialist discipline neither of us could shortcut. Almost none of what made this difficult was code.

A Closed Research Discipline
Domain

Evidence synthesis is a specialist craft with its own rules, and getting it wrong means the output can’t be trusted. We had to learn how their researchers actually work, what grey literature is and why it’s hard to source, how inclusion and exclusion criteria get applied, and what a librarian needs to see before they’ll rely on a tool for real work.

What this meantWe spent real time learning the discipline before we designed anything.

The Models Weren’t Ready Yet
Core

We began in 2024, before the current generation of reasoning models existed. One of the projects we took on, automating the judgment-heavy screening and coding work, ran into a wall the models of that moment couldn’t clear, so we held it at the pilot stage rather than ship something unreliable into a research organisation. When reasoning-focused models arrived not long after, that work became possible.

What this meantKnowing when to wait was as important as knowing what to build.

Messy, Multilingual, Real-World Sources
Complex

The source material was not clean text. Scanned PDFs with no readable layer, documents in languages the reader didn’t speak, formats that changed from one repository to the next. Before anything could be searched it had to be pulled in, read by OCR where needed, and translated into one common language, so that a single question could reach across the whole collection.

What this meantThe ingestion pipeline had to make messy documents behave like clean ones.

Building for Experts New to AI
Audience

The people we were building for were deep domain experts who hadn’t yet had reason to work closely with AI. That changed the job. Part of the delivery was showing them what these tools could and couldn’t do, setting honest expectations, and designing something that felt trustworthy to someone encountering it for the first time.

What this meantWe were teaching as much as we were shipping.

A closed discipline, a technology that was not ready, and a room of experts who had to trust the thing before they would use it. The engineering was the tractable part.

Opening illustration for the philosophy behind the Education.org build
The Philosophy

Adoption Before Automation

Rolling out AI inside an organisation that is excellent at one thing is not mainly a technical problem. The people who have to trust the tool are often the ones with the least reason to have kept up with AI, because they were busy being expert at the actual work. Education.org was a room full of exactly that: sharp researchers who simply hadn’t had time to dig into what these models do or don’t do.

So the job was never just to automate a workflow. It was to bring a team of experts along at a pace they could trust, while building the thing that would earn that trust. In practice that meant three things:

  • Understand the knowledge work before proposing anything. We mapped how their researchers actually work before we designed a single screen.
  • Be candid about what the models of the day could and couldn’t do. Honest limits set early are cheaper than disappointment set late.
  • Scope to what would genuinely work, not what would demo well. Including the moment we recommended pausing, because the technology was not ready.

The principle travels to any expert organisation adopting AI. Understand the work first, set honest expectations, and build for the people who have to adopt it.

The two halves of the job AUTOMATION THE EASY HALF ADOPTION THE HALF THAT DECIDES

The automation is the easy half. Getting a room of experts to rely on it is the half that decides whether the thing gets used at all.

That second half showed up in how we worked, not just what we built. We walked in assuming certain things were obvious, that everyone understood what these models could and couldn’t do. We were wrong. These were deeply smart researchers with real depth in their field who simply hadn’t had time to dig into AI, and teaching the technology’s limits was a prerequisite for the engineering to land, not a nice-to-have.

It shaped how we communicated too. We didn’t ask open-ended questions, we came with options, framing decisions as A, B or C rather than “what do you want,” explaining the tradeoffs, and making sure no one had to chase us for an update. That is how you keep a non-technical stakeholder in control of a deeply technical process, and over time it changes what they can do.

The same instinct applies once a product is live. When the tools underneath an AI system move on, the right response isn’t to wait for the client to notice. It is to come back and say plainly: here is what we built, here is what has changed, here is what it means for you. Framed as education rather than a sales pitch, that is what turns a delivered project into a lasting one.

Whose hands stay on the wheel WE CHART IT THEY STEER

“Here’s what we did before, here is what the new technology is right now. For a client that it is working, that’s actually better for them in the long run.”

Product Walkthrough

The Platform in Action

From a plain-language question to a cited answer, then the library it draws from and the site where the public meets it. Here is the Ask Evidence platform in action.

01 / Ask

Your education questions, answered

The tool opens on a single prompt: ask a question about education and get a research-backed answer drawn from Education.org’s own verified resources. A “Popular Topics” row offers real starting points, from education policy for equity to the role of NGOs, so a first-time user isn’t staring at an empty box. One question, the whole evidence base behind it.

The Ask the Evidence landing page, with a single question prompt and a row of popular topics
02 / Answer

Cited to the source

Ask something like what role NGOs play in education policy, and the answer comes back written out, with a Sources card naming the exact document it drew from and inline markers tying each claim to it. Underneath, a Related Questions list suggests where to go next. Every statement traces back to a real study, which is what makes it usable for research rather than just readable.

An answer about the role of NGOs in education policy, with a named source document, inline citation markers and a list of related questions
03 / Library

The evidence behind the answers

The Library holds every document the tool answers from, in one place. National education plans from Guinea-Bissau, Iraq, Malawi, Zimbabwe, Myanmar and more sit alongside research reports, each a real source the answers can cite. This is the corpus, kept in view and managed by the team rather than buried in a database.

The Library view listing national education plans and research reports the tool answers from
04 / In Place

Where the public meets it

The same tool lives on Education.org’s own site, reachable from an “Ask the Evidence” button on every page. A visitor reading about the organisation’s work can put a question to its research without leaving, which is how the evidence reaches the professionals it’s meant for.

Education.org's public site with the Ask the Evidence button in the corner of the page
The Solution

What We Built

We started by mapping Education.org’s whole research workflow, the chain that ran from searching for sources, through cataloguing and screening them, to coding and pulling findings together. Then we went stage by stage and built to replace the manual work at each one. That produced three projects: two focused tools that took on specific stages, and one larger platform that became the centrepiece. Ionio owned the full arc of each, product strategy, wireframes, UI and UX, the build, deployment and handover.

Project 01

Search Automation

The workflow began where every piece of research does, with finding sources, and that was the first stage we automated. A researcher used to work through repositories one at a time, hunting across libraries, journal databases and open repositories, much of it grey literature that is awkward to find and awkward to read.

We built scrapers that went out to the sources that mattered, among them the World Bank and the African Education Database, and pulled the documents in automatically. Across the target sources this gathered more than a thousand documents into one place, turning days of manual desk searching into a process that ran on its own.

Where the Documents Come From

Scraped, Not Searched
WHERE THE SOURCES LIVE World Bank African Education DB Journal Databases Open Repositories Grey Literature ONE PLACE SCRAPERS RUN THE SEARCH · NOBODY OPENS A REPOSITORY BY HAND
1,000+Documents gathered across the target sources
5 source typesLibraries, journals, open repositories, grey literature
UnattendedThe stage runs without anyone opening a repository

Days of desk searching, replaced by a process that runs on its own.

Project 02

Screening and Coding

The next stage was the hardest, and the most revealing. Once sources are found, researchers screen each one for inclusion against a set of criteria, then code it, work that is slow, expert, and done entirely by hand. We built to automate it, and in doing so we found the real limit was not the tooling but the moment.

The coding framework leaned on subjective human judgment that the models of the day could not reproduce reliably, so we held this at the pilot stage rather than ship something a research team couldn’t trust. It became the clearest lesson of the engagement: sometimes the right build is the one you wait to finish.

Where the Five Passes Landed

Spread, Not Accuracy
THE CODE ASSIGNEDTHE ANSWER YOU CAN RELY ON ABCD0102030405 RESEARCHER · ONE CODEMODEL · FOUR CODES

A research team cannot use an answer that changes every time you ask.

Project 03

The Ask Evidence Platform

The third project grew into the main one. Rather than speed up one stage, it changed the question a researcher had to ask at all: instead of hunting for documents, they could ask the research directly and get an answer back cited to the studies behind it. This is the platform the walkthrough above steps through, and it comes down to a few core systems.

Four Systems Around One Shelf

Everything Touches the Middle
EVIDENCE LIBRARY CONTRIBUTION INGESTION ASK EVIDENCE USAGE TWO SYSTEMS FEED IT · TWO SYSTEMS DRAW FROM IT

Not hunting for documents. Asking the research directly.

  • The Ask Evidence tool. The part people touch. A plain-language question goes in, and a written answer comes back drawn from Education.org’s own research, cited to the studies behind it. It runs the same way for the internal team and for the public they serve.
  • The Evidence Library and ingestion pipeline. Every document that enters the platform gets processed on the way in. Language is detected, scanned files are read with OCR where needed, and non-English material is translated through DeepL into one common knowledge base, so a single question can reach across the entire collection whatever language each source started in.
  • The Contribution System. Researchers add their own documents to the central library, with the metadata that makes them findable. The collection grows with the team, and everything added is processed and searchable straight away.
  • The CRM and usage view. Behind the tool, every session is kept and readable, listed by the person who asked, so the organisation can see what’s being asked and how the platform is being used rather than guess at adoption.
Internal Infrastructure

Technical Infrastructure

The platform runs on a conventional, staffable stack, chosen so Education.org could maintain and extend it after handover.

Frontend

React

Carries the query interface and the contribution and admin views. Questions, answers and document uploads happen without page reloads.

Backend

Node.js and Express

Runs the platform’s core modules: users, documents, the library, the processing workers and the chat layer. Mature, well-documented, and easy to hire into.

Vector Database

Qdrant

Stores the embedded representation of every document, so a plain-language question retrieves the passages that actually match its meaning, not just its keywords.

Database

Postgres

Holds accounts, document metadata, usage history and the structured data behind the dashboard.

Translation

DeepL API

Translates non-English documents during ingestion, so the whole library lives in one searchable language.

Language Model

GPT-4o

Generates the written answer from the retrieved evidence, grounded in the documents rather than in the model’s general knowledge.

Document Processing

Worker and job queue

Ingestion runs as background jobs that handle many documents at once, with OCR, language detection and translation happening in sequence before a document joins the library.

Cloud

Docker and AWS

Every component containerised and deployed on AWS across RDS, EC2 and S3, so the platform runs consistently and scales without re-architecting.

The Results

The Business Impact

Most of what Education.org owns is research it already paid to produce, and before this engagement much of it sat idle, out of reach because getting to it cost days. What we built changed what that asset is worth. Research that once took days to search now answers a question in seconds, so the same collection gets used far more often, in more decisions, and nothing new had to be commissioned for that value to appear. For an organisation whose whole output is evidence, the cheapest research it can produce is the research it already owns and can finally reach.

ONE SOURCE, LIT AGAIN AND AGAIN, NEVER SPENT THE ARCHIVEALREADY PAID FOR THE RESEARCH TEAM THE PEOPLE THEY SERVE REACHING THE RESEARCH DIRECTLY

The first flame is no smaller for having lit the rest.

It also changed what kind of organisation they are. They used to publish reports and hope the right practitioner found the right one; now they run a public tool that answers questions straight from their library, cited to the studies behind it, in whatever language the question arrives in. The same shift hands their researchers back the hours that sourcing, translating and screening used to eat, and it turns every question the field asks into something they can finally see rather than guess at.

And it gave them a proof point. Education.org became one of the very few NGOs in their field with AI built into how they actually work, at a moment when almost none had. In a space where credibility and being early both count, that is something they can put in front of a funder.

What It Added Up To

What Each Side Gained

01Research Value

What the Archive Gained

  • Days → Secondsto reach an answer inside the same collection
  • Nothing newhad to be commissioned for that value to appear
  • Any languagea question arrives in, answered from one library
02Organisational Value

What Education.org Gained

  • Hours backfor the judgment only researchers can supply
  • Every questionthe field asks, now visible rather than guessed at
  • Among the firstNGOs with AI built into how they work

Sallyann had spent a year telling NGO leaders they needed to become AI-first. When we finished, she could point at Education.org and show them one that was.

Project Roadmap

Our Approach

The engagement ran as a discovery phase, a hard pause, and then a focused build.

Execution Timeline (2024–2025)
Discovery
The Pause
The Build
Discovery & Workflow Mapping 2024
  • Learned how Education.org works, mapping the research workflow from searching through screening and reporting
  • Established where AI could help and where it could not yet
  • Defined the projects worth taking on: search automation, then screening and coding, then the platform

Everything we built afterwards, and the one thing we did not, traces back to that map.

The Pilot That Waited 2024
  • Screening and coding ran into the limit of the models available at the time
  • The task demanded subjective judgment the models could not reproduce reliably
  • Held at the pilot stage rather than shipped unreliable into a research organisation

It was the right call, and it is one we would make again.

The Build 2024 to 2025
  • Reasoning-focused models arrived and the shelved work became possible
  • Ingestion pipeline first, because everything else depended on it
  • Then the Ask Evidence tool, the contribution system and the CRM
  • Tested and adopted by the Education.org team, then opened to the public they serve
Retrospective

What We Learned

Four lessons out of two years with a research organisation. Two of them are about knowing when not to build.

I.

Educating the client is part of the build

Working with experts new to AI, we learned not to assume what was obvious to us was obvious to them. Setting honest expectations about what the models could and couldn’t do wasn’t a side conversation, it was part of delivering the product.

II.

Sometimes the right move is to wait

The harder automation wasn’t reliable on the models we started with. Pausing rather than forcing it, then resuming when the technology caught up, produced a better product than pushing through would have. Timing is a design decision.

III.

The bottleneck can be the work, not the model

When the harder automation stalled, the real problem wasn’t the model, it was that the task leaned on subjective human judgment no model of the day could reliably apply. The lesson was to look hard at the process itself before blaming the tool.

IV.

A search is only as good as what’s been ingested

Most of the real engineering went into the unglamorous work of getting messy, multilingual, scanned documents into clean, searchable shape. The quality of every answer traces back to that pipeline.

Evidence that nobody can reach is evidence that may as well not exist.

That is the whole point of putting a system over a body of knowledge. An organisation spends years producing research, and we build the thing that puts it in front of the people who have to act on it, in seconds, cited to the study behind it. Education.org did that with an archive it already owned and had already paid for.

We mapped how their researchers actually worked before proposing anything, said plainly what the models of the day could and could not do, held back the one project that wasn’t ready, and built the rest. What it left behind is an NGO that runs on AI and can point at the thing to prove it.

How We Operate
01 // Integration
We embed with your technical team. The work you just read represents how we operate. We build production systems that connect into your existing architecture, and transfer the knowledge so you own what we build.
02 // Acceleration
We do not start from zero. The tooling we have developed across our engagements, the pipelines, the agent frameworks, the eval systems, accelerates every project we take on. You are not paying us to learn on your time.
03 // Experience
We know what works. We have been building AI systems for a decade. We shipped architectures before they became mainstream, and we know the pitfalls because we have made the mistakes already.

When to Talk to Us

You hold a body of expertise or evidence that is expensive to reach, the way Education.org’s research was, and you can see that the next version of your organisation runs on it rather than around it.

You want a partner who understands the business model as well as the technology, not a dev shop that only builds what you specify.

When We’re Not a Fit

You want a chatbot for your dashboard, AI for the press release, or features a foundation model will commoditize in six months.

We will tell you that directly.

Next Step

Let’s See If There’s a Fit

30 minutes. No deck. We will talk through your challenge, share some relevant work, and see if it makes sense to work together. Or honestly, just grab coffee and chat, no agenda needed.

Book an Intro Call →

Prefer email? contact@ionio.ai