All Knowledge Is Not the Same
Finding something is not the same as knowing something
I read a genuinely good AWS article recently about building an AI-powered knowledge management system. It is solid engineering and worth your time: Bedrock Knowledge Bases doing the RAG orchestration, documents in S3, OpenSearch Serverless as the vector store, Titan embeddings, a voice-first interface so frontline workers can use it hands-free, and the whole thing deployable in hours.1 No notes. That is a well-built system.
What I liked most was how they framed the problem, because it is the same one I keep coming back to. Their words: "This 'tribal knowledge' often disappears when key personnel leave, creating knowledge gaps that impact efficiency and innovation."1 I closed an earlier post with almost exactly that thought, so we are clearly worrying about the same thing.
The basic idea is simple enough. An organisation has a mountain of documents lying around: policies, procedures, manuals, reports, old project material, years and years of accumulated stuff that contains useful information but is painful for a human to find. So you put it somewhere an AI can search. Ask a question, retrieve the relevant passages, hand them to a model, get an answer. That is useful, and I am not about to argue otherwise.
But it set us off on something slightly different, which is that all knowledge is not the same.
Two questions that look similar and are not
Give me five documents and ask me this:
"What do these documents say about Acme?"
That is answerable. Search them, find the relevant passages, summarise. RAG does this well, and the AWS system would do it well.
Now ask me this instead:
"What do we actually know about Acme?"
Suddenly it is a much harder question, and the difficulty has nothing to do with retrieval. Perhaps three of those documents are about the same Acme and two are about a different company with a similar name. Perhaps one uses a trading name that changed in 2023. Perhaps two of them flatly disagree, and one is four years older than the other. Perhaps one contains a definition that governs how you should read the other four. Perhaps one contains evidence that overturns what the others assert. And perhaps, having looked at all of it properly, the honest answer is that there is not enough here to say.
None of that is a search problem. Every one of those is a knowledge problem, and no amount of better embeddings will touch them.
Semantic similarity is very clever right up until you need to know whether something is actually true.
A pile of Lego is not a model
The analogy I keep using is Lego. The documents are the bricks. Search is very good at finding the right bricks, and a language model is genuinely impressive at describing what those bricks might make. But somebody still has to build the thing. And once it is built, you want to know which bricks went into it, because otherwise you cannot tell whether it will hold together when someone leans on it.
That assembly step is what we call knowledge compilation, and it is the part almost nobody is building. Take evidence, work out what that evidence actually supports, join the pieces together, keep hold of where every piece came from, and produce a representation of the knowledge that can be relied on rather than merely quoted.
To be clear about where we are with this, because I would rather be accurate than mysterious: we have already built it. It came out of the work on Scout, where we had to stop treating everything a model emitted as knowledge. Evidence had to stay evidence. Concepts had to be concepts. Relationships needed grounding. Conclusions had to be attributable to something specific. That sounds like basic hygiene and it is not, because we spent months discovering the many inventive ways you can accidentally promote a model's interpretation into an established fact.
The compiler has since proven useful well beyond Scout, which is why it is becoming a capability in its own right rather than a component buried inside an agent.
This is not an argument against RAG
I want to be careful here, because it would be easy to read this as a swipe at retrieval, and it is not. RAG is genuinely excellent when the job is "go and find me the information that answers this question". That is a real job, most organisations cannot do it today, and a system like the AWS one solves it properly.
The gap opens when the job changes to "build me a trusted representation of what we actually know". Notice that the AWS write-up is honest about the boundary itself: among its stated limitations is that grounding reduces but does not eliminate hallucination risk.1 That is not a criticism of their design, it is an accurate statement about what retrieval plus generation can promise. Grounding makes an answer more likely to be right. It does not establish that it is right, and it does not tell you how confident to be.
So we think there is another layer above retrieval, and the two are complementary rather than competing. Retrieval finds the material. Compilation decides what the material establishes.
Why this matters more as agents start acting
This distinction is fairly academic while AI is answering questions for a human who will sanity-check the response. It stops being academic the moment agents start making decisions and taking actions, which is exactly where the industry is heading.
An agent about to do something consequential should not have to go back to a pile of documents and work out what they mean from first principles, every time, hoping it lands in the same place it did yesterday. Sometimes it needs to be handed something far simpler and far harder to produce:
This is what we know.
This is why we know it.
This is how confident we are.
This is what we do not know.
That fourth line is the one that gets left out, and it is the one that matters most. A system that cannot tell you what it does not know will confidently fill the gap, because that is what these models are built to do. Being handed an explicit boundary is what lets an agent stop rather than improvise.
Teaching machines to know what they know
We have spent an enormous amount of collective effort teaching machines how to find information, and we have got very good at it. Retrieval is close to a solved commodity. The next problem is considerably less tractable and considerably more valuable: teaching them to know what they know, to show why they know it, and to say plainly when they do not.
All knowledge is not the same. A passage that mentions a customer, a definition that governs how you read that passage, a piece of evidence that contradicts it, and a conclusion somebody drew from all three are four completely different kinds of thing. Flatten them into one undifferentiated pile of text and you have thrown away the distinctions that determine whether the answer is trustworthy.
Which is the whole point. Finding something is not the same as knowing something, and the gap between the two is where the next few years of this actually get interesting.
- Democratizing institutional knowledge: building an AI-powered knowledge management system with AWS - AWS Machine Learning Blog. Source of the architecture, the tribal-knowledge framing and the stated limitations.