System 05 · Agentic RAG

AI knowledge base that answers from your documents.

An AI knowledge base and chatbot for your business, built on retrieval-augmented generation (RAG): an agentic assistant that answers from your own material instead of guessing. It cites what it used, and it says it does not know rather than inventing something plausible. When a question falls outside the corpus, it escalates to a human instead of answering, which is what makes it deployable where wrong answers cost money.

RetrievalEmbedded, over your corpus
GroundingAnswers cite their source
RefusalSays "not in the material"
HostingOur cloud, or yours

Reviewed by Martin Kelly, founder of Botonomy Updated

The short answer

What is an AI knowledge base?

An AI knowledge base is a system that lets people ask questions of your own documents in plain language and get an answer with its source. It works by retrieval-augmented generation (RAG): it finds the relevant passages in your material first, then answers only from those. A RAG chatbot is the same system with a chat window on the front.

A general AI chatbot answers from what it was trained on, which does not include your prices, your policies or last month's product change. That is where confident wrong answers come from. An AI knowledge base is limited to your material on purpose, and shows which passage each answer came from.

The agentic part is the judgement about when not to answer. Below a confidence floor it says it does not know, offers a person, and logs the question, so the gaps in your material become a list you can fix.

Agentic RAG, defined. Agentic RAG is retrieval-augmented generation in which the assistant decides how to answer as well as what to retrieve: it searches your documents, answers with citations when the material supports it, and declines or escalates to a person when it does not.

A real answer from Aria, the AI knowledge base chatbot on botonomy.ai, explaining the difference between the Operate and Accelerate plans from Botonomy's own pricing material
A real answer from the assistant on this site, which runs on this system.
Architecture

Three layers that scale answers — deterministically.

Directive

What to do

The SOP, written down. Goals, inputs, the tools to use, the outputs and the edge cases — in plain language, versioned like code. When we learn an API limit the hard way it gets written here, so it is never learned twice.

Orchestration

Decision-making

The only probabilistic layer. It reads the directive, calls the tools in order, handles the errors and asks when something is genuinely ambiguous. It does not do the work itself.

Execution

Doing the work

Deterministic scripts. Same input, same output, every run. Ninety per cent accuracy per step is fifty-nine per cent over five, so anything that must be right lives down here instead.

Data flow

In, through, out.

What it reads
Your documents, pages and knowledge base
Product and pricing material
Past support and sales conversations
Problem → solution
The same question answered by hand, dailyAnswered once, then answered on your behalf
A chatbot that invents what it does not knowAnswers with the receipts, or an honest no
Material nobody has time to readYour own documents, finally reachable in a sentence
What you get
Embedded, queryable index
Grounded answer with citations
Captured lead from the conversation
Gap report: what it could not answer

The part clients end up valuing most is the list of questions it could not answer. Every one of them is a gap in your material that somebody was actively looking for — told to you, rather than guessed at.

Diagram of an agentic RAG system: ingest your documents, retrieve the relevant passages, answer with the source or decline and hand off, and log the gaps
Below the confidence floor it declines, offers a person, and logs the question.
Where it fits

General AI chatbot vs help-centre search vs AI knowledge base

Three ways to let a customer or a colleague find an answer without asking a person.

General AI chatbotHelp-centre searchAI knowledge base (RAG)
Answers fromIts training dataKeyword matchesPassages in your own documents
Shows its sourceNoA list of articlesYes, for every answer
When it does not knowOften guessesReturns nothingSays so and offers a person
Kept currentNot with your changesWhen articles are editedRe-read from your material on a schedule
What you learnNothingSearch termsEvery question it could not answer
Where your data livesWith the vendorYour help centreYour choice, including your own infrastructure
Integrations

It reads your stack. It doesn't replace it.

  • ClaudeGrounded answering
  • OpenAIEmbeddings + fallback
  • PineconeVector retrieval
  • QdrantSelf-hosted vectors
  • PostgreSQLSource of record
  • LangChainRetrieval orchestration

Also connects to

Your CMSGoogle DriveNotionClickUpSlackAnalytics 4

Read-only wherever read-only is enough. Nothing gets write access it doesn't need, and every write is logged.

Quality gates

Checked before it ever reaches you.

Every answer is traceable to retrieved passages — no ungrounded generation.

Below the confidence floor it refuses instead of guessing.

Unanswered questions are logged rather than silently dropped.

That consistency is the whole point. The same checks and the same artifacts on every run, so when a number moves you know it moved because your site moved — not because the method did.

Datasheet

The boring details.

SpecStandardDeployable
Where it runsOur infrastructureYours
Who holds the API keysUs, scoped per clientYou
CadenceMonthly, or weeklyAny schedule
Raw data retentionRebuilt each runYour policy
Artifact deliveryYour Drive and trackerYour choice
Source code access—Full
Runs unattendedYesYes
Failure alertingSlack and emailYour channels
The other systems

Same architecture, different job.

Want it in your own stack?

Deployable hands over the software, the hosting and the keys. You run it; we stay on for support.

Talk scope
Before you start

Questions about the AI knowledge base.

What is an AI knowledge base?

An AI knowledge base is a system that lets people ask questions of your own documents in plain language and get an answer with its source. It finds the relevant passages in your material first, then answers only from those, which is what keeps it from making things up.

How is a RAG chatbot different from a general AI chatbot?

A general AI chatbot answers from its training data, which does not include your prices, policies or products. A RAG chatbot is limited to your own material and shows which passage each answer came from. When your material does not cover the question, it says so.

Can it go on our website as a chatbot for customers?

Yes. The same knowledge base can sit behind a chat window on your site, answer questions before a sale and capture the lead when someone asks to talk to a person. The assistant on this site runs on the same system.

How does it stay up to date when our documents change?

The source material is re-read on a schedule and the index is rebuilt when the content has changed. Each document carries a review date, and one that is past its date is flagged, because an out-of-date document produces a confident wrong answer that no chatbot check will catch.

What is a RAG system?

A RAG system is an assistant that answers from your own documents instead of from general training data. Retrieval-augmented generation means it finds the relevant passages in your material first, then answers from them, and cites what it used.

What does it do when it does not know?

It says so. The system is built to decline rather than produce a plausible answer it cannot ground in your material, because a confident wrong answer costs more to undo than no answer costs to escalate.

Where does our data live?

Wherever you choose. On Deployable the system runs inside your own infrastructure with your own keys, so your documents never leave your environment.