Demo project
Knowledge base agent that quotes its sources
An answer is built only from the retrieved document fragments and carries a verbatim quote naming its source.
development · Node.js · SQLite · OpenRouter API · RAG · bge-m3

The task
Every company wants a chat over its own documents: staff and customers ask in plain words instead of scrolling a forty-page manual. One thing gets in the way. A thin wrapper around a language model answers confidently no matter what is asked, including questions the documents never covered. An invented warranty term or an invented price costs far more than an honest “that is not in the knowledge base”.
The demo runs on eight documents of a fictional joinery workshop: production rules, materials and finishes, price list, delivery and carrying, warranty, product care, measuring and payment, and a FAQ. The company is invented outright and the files hold no contacts and no personal data - this demonstrates the mechanics, not somebody’s real paperwork.
The solution
The path from a document to an answer.
A document is cut by sections, each section by paragraphs, and paragraphs are packed into fragments of up to 900 characters. A fragment remembers which document and section it came from and its own position - fragment 2 of 6 - and that is what the caption under a quote is later built from. Eight files produced 53 fragments, 315 characters on average, 821 at the longest.
The embedding model receives “document. section. text” rather than the bare paragraph: stripped of its headings, a paragraph titled “Terms” means nothing at all. Vectors are normalised, so similarity is a plain dot product.
The index lives in SQLite: three tables, the vector stored as a blob. For 53 fragments a separate vector database would be ceremony, while a file survives a restart and opens in any client.
The question becomes a vector, is compared against every fragment and cut off at a similarity threshold; the best five go into the answer. On top of that sits a cap of three fragments per document - insurance for the case where one file takes the whole result set. That the cap works is shown by the self-check, on a small separate index.
The model receives the retrieved fragments labelled with document and section, plus one rule: answer only from these, put the fragment number after every statement, quote verbatim. The reply comes back as JSON.
Each quote is then looked up inside its own fragment as a substring. If it does not match, the quote is discarded, and the page shows how many were discarded.
Three barriers against invention
The first is the similarity threshold. A question that misses the knowledge base retrieves nothing at all, and the answering model is never called: it has nothing to improvise from.
The second is the found flag in the model’s structured reply. Fragments were retrieved but hold no answer - the agent says so plainly. The recorded run shows this on a question about instalment plans: five fragments cleared the threshold, including the payment terms, and the agent read them and declined to invent financing.
The third is quote checking. It guards not against “not found” but against an elegant fabrication attached to a real document.
Details that are easy to miss
- A self-check runs 14 tests of the mechanics with no model call and no network: a doctored quote fails verification, a quote stitched from two places in one fragment fails, a reference to a fragment outside the retrieved set fails, the threshold cuts off distant matches, and no single document takes the whole result set. 14 of 14 pass.
- Verification normalises spacing, dash forms and quotation marks on both sides, so the check trips on facts rather than on typography.
- Answer temperature is set to zero: nothing in the request to the model deliberately adds spread to the wording.
- One compound question was answered from three documents at once: the price came from the price list, the delivery fee from the delivery document, and the panel weight from the materials document, so the agent could work out that the item exceeds 30 kg and carrying is billed separately. No single document answers that question.
- API calls run with a timeout and three attempts with growing delays; the key is read from an environment file, never hardcoded and never printed.
- There are no external dependencies - Node’s built-in fetch and SQLite are enough. Models and the similarity threshold are switched through environment variables.
- Every question asked on the page is appended to a run log, so a screenshot always captures a real answer rather than a mock-up.
- The demo is served locally, and every question on the page is a paid model call. That is why it has no open live link and the write-up shows recorded runs instead.
- Timings come from one recorded run and depend on machine and network: question vector 224-921 ms, search across 53 fragments 2-8 ms, model answer 1.1-1.8 s.
Screenshots




