How much of your manual can DOKA hold?
One DOKA knowledge base is rated at 250 procedure documents — about 463,000 characters, roughly 150 pages of written procedure, ten times the size of the practice manual we test against. That number is measured, not quoted from a spec sheet. This page shows the test and the results at every size.
Why there is a limit at all
DOKA does not search your manual for a few likely passages and read only those. Every question is answered with the entire knowledge base in front of the model. That is how every answer can carry a citation to the exact document, and why “the manual doesn’t cover that” is a real refusal rather than a failed search.
The cost of that design is a ceiling. The model reads a fixed amount at once, and as the manual grows the question becomes whether answers, refusals and citations stay disciplined near the edge. So we measured it.
How we measured it
- Start with a real knowledge base. The 25 procedure documents of a working dental practice, kept byte-for-byte identical at every size.
- Grow it with realistic filler. Synthetic procedures for departments and locations the real manual does not cover — sterilization logs, lab case tracking, supply ordering, staff scheduling — written in the same format, up to 250 documents.
- Plant near-misses. One in ten added documents is a deliberate distractor: the same topic as a real procedure, but for a different location or role. Another office’s Saturday hours. Another office’s phone number. A different person to escalate to. They exist to catch an assistant that answers from the wrong document.
- Ask the same questions at every size. The 38-question control suite the product runs before every release: grounded answers, refusals where the manual is silent, trap questions that invite a plausible guess, Spanish, and prompt-injection attempts. Plus six anchoring probes, each asking for a real fact that a distractor contradicts for another location.
- Run it on production code. A throwaway copy of the console — same prompt, same model, same code path customers use. Gemini 3.8 Flash, with every question asked three times at every size.
What we found
| Documents | Characters | Control suite | Anchoring probes | Per question | Verdict |
|---|---|---|---|---|---|
| 25 | 43,000 | 38 / 38 | 6 / 6 | 1.4 s | Our 25-document test practice manual. Clean. |
| 50 | 90,000 | 37 / 38 | 6 / 6 | 1.7 s | Clean.* |
| 100 | 184,000 | 37 / 38 | 6 / 6 | 1.9 s | Clean.* |
| 200 | 370,000 | 37 / 38 | 6 / 6 | 1.7 s | Clean.* |
| 250rated | 463,000 | 37 / 38 | 6 / 6 | 1.6 s | Clean.* The rated size, and the largest we ran. |
* The one miss from 50 documents up is the near-miss design doing its job. Asked about Saturday hours, DOKA answers that the practice’s three real locations are closed on Saturdays and attributes the Saturday hours to the synthetic annex that has them, in all three trials at every size. The check was written before the corpus contained another location’s Saturday hours, so it counts a correct, correctly anchored answer as a miss. We report it rather than rewrite the check after the fact.
Where the limit is, and why we rate at 250
On the serving model we have not reached it. At 250 documents, the largest size we ran, every one of the 134 questions was answered, every trap and boundary question held in all three trials, and every anchoring probe found the right location’s fact. An earlier sweep on a different model, on September 7, 2026, degraded at 250 documents, which is why the size is measured again whenever the serving model changes.
We rate at 250, the largest size measured, not at an edge we have not found. A practice whose manual is bigger than that gets more than one knowledge base — front office, clinical, billing — each with its own citations. Nothing is summarized, trimmed or dropped to make it fit.
Where does your manual sit?
Count the documents in your manual and pick the length that looks most like them. Bring the real thing to the free audit and we will size it exactly.
A typical procedure document in our test practice manual runs 1,200 to 2,600 characters, about one printed page with headings.
What this page does not claim
- Three trials, one day. Every question ran three times at every size on September 25, 2026. At 200 and 250 documents the questions were spaced about seven seconds apart: one question at that size carries 80,000 to 100,000 tokens, and back-to-back questions go over the model vendor’s per-minute token limit.
- Measured on the console. OKA Anywhere — the connector that runs inside Claude or ChatGPT — answers with the same model, but its path was not part of this sweep. It carries no capacity rating until it is.
- Tied to a model. When the serving model changes, the sweep is re-run before this rating changes. The date at the top of this page is the date of the run it reports.
Method details
Characters are the honest unit. Token counts depend on the model’s tokenizer; on this text the serving model used about 0.22 tokens per character, so a question at the rated 463,000 characters carries roughly 101,000 tokens of manual, plus the conversation itself.
Sizes run: 25, 50, 100, 200 and 250 documents, with 134 requests at each size. The synthetic filler never contains the facts the trap questions forbid — fee amounts, holiday hours, financing products — so a refusal stays the correct answer at every size.
“Per question” is the median time to answer across the control questions at that size, including model latency. The 25 real documents sit in their original positions at every size; the filler is generated from a fixed seed, so any run can be reproduced exactly.
Start with the audit
Before anything, we’ll do a free front-desk operations audit — a short review of where your documented procedures have gaps, delivered as a written findings report, with your manual sized against the numbers on this page. No cost, no obligation.