What we tested, and what it scored.
Every score on this page comes from a test that ran on the live console. The page is generated from our results archive: the newest full run of each evaluation suite, its score, the number of trials, the model it ran on, and the date. Failures are listed by test ID, not hidden.
The apparatus behind these scores is written up in the paper: Refuse or Cite.
The newest run of every suite on the live console
| Product | Suite | Score | Trials | Model | Date | Open items |
|---|---|---|---|---|---|---|
| DOKA | Control suite | 38 / 38 | 5 | Gemini 3.8 Flash, console | 2026-10-01 | — |
| OKA | Acceptance suite | 48 / 48 | 5 | Gemini 3.8 Flash, console | 2026-09-30 | — |
| OKA | Safety suite | 16 / 16 | 5 | Gemini 3.8 Flash, console | 2026-09-29 | — |
| LOKA | Acceptance suite | 51 / 51 | 5 | Gemini 3.8 Flash, console | 2026-09-30 | — |
| LOKA | Safety suite | 16 / 16 | 5 | Gemini 3.8 Flash, console | 2026-09-30 | — |
| VOKA | Acceptance suite | 51 / 51 | 5 | Gemini 3.8 Flash, console | 2026-09-30 | — |
| VOKA | Safety suite | 16 / 16 | 5 | Gemini 3.8 Flash, console | 2026-09-30 | — |
| MOKA | Acceptance suite | 50 / 50 | 5 | Gemini 3.8 Flash, console | 2026-10-01 | — |
| MOKA | Safety suite | 24 / 24 | 5 | Gemini 3.8 Flash, console | 2026-10-01 | — |
| Backup path | ||||||
| DOKA | Control suite | 37 / 38 | 5 | GPT-5.6 Terra | 2026-09-30 | RT-005 |
| OKA | Safety suite | 16 / 16 | 5 | GPT-5.6 Terra | 2026-09-29 | — |
A test counts as passed only when every trial passes. "Open items" are the test IDs that did not; each is tracked in our calibration record with its cause. A score is never edited after the fact: if a check was wrong, we say so beside it rather than rewrite it.
OKA's tests were recalibrated on September 29 to accept the procedures' own wording; every change was checked against all recorded answers first.
What the suites are
- Control suite. The 38 questions the dental product must get right before any release: grounded answers with the source cited, refusals where the manual is silent, trap questions that invite a plausible guess, Spanish, and prompt-injection attempts.
- Safety suite. One per vertical: the questions that must be refused and routed, never answered — clinical, legal, financial, and false-authority attempts to override the documents.
- Acceptance suite. Generated from each corpus by the same method we use for a customer's own documents, and validated by execution before it counts: grounding, vacuity, and adversarial wrong-answer checks.
Languages
Spanish is part of the control suite. On September 16, 2026 we asked the dental assistant three questions each in Somali and in Hmong — the cancellation policy, what to collect before scheduling a new patient, and a clinical question about antibiotics. All six answered in the caller’s language with the source document cited, and both clinical questions were refused in that language and routed to the dentist. Six of six, on Gemini 3.8 Flash.
Capacity
How large a manual one knowledge base holds, measured rather than quoted, with the results at every size: the capacity page.
How to read this
- Tied to a model. Each row names the model it ran on. When the serving model changes, the release gate re-runs before the next release, and this page changes with it.
- Release gate, not calendar. The suites run before every release and whenever the serving model changes. A date here is the date of the run it reports.
- Generated, not written. The table is produced from the results files by a script. Nobody types a score into this page.
- Backup path. The model the console switches to when the serving model cannot be reached. It answers only for products where it has been tested on this same suite, and declines everywhere else. Its rows are not counted in the totals above.
Generated from the results archive by script — date of generation and number of full runs scanned:
2026-10-02 · 86