FORGE AGENTS

Paper

Refuse or Cite

A source-grounded assistant for front-desk procedures, and the test apparatus that keeps it honest. Dustin Aldridge, Forge Agents LLC. Version 2.0, September 26, 2026.

Read the paper (PDF, 13 pages, 197 KB). Archived on Zenodo: DOI 10.5281/zenodo.22985559.

Abstract

A hosted assistant answers front-desk questions only from an office's own procedure documents. Every answer either quotes a passage checked against the named document, says the answer is not documented and routes the question to a person, or asks one clarifying question. It serves on Gemini 3.8 Flash, with GPT-5.6 Terra as the backup for the dental and general-office products. All testing used five demonstration corpora that Forge Agents wrote for representative offices.

Version 1.0 reported 310 of 310 tests for the serving model and 110 of 110 for the backup, graded under test definitions that had been edited that same day to accept that day's answers. For this version, every test, the grader, the prompts and the corpora were frozen under a hash manifest before any new answer existed, and none of them changed afterwards. Fresh answers from the serving model passed 306 of 310 tests at three trials each (923 of 930 trials). Of the seven failed trials, five were correct answers the grader missed on wording and two were a minor model error. Under the tests as they stood before the September 20 edits, the same fresh answers pass 296 of 310. The backup passed 107 of 110 fresh.

This version also corrects version 1.0's prompt-evolution result: the loop never served its candidate prompts. Served for the first time, the candidate version 1.0 reported as an improvement scored 36 of 38 against 38 of 38 for the production prompt and was not adopted. The dental knowledge base was padded with other offices' procedures up to 250 documents and still met a capacity rule written before the run, at three trials per test. That rule allows one known miss, which recurs at every size from 50 documents up. A separate judge model found 2 of 3,659 factual claims in fresh answers supported by no document.

What is in it

Version 1.0, measured again under frozen tests.

The results open with one table that sets what version 1.0 reported beside what its runs actually recorded, the same answers regraded under the frozen tests, the fresh answers, and the fresh answers under the tests as they stood before version 1.0's edits. The paper then covers the serving and backup models suite by suite, the corrected prompt-loop result, retrieval, capacity, claim-level entailment, prompt injection, the failover drill and the public proof page. An appendix explains, topic by topic, each statement of version 1.0 that the new measurements correct.

Records

The scores are public; the records are on request.

The public scores, with every failing test named and a note on each, are on the proof page, which is rebuilt hourly. The stored answers, frozen tests, prompts and corpora behind the paper are available on request. The founder's account of why the product is built this way is on the founder page.

To cite: Aldridge, D. (2026). Refuse or Cite: A Source-Grounded Assistant for Front-Desk Procedures and the Test Apparatus That Keeps It Honest. Version 2.0. Forge Agents LLC, Saint Paul, Minnesota. https://doi.org/10.5281/zenodo.22985559

Version 1.0 of September 20, 2026 remains available as a PDF and on Zenodo as 10.5281/zenodo.22866053. The DOI for all versions is 10.5281/zenodo.22866052; it always resolves to the newest version.