We change a prompt and hope nothing broke. I want a test suite that tells us.
Test cases built from real use, run on your live system, and a written report of what misses and why.
Who it is for
Startups and product teams shipping LLM features: assistants, retrieval and search, support bots, agents. You have users, you are changing prompts and models, and you do not have a test suite you trust.
What you get
- Up to 50 test cases built from real use: your logs, tickets, or real questions, plus the edge cases your team already worries about.
- A grader that checks substance. It is built to accept correct rewordings, recognize a correct refusal, and fail a confident wrong answer, and we calibrate it before we rely on it.
- Every case run 5 times on your live system, so flaky behavior shows up.
- A written report: each miss, the likely cause (retrieval, prompt, model, or missing content), and a ranked list of fixes.
- The suite and grader files, yours to keep and rerun.
The working example
We test our own product this way, and the results are public. Our proof page shows 9 suites and 310 tests, each run 5 times, all passing. The backup model is shown separately with its one miss. We publish the miss because a report that hides failures is not worth reading. Your report lists misses just as plainly.
How it works
Scope call
What the feature does, what a wrong answer costs you, and how we reach your system: an endpoint, staging, or a harness you provide.
Build the cases
We draft cases from real use. You approve the expected behavior for each one.
Calibrate the grader
We check the grader against known good and bad answers, including rewordings and refusals, before we trust it.
Run and report
5 trials per case on your live system, a written report, and a walkthrough call. About two weeks in total.
Price
- Monthly regression runs: $500 a month. We rerun the suite each month and report what changed.
- Custom work: $95 an hour. More cases, new suites, or help fixing what the report found.
Who does the work
Dustin Aldridge, founder of Forge Agents in Saint Paul, Minnesota. Army veteran with nine years in IT. He builds and tests Forge’s grounded assistants, which quote their source and say so when the documents do not cover a question. We are a small team, and we will tell you plainly what we can and cannot test.
Questions
Do you need our production data?
No. We can work from staging and from cases you describe. If you share real examples, remove personal data first.
Which models and stacks do you work with?
We need a system we can call with a test input and read the answer from. The grader does not depend on your model vendor.
Can we run the suite ourselves afterward?
Yes. You keep the cases, the grader, and instructions to rerun them. The monthly plan is there if you would rather we do it.
Talk to us
Book a 20-minute call or email daldridge@forgeagentsio.com with a short description of your feature.