Forge Agents ← Home
AI evaluation

We change a prompt and hope nothing broke. I want a test suite that tells us.

Test cases built from real use, run on your live system, and a written report of what misses and why.

Who it is for

Startups and product teams shipping LLM features: assistants, retrieval and search, support bots, agents. You have users, you are changing prompts and models, and you do not have a test suite you trust.

What you get

The working example

We test our own product this way, and the results are public. Our proof page shows 9 suites and 310 tests, each run 5 times, all passing. The backup model is shown separately with its one miss. We publish the miss because a report that hides failures is not worth reading. Your report lists misses just as plainly.

How it works

STEP 1

Scope call

What the feature does, what a wrong answer costs you, and how we reach your system: an endpoint, staging, or a harness you provide.

STEP 2

Build the cases

We draft cases from real use. You approve the expected behavior for each one.

STEP 3

Calibrate the grader

We check the grader against known good and bad answers, including rewordings and refusals, before we trust it.

STEP 4

Run and report

5 trials per case on your live system, a written report, and a walkthrough call. About two weeks in total.

Price

$2,500Eval starter, fixed. Up to 50 cases, grader, 5-trial run on your live system, written report. About two weeks.

Who does the work

Dustin Aldridge, founder of Forge Agents in Saint Paul, Minnesota. Army veteran with nine years in IT. He builds and tests Forge’s grounded assistants, which quote their source and say so when the documents do not cover a question. We are a small team, and we will tell you plainly what we can and cannot test.

Questions

Do you need our production data?

No. We can work from staging and from cases you describe. If you share real examples, remove personal data first.

Which models and stacks do you work with?

We need a system we can call with a test input and read the answer from. The grader does not depend on your model vendor.

Can we run the suite ourselves afterward?

Yes. You keep the cases, the grader, and instructions to rerun them. The monthly plan is there if you would rather we do it.

Talk to us

Book a 20-minute call or email daldridge@forgeagentsio.com with a short description of your feature.

Book a 20-minute call Email daldridge@forgeagentsio.com