A deterministic feature is finished when it does the thing. A probabilistic one isn’t, and every fixed-price AI contract that pretends otherwise ends in an argument about whether a wrong answer is a bug. So we settle it in week one, before any model work.
The Acceptance Set
A frozen batch of your own real cases with the correct answers graded by your subject-matter expert, not by me. I build the harness, you grade and sign. That set is the contract’s definition of correct.
Then three numbers go in the statement of work — how often it’s right, how often it fails safely into a review queue instead of answering confidently wrong, and what one run costs at the ninety-fifth percentile. The safe-failure number is always set higher than the accuracy number, because a wrong answer is a cost and a confident wrong answer is the defect.
And if it misses
I work to threshold at no additional fee for up to three weeks. If it still misses, we stop. You keep the code, the Acceptance Set and the eval harness, and the final quarter of the fee is never billed. I’d rather eat twenty-five percent than argue with you about whether a wrong answer is a bug.
One more thing, because it’s the honest part: the eval harness runs without me. If you fire me, you can still tell whether it’s working.
The pass bands and how the money moves2 tables
The bands, published
| Shape | Pass rate |
|---|---|
| Extraction and classification | 92–97% |
| Agentic workflow with human approval — first-pass draft acceptance | 85–93% |
| Grounded assistant — correct and correctly cited | 80–90% |
| Safe failure — declines into a review queue rather than answering confidently wrong | ≥98% |
Set at signing rather than discovered in month four. The acceptance event is three consecutive green runs on three different days, then one live shadow week where the system runs alongside the human process and nothing it produces gets acted on.
How the money moves
| Milestone | Share |
|---|---|
| Kickoff, released on delivery of the signed Acceptance Set and eval harness | 35% |
| Build complete | 40% |
| Acceptance | 25% |
That last 25% is genuinely at risk on the three numbers, which is the only version of a probabilistic fixed price a finance department should ever sign.
The eval suite is yours, in your repo
Not a slide about evals — a bill of materials that lands in your repository and runs on your CI:
- The scorers, one per number in the statement of work
- Permission-leak probes, if the system knows who is asking
- A red-team set drawn from the OWASP Top 10 for LLM Applications
- A regression runner wired into CI, so a prompt change can fail a build
- A per-run cost and latency report
- A one-page scorecard your team can run forever
And a refusal that belongs here rather than in month four: if your process needs 99.5% and cannot tolerate a review queue, this isn’t an AI problem. I’ll tell you that in the AI Test Pattern — $7,500 flat, and a recommendation that may be ‘do not build this’ — not after you’ve spent a build budget finding out.
None of this is worth reading until there is something to build. That starts at the AI Test Pattern — $7,500 flat, and a recommendation that may be ‘do not build this’. Every price is on the AI page.