Where should the confidence gate sit?
A gate trades coverage against errors reaching customers, and the right setting depends on what a wrong answer costs you. Enter two settings you have measured and see what each one actually costs a week.
Working it out
Adjust the inputs on the left.
Weekly cost of each setting
| Setting | Answered | Escalated | Wrong out | Escalation cost | Error cost | Total |
|---|---|---|---|---|---|---|
| Escalate everything | - | - | - | - | - | - |
| A - looser | - | - | - | - | - | - |
| B - tighter | - | - | - | - | - | - |
Both settings need to be numbers you have measured on the same traffic. Two guesses compared against each other produce a confident answer to a question you have not asked yet. Nothing you type is sent anywhere.
How this is calculated
For each setting: answered = volume × share handled, everything else escalates. Wrong answers reaching customers = answered × (1 − accuracy). Weekly cost is escalations at your handling cost plus wrong answers at your error cost.
The escalate everything row is the honest baseline. It has zero wrong answers and the maximum human cost, and if it beats both of your settings then the gate is not earning its place at any threshold you have tried.
The break-even figure is the useful one
Cost per wrong answer is the input nobody can pin down, so the calculator inverts the question: at what error cost do the two settings cost the same? Above that number the tighter gate wins, below it the looser one does. Rather than defending a precise figure, you only have to decide which side of the break-even you are on, which is a far easier conversation to have with a team.
What this deliberately does not model
A curve. There is no accuracy-versus-threshold relationship built in, because it is specific to your system and inventing one would produce confident nonsense. You supply two points you have measured and the tool compares them.
Topic and risk rules. Anything touching money, health, legal or safety should route to a person regardless of confidence. That is a hard rule, not a score, and it sits outside this arithmetic - a confident wrong answer about a refund is exactly the case a threshold does not catch.
The cost of a bad escalation. An escalation into an unmonitored queue is worse than no gate at all, and it costs more than the handling figure here assumes.
On which signals are worth gating on - and why the model's own stated confidence is the weakest of them - see Confidence thresholds and escalation design.
Let's talk
Do not have the two measurements yet?
Measuring deflection and accuracy at a given gate is a few days of work, and it converts this from a debate into arithmetic. That is where the ai customer support agents engagement starts.