Sigma Logic AI Lead with AI. Thrive with Innovation.
Evaluation

We shipped without an eval set

What the first year looks like when nobody measured: decisions made on anecdote, changes nobody can justify, and the cost of retrofitting too late.

On this page 11 sections
  1. Key takeaways
  2. Who this applies to
  3. How it gets cut
  4. The first year
  5. What is actually lost
  6. The retrofit
  7. What to do if you are here now
  8. What we do about this in proposals
  9. When you genuinely do not need one
  10. Frequently asked questions
  11. Next step

Skipping evaluation does not produce a worse system on day one. It produces a system nobody can reason about by month six - where every quality question is settled by whoever is most senior, every change is justified by a handful of examples, and the launch-day number is the only number that has ever existed. The retrofit costs more than the original would have, and it can never recover the baseline you did not take.

Written as a pattern rather than one incident. It follows the same course often enough to describe as a sequence.

Key takeaways

  • Evaluation is cut because it produces no visible feature, not because anyone decided against it.
  • Without a baseline, “better” becomes an opinion and the loudest person wins.
  • Anecdote-driven tuning fixes the cases people complain about and breaks quieter ones.
  • The launch-day baseline cannot be recovered later. That loss is permanent.
  • Retrofitting costs more than building it, because you also relitigate every past decision.

Who this applies to

You have a system in production with no evaluation set, or you are about to accept delivery of one. Also relevant if you are weighing whether to cut it from a proposal to fit a budget.

How it gets cut

Almost never as a decision. It is 15-20% of a build, it produces nothing anyone can see, and it competes against features on a fixed budget.

The proposal says “thorough testing”. The delivery timeline gets tight. Testing becomes a person trying thirty cases by hand the week before launch, which produces a satisfying feeling and no artefact. Everyone moves on.

Note what has happened: the cost was not removed. It moved to you, later, at a worse moment. That is the same shape as the cheaper proposal that omits evaluation to look cheaper - see how to evaluate an AI agency proposal.

The first year

Month one. It works. The demo cases pass, early feedback is positive, and there is no reason to think anything is wrong.

Month two. Someone reports a bad answer. The prompt is edited to fix it. Nobody can check what else that edit affected, because there is nothing to check against. The complaint stops, so it is treated as resolved.

Month four. Four more edits have accumulated by the same mechanism. Two of them are in tension. The prompt is now long and nobody will remove anything, because nobody knows which instructions are load-bearing.

Month six. Support says it feels worse than it was. Nobody can produce evidence either way. A meeting is held. The most senior person’s impression carries, because impressions are the only input available.

Month eight. A provider ships a model update. Behaviour shifts. Nothing detects it. The eventual investigation cannot establish whether quality dropped, because there is no before.

Month ten. Someone proposes switching to a cheaper model to reduce spend. The conversation cannot resolve - nobody can quantify what would be given up - so it is either blocked by fear or approved on hope.

Month twelve. Budget review asks whether the system is working. The honest answer is that nobody knows.

Each step is reasonable in isolation. The compounding effect is a system that cannot be improved, defended or safely changed.

What is actually lost

Three things, and one of them is permanent.

The ability to change anything safely. Every modification is a gamble. Teams respond by freezing the system, which means it stops improving and starts drifting - see the prompt that worked until the input changed.

The ability to detect external change. Model updates and input drift both arrive without any signal on your side. You find out through complaints, weeks late.

The baseline. Permanently. This is the one that cannot be recovered. You can build an evaluation set today and measure current performance. You cannot measure what it was at launch. So the question everyone eventually asks - has this got worse - has no answer, ever. That is the real cost of skipping it, and it is invisible at the time.

Measurement started at month twelve cannot recover the launch figure A system's true quality varies across its first year, but with no evaluation set nothing is recorded. Building the set at month twelve produces a first data point at month twelve. Every earlier month remains unmeasurable, so the question of whether quality got worse has no answer.

What you can know, and when you started looking

launchmonth 4 month 8month 12

True quality - real, and unrecorded

First measurement

TWELVE MONTHS THAT CANNOT BE RECONSTRUCTED

You can build the set today and learn where you are. You cannot learn where you were, so “has this got worse” never gets an answer.

Two of the three things lost by skipping evaluation can be bought back later at a cost. The baseline cannot - and it is the one nobody misses until the first budget review asks whether the system is working.

The retrofit

It is worth doing. It costs more than the original would have, and here is what it actually involves.

Build the set from real cases, stratified, including cases the system should refuse. Three to five days of domain-expert time. Same work as it would have been - see building an eval set from real tickets.

Measure where you are. This is the uncomfortable step. The number is frequently lower than everyone believed, and there is no earlier number to compare it against, so it lands as bad news rather than as a diagnosis.

Establish run-to-run variance, so future changes can be judged.

Untangle the prompt. The genuinely additional cost. A year of accumulated instructions, several contradictory, and now you can finally test which matter - remove a group, run the set, observe. Half a day of this typically removes a third of the prompt, and it is only possible once the set exists.

Relitigate the deferred decisions. The model switch that was blocked, the threshold nobody would touch, the instruction everyone suspected was harmful. Each becomes a measurable question, and each has an accumulated cost from having been unanswerable.

That last item is the hidden expense. It is not just building the missing artefact; it is working through a backlog of decisions that were parked because they could not be resolved.

What to do if you are here now

In order.

  1. Build the set. Do not wait for a project to justify it.
  2. Measure and write the number down. Today’s baseline is worth having even though it is late. It is the earliest one you will ever have.
  3. Establish variance before judging any change.
  4. Add a case for every production failure from now on.
  5. Then untangle the prompt, with the set as your safety net.

And resist the temptation to fix things first. Measuring after fixing gives you a number with nothing to compare it to, which is how you arrive back here in another six months.

What we do about this in proposals

Evaluation is a named line item with an owner and a deliverable, not a paragraph about methodology, and we quote it as part of the build rather than as an option.

The reason is precisely that it is easy to cut a paragraph and awkward to cut a line item with a name. That is a deliberate proposal-writing choice, and it makes our number look higher against a competitor whose testing is a sentence.

When a client wants it removed to fit a budget, our position is that we would rather reduce scope elsewhere - fewer intents, one channel instead of two - than ship something unmeasurable. A narrower system you can improve beats a broader one you cannot. We have lost work over this, and we have also been called back later to do the retrofit described above, which is a poor outcome for the client even though it is more revenue.

The claim we are careful not to make: evaluation does not make the system better on day one. It usually makes the launch number look worse, because measuring properly surfaces failures that thirty hand-checked cases missed. What it buys is every subsequent month.

When you genuinely do not need one

Prototypes that may be discarded. Thirty cases in a spreadsheet is right for deciding whether an approach is viable.

Deterministic systems. If output correctness is programmatically checkable, you need tests, not an evaluation set.

Very low stakes and very low volume, where a human reads every output and would notice immediately. They are the evaluation.

If none of those describe you, the set is not optional - it is deferred.

Frequently asked questions

How much does retrofitting cost?

The original three to five days of labelling, plus the prompt untangling and the backlog of deferred decisions. Call it twice the original, and the permanent loss of the launch baseline on top.

Can we reconstruct the original baseline from logs?

Partially, if you retained inputs and outputs and can score them now against a rubric you write today. Better than nothing, and not the same as having measured at the time with a rubric agreed then.

Is it worth building one for a system we might replace?

If the replacement decision is within weeks, no. If it is a maybe-next-year, yes - and the set transfers to the replacement, which is the strongest argument for building it now.

What if the retrofit number is bad?

That is information, and it is the most useful information you have had about the system. It also gives you the baseline to improve from, which is the point.

How do we stop this happening again?

Make the evaluation run a scheduled deliverable with an owner, not a good intention. The set that gets neglected is the one nobody is accountable for.

Next step

If you have a system with no measurement, building the set is a few days and it converts every future quality argument into a question with an answer. The AI evaluation and QA engagement builds it from your real cases with a measured baseline and CI integration.

Related: Building an eval set from real tickets · How to measure whether an AI system works · How to evaluate an AI agency proposal · AI evaluation and QA

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.