11 articles
Measuring whether it works
Evaluation sets, baselines, confidence thresholds and latency budgets. The measurement work that separates a system you can improve from one you can only have opinions about.
-
We shipped without an eval set
What the first year looks like when nobody measured: decisions made on anecdote, changes nobody can justify, and the cost of retrofitting too late.
-
Testing an AI system against the human baseline
How to measure what your people currently achieve on the same cases, why that number is usually lower than everyone assumes, and how to compare honestly.
-
Latency budgets for conversational AI
Where the seconds actually go in a grounded answer, what users tolerate in chat against voice, and how to buy back time without losing accuracy.
-
The prompt that worked until the input changed
A failure pattern where nothing on your side changed: the inputs drifted, the prompt stayed, and accuracy fell for months before anyone connected the two.
-
Regression testing prompts like code
How to put prompts under version control and CI so a change that improves one case cannot silently break nine others.
-
Your provider changed the model - what broke
A post-mortem pattern: the AI system that degraded without a deploy. Why model updates break working systems, how to detect it, and what to pin.
-
Hallucination rate: what number is acceptable
Why a single hallucination rate is the wrong target, how to measure the fabrications that actually cost you money, and what thresholds are defensible by category.
-
Golden datasets: how big is big enough
Why 150-400 cases is usually right, what the confidence interval actually looks like at each size, and why a smaller maintained set beats a larger neglected one.
-
Building an eval set from real tickets
How to turn a support queue into a test set that catches regressions: sampling, stratification, labelling, and the cases most teams forget to include.
-
Confidence thresholds and escalation design
Why model confidence is a poor escalation signal on its own, what to combine it with, and how to tune the gate that decides whether a customer meets a person.
-
How to measure whether an AI system works
A practical method for evaluating a production AI system: choosing the metric, building the test set, setting a baseline, and knowing when a score has moved.
The other folders
- What will this cost, and who should build it? Buying and budgeting AI 10 articles
- Which tool, and why do these things break? Workflow automation 13 articles
- How should the system be put together? Building and running it 11 articles
- What do we have to be able to show? Governance and regulation 4 articles
- Why does ChatGPT name a competitor? AI search visibility 4 articles
Let's talk
Prefer a conversation to an article?
Most of what is written here started as a question a client asked on a call. If you have one, bring it - we answer questions before contracts.