Sigma Logic AI Lead with AI. Thrive with Innovation.
Delivery

Pilot to production: what the second invoice covers

Why the gap between a working AI pilot and a production system is usually two to four times the pilot cost, and what that money actually buys.

On this page 11 sections
  1. Key takeaways
  2. Who this applies to
  3. What a pilot deliberately does not do
  4. Where the money goes
  5. Why “it already works” is misleading
  6. How to tell a fair estimate from a padded one
  7. The sequencing that avoids the surprise
  8. Why we structure it this way
  9. When productionisation is not worth it
  10. Frequently asked questions
  11. Next step

A pilot proves an AI approach works on your data. Production makes it survive real users, real permissions, real failure and real change. That second step typically costs two to four times the pilot, and the money goes to integration, access control, evaluation, monitoring and rollback - none of which a pilot needs.

The uncomfortable version: a pilot that impresses everyone is roughly 20-30% of the work. Teams who do not know that read the second quote as opportunism.

Pilot against production, as a share of total effort A pilot is roughly a quarter of the total work. Production is the remaining three quarters, spent on integration and data reality at 25 to 35 percent of that phase, permissions and audit at 20 to 30, evaluation at 15 to 20, monitoring at 10 to 15, and rollout and handover at 5 to 10 each. PilotProduction 20-30%70-80% What the second invoice buys, as a share of that phase Integration and data reality Permissions and audit Evaluation Monitoring, alerting and cost control Rollout, rollback and handover 25-35%20-30%15-20% 10-15%10-20%
None of the lower work existed in the pilot - that is what made the pilot fast, not what made it cheap. Permissions and audit is the most underestimated line, because in a pilot one person sees everything.

Key takeaways

  • A pilot deliberately skips the expensive parts. That is what makes it fast, not what makes it cheap.
  • Permissions and audit are usually the single largest line in productionisation.
  • If the pilot had no evaluation set, the production phase pays for one and also pays to discover what the pilot actually scored.
  • Ask for the production estimate before approving the pilot, even as a range.
  • A pilot that cannot state a number should not proceed to production.

Who this applies to

You ran a pilot - internally, or with a vendor - it worked, and you have been handed a number for making it real that is much larger than you expected. This explains what that number is for and how to tell a fair one from a padded one.

What a pilot deliberately does not do

A good pilot answers one question: does this approach produce acceptable output on our actual data? To answer it fast, it legitimately omits almost everything else.

ConcernPilotProduction
Who is allowed to see whatOne test account, or ignoredEnforced per user at query time
FailureSomeone reruns itRetries, dead letters, alerting, fallback
Wrong outputNoticed in the roomDetected, logged, escalated, reviewed
DataA clean extractLive systems, partial records, edge cases
ChangeFrozenContent drifts, models update, product changes
LoadOne person clickingConcurrency, rate limits, cost ceilings
RollbackClose the notebookFeature flags, staged rollout, kill switch
OwnershipThe person who built itA named owner and a runbook

Every row on the right is work that did not exist on the left. That is the second invoice.

Where the money goes

Rough allocation for a typical productionisation, as a share of that second phase:

Integration and data reality: 25-35%. The pilot ran on an export. Production reads live systems that contain records missing fields, duplicates, an encoding nobody documented, and a legacy status value that means something different before 2021. This is where estimates slip most.

Permissions and audit: 20-30%. The single most underestimated line. In a pilot, one person sees everything. In production, an assistant must return only what this user is allowed to see, enforced at query time rather than filtered afterwards. Getting that wrong is not a bug, it is an incident, and building it correctly means modelling identity, groups and inheritance across every source.

Evaluation: 15-20%. Building a real test set from real cases, a scoring harness, and a measured baseline. If the pilot did not have one, part of this budget is spent discovering that the pilot’s apparent quality was measured on the easy cases.

Monitoring, alerting and cost control: 10-15%. Failure detection, latency, cost per transaction, a ceiling that stops a runaway loop billing you overnight.

Rollout and rollback: 5-10%. Staged release, feature flags, a way to turn it off in seconds without a deploy.

Handover: 5-10%. Runbook, training, documentation. Skipped often; the reason systems become unmaintainable within a year.

Why “it already works” is misleading

The sentence that causes most of the friction is “but it already works”. It does - on the path it was shown.

Pilots are demonstrated on representative cases, which in practice means cases where the data is complete and the question is well formed. Production traffic includes the customer with two accounts, the invoice that was scanned at an angle, the employee who left last week whose permissions were never revoked, and the question phrased in a way nobody anticipated.

The pilot did not fail on these. It never met them.

How to tell a fair estimate from a padded one

Five checks:

  1. Is the estimate itemised by concern? Integration, permissions, evaluation, monitoring, rollout. A single number for “productionisation” is not reviewable.
  2. Does it name the systems it must integrate with, and their state? A vendor who has looked at your API documentation writes different sentences than one who has not.
  3. Does it include a measured baseline from the pilot? If the pilot’s quality was never quantified, the production estimate is partly a guess, and it should say so.
  4. Does it separate fixed from variable? Integration with a known API is fixed-priceable. Data cleanup on a system nobody has inspected is not, and pretending otherwise means someone is pricing in risk.
  5. Is running cost quoted alongside? A production quote without a monthly running figure is incomplete, because a system nobody has budgeted to operate is one nobody will operate.

A fair estimate often includes an honest unknown: “we cannot price the data remediation until we have read 500 real records; here is a fixed fee to do that and then a firm number.” That is a good sign, not evasion.

The sequencing that avoids the surprise

The reason the second invoice shocks people is almost always that the pilot was approved without a production estimate existing - which is what a scope drawn around a capability rather than a workflow produces, see scoping an AI project that actually ships.

Better sequence:

  1. Diagnose first, briefly. Establish the metric, the baseline and the rough production shape before building anything.
  2. Get a production range with the pilot proposal. Not a quote - a range, with the variables named. “Production is $40,000-90,000 depending on whether permissions must be enforced per user and whether the CRM data needs remediation.”
  3. Run the pilot with an evaluation set, even a small one. This is the difference between “it seemed good” and “it scored 68% on 200 real cases”.
  4. Decide at the boundary. With a number from step three and a range from step two, the production decision is an arithmetic problem rather than a sunk-cost argument.

We run this as four phases with a stop point at the end of each - diagnose, prove, integrate, operate - specifically so the expensive commitment is made with evidence rather than momentum.

Why we structure it this way

Our delivery model puts a decision point at the end of every phase, and the one that matters most is the boundary between prove and integrate.

That is the phase where cost multiplies, and it is the phase most likely to be entered because a demo went well in a room rather than because a number justified it. Making it a formal stop, with the evaluation report in hand, changes the conversation from “we have come this far” to “here is what it scores and here is what production costs”.

We have also had clients stop there, correctly. A pilot that scores 45% on real cases when the business case needed 70% has done its job: it cost a fraction of production and it prevented a much larger mistake. A pilot that cannot be stopped is not a pilot, it is the first instalment of a decision already made.

The corollary is that we quote the integrate phase as a range at proposal time, before the pilot runs, with the variables named. It makes our early number look larger than a competitor quoting only the pilot. That is the trade, and we would rather lose that comparison than have the month-four conversation.

When productionisation is not worth it

When the pilot scored below what the business case needs. Production makes a system reliable. It does not make it more accurate. If the pilot resolves 40% and you needed 70%, the answer is more work on the approach, or no.

When the process should be redesigned instead. Sometimes the pilot reveals that the underlying process is the problem - too many handoffs, an approval nobody needs. Automating it faithfully preserves a bad design at speed.

When the volume does not justify the fixed cost. A production system carries a permanent operating cost. Below a certain throughput a human doing it manually is genuinely cheaper and considerably more flexible.

When no one will own it. Repeated across these articles because it remains the most common cause of expensive systems quietly dying.

Frequently asked questions

Is the ratio always two to four times?

It is a common range, not a law. It compresses when the pilot already ran against live systems with real permissions. It expands when the pilot used a clean export and production must touch several messy sources.

Can we skip the pilot and go straight to production?

Occasionally, for well-understood patterns on clean data. Usually a bad idea, because the pilot is the cheap place to discover that the approach does not fit. Skipping it does not remove that discovery, it just relocates it to the expensive phase.

Should the same vendor do both?

Not necessarily, but transitions cost time. If you plan to switch, insist the pilot ships with its evaluation set and documented assumptions, otherwise the next team re-derives them at your expense.

What if the vendor will not give a production range?

Ask why. “We need to inspect the CRM data first” is a legitimate answer, and should come with a fixed fee to do that. “It depends on too many factors” without any attempt to name them is not.

How long does productionisation take?

For a system of this size, typically four to twelve weeks after the pilot. Permissions work and data remediation are the two lines that most often extend it.

Next step

If you have a pilot and no credible production number, the two-day diagnosis produces a costed roadmap you keep - including the case where the honest recommendation is to stop at the pilot.

Related: How to evaluate an AI agency proposal · Fixed price or time and materials for AI projects · How to measure whether an AI system works · How we work

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.