Sigma Logic AI Lead with AI. Thrive with Innovation.
Buying AI

How to evaluate an AI agency proposal

Fourteen questions that separate an AI proposal that will ship from one that will produce a demo, plus the answers that should end the conversation.

On this page 10 sections
  1. Key takeaways
  2. Who this applies to
  3. The four things that actually predict success
  4. Fourteen questions, and what the answers mean
  5. Red flags that are commonly mistaken for credentials
  6. The proposal structure that usually indicates experience
  7. How we structure our own, and why
  8. When you should not run this evaluation at all
  9. Frequently asked questions
  10. Next step

Judge an AI proposal on four things: whether it names a target metric and a current baseline, whether evaluation is a deliverable rather than a promise, whether it says what happens after launch, and whether it tells you what it will not do. Proposals that skip all four are demos with an invoice attached, regardless of price.

Most AI proposals are hard to compare because they describe technology rather than outcomes. The questions below make them comparable, and several of them will visibly change how a vendor talks to you.

Key takeaways

  • A proposal without a baseline measurement cannot be held to a result.
  • “We will fine-tune a model” in a proposal is usually a red flag, not a credential.
  • The maintenance section tells you more about a vendor than the methodology section.
  • Ownership of prompts, evaluation sets and integration code should be stated, not assumed.
  • A vendor who cannot describe when you should not do the project has only one answer available.

Who this applies to

You have one to four proposals for an AI build - a support agent, a document pipeline, an internal assistant, a forecasting model - somewhere between $10,000 and $150,000, and they are hard to compare because each describes a different thing.

This is not about procurement process for enterprise software. It is about telling a team that has shipped this before from one that has shipped a demo before.

The four things that actually predict success

1. A named metric and a measured baseline

The single strongest predictor. A proposal should say what number is expected to move, what it is today, and how it will be measured after.

Good: “Current first-contact resolution on billing tickets is 41% measured across 90 days of history. Target 65% on the same intents, measured on a held-out set of 300 real tickets.”

Bad: “Improve customer experience and reduce support workload.”

The second is not a softer version of the first. It is a proposal that cannot fail, which means it also cannot succeed. A project without a target metric does not get cancelled when it stops working - it just continues.

If no baseline exists, that is fine, but the proposal should include establishing one as the first deliverable and price it.

2. Evaluation as a deliverable

Ask directly: is an evaluation set built from our own real cases, and does it ship to us?

This matters more than any architectural choice in the document. An evaluation set is what converts “it seems to be working” into a number, and it is the only mechanism that catches the silent regressions that arrive when a provider updates a model underneath you.

It also costs real money - typically 15-20% of a build. Which is precisely why proposals leave it out to look cheaper. A quote without evaluation is not less expensive; it has moved that cost to you, later, at a worse time.

See how to measure whether an AI system works for what a real one contains.

3. What happens after launch

The maintenance section is the most revealing part of any AI proposal, and it is usually the shortest.

Things that should be there: who monitors it, what alerting exists, how often the evaluation runs, what happens when the model provider ships a new version, what the incident path is, and what it costs per month. Ongoing support for a mid-sized system typically runs $1,000-3,000 a month.

A proposal that ends at go-live is describing a handover of a system nobody has agreed to keep alive. Systems like this degrade. Your content drifts, your product changes, the model changes. Six months of no ownership is usually enough to make one quietly unusable.

4. An explicit statement of what it will not do

Every real AI system has a scope boundary, and the vendors who have shipped several know exactly where theirs is. Ask what the system will not handle, and what happens when it meets those cases.

A confident, specific answer - “it will not handle disputes involving more than one order, those route to a person with the full thread attached” - is evidence of experience. A vendor who says it will handle everything has either not built one or is not telling you about the cases where it fails.

The four things in a proposal that predict whether the project works Four gates in a row: a named metric with a measured baseline, an evaluation set as a deliverable, a plan for after launch, and an explicit statement of what the system will not do. Under each, the phrase you get instead when it is absent: efficiency, we test a lot, launch is the end, handles it all. Named metric and baseline IF ABSENT, YOU GET: "EFFICIENCY" Evaluation set as a deliverable IF ABSENT, YOU GET: "WE TEST A LOT" A plan for after launch IF ABSENT, YOU GET: LAUNCH = END What it will not do IF ABSENT, YOU GET: "HANDLES IT ALL" FOUR YES ANSWERS: READ THE REST. ONE NO: THE REST IS DECORATION.
Everything else in a proposal is negotiable. These four are the difference between a build and a demo, and their absence is not something the fourteen questions below can repair.

Fourteen questions, and what the answers mean

QuestionAnswer that reassuresAnswer that should worry you
What metric moves, and what is it today?A number, a measurement window, a method“Efficiency”, “experience”
Is an evaluation set included and do we keep it?Yes, built from our cases, ours to keep“We test thoroughly”
What is the escalation or exception path?Specific triggers and destinations“It handles the full range”
Who owns the prompts and code afterwards?We do, in our repositorySilence, or “our platform”
What is the monthly running cost at our volume?Arithmetic shownA single figure with no working
What happens when the provider changes the model?Eval suite re-runs, regression triggers a fix“We monitor closely”
Which systems will it write to?A named list of actions“Full integration”
How is PII handled and where does data go?Named subprocessors and a data flow“It is secure”
What is the rollback plan?Feature flag, staged rollout, kill switchNot addressed
What would make you tell us not to proceed?Concrete conditions“We are confident this fits”
Who exactly will do the work?Named people, availability“Our team of experts”
What is the fixed scope vs the variable part?Explicit boundary and change processEverything fixed, or everything hourly
Can we see a system you built running?Yes, or a detailed anonymised accountOnly a demo environment
What went wrong on your last project?A real answer“Nothing significant”

That last question is worth more than it looks. Everyone who has shipped three of these has a story about a model version change, a permissions bug, or a process they automated that should have been redesigned instead. A vendor with no failures has either not shipped, or is willing to be untruthful in a sales conversation, and both are disqualifying.

Red flags that are commonly mistaken for credentials

“We will fine-tune a custom model for you.” Sometimes correct, usually not. For most business tasks, retrieval against your own content plus a good general model outperforms fine-tuning, costs far less, and updates instantly when your content changes. Fine-tuning is appropriate for narrow format or tone problems, and for genuinely specialised domains. If a proposal opens with it as the headline, ask why retrieval was rejected. A vendor who cannot answer is selling effort, not judgement.

A very large model list. Naming six providers signals flexibility to a buyer and indecision to an engineer. Ask which one they would use for your case and why.

Accuracy claims without a denominator. “95% accurate” is meaningless without knowing on what, measured how, over which cases. Accuracy on the easy 60% of your tickets is not a system.

A fixed price with no discovery. Confidence before information is not a good sign. The credible pattern is a small paid diagnosis, then a fixed price for a known scope.

Timelines in days for production systems. A demo takes days. A production system with permissions, evaluation and rollback does not.

The proposal structure that usually indicates experience

Not a rule, but a pattern worth recognising. Proposals from teams who have delivered several of these tend to be organised around risk rather than around technology:

  1. What we understand the problem to be, in your language
  2. The metric, the baseline and how we will measure
  3. Phase one, with a decision point at the end
  4. What we will not do in phase one, and why
  5. What it costs to run, monthly, with arithmetic
  6. What you own at the end
  7. The conditions under which we would advise stopping

Proposals organised around architecture diagrams and model comparisons are often written by people who have built the interesting part but not the part that survives.

How we structure our own, and why

We sell a two-day fixed-fee diagnosis before any build quote, and it is worth being direct that this is partly a commercial decision and partly a technical one.

The technical reason: the three things that most affect price on an AI build - the state of your content, how many systems must agree, and whether the system may write as well as read - are not reliably knowable from a sales call. A fixed price quoted without them is either padded to cover the unknown or will be revised.

The commercial reason: it filters. Companies unwilling to spend a small fixed fee to find out whether a project is worth doing are usually not ready to run one.

What you keep from it is the ranked opportunity map and the costed roadmap, whether or not you continue with us. That is deliberate. If the roadmap is only usable by the people who wrote it, it is a lock-in device rather than a deliverable.

We will also tell you when the answer is not to build. The most common versions are: your volume is too low for the economics to work, the problem is upstream of where you want to apply AI, or an off-the-shelf product does 80% of this for a tenth of the cost.

When you should not run this evaluation at all

When you have no internal owner. Before comparing proposals, identify the person who will own this system after launch. If there isn’t one, the evaluation is premature - the best proposal in the world produces a system that decays within two quarters.

When the requirement is not yet a requirement. If the brief is “we should do something with AI”, no proposal can be evaluated because there is no criterion. Do the diagnosis work first, internally or paid.

When the budget only covers the build. If there is no room for the running cost, do not start. A system you cannot afford to maintain is a liability you paid for.

Frequently asked questions

Should we ask for a fixed price or time and materials?

Fixed for a known scope, time and materials for genuine research. The trap is fixed price on an unknown scope, which prices the vendor’s risk into your invoice. See fixed price or time and materials for AI projects.

How many vendors should we ask?

Two or three. More than that and the comparison cost exceeds the price difference, and you will start selecting on document quality rather than delivery capability.

Is the cheapest proposal ever the right one?

Yes, when it is cheaper because the scope is genuinely smaller and it says so. No, when it is cheaper because evaluation, maintenance and integration have quietly left the document. Normalise the scope before comparing the numbers.

What if we already have an internal team?

Then the question is different: what specifically is missing - capacity, a skill, or an outside view? A proposal that does not distinguish those is not aimed at your situation. A fractional AI lead is often the better shape here than a build engagement.

Should the vendor sign our security questionnaire before proposing?

For anything touching customer data, yes, and their willingness to do it early is itself a signal. Vendors who have passed several will have the answers ready.

Next step

If you want an independent read on proposals you already have, that is a normal use of the AI strategy engagement - including when the recommendation is to accept one of them rather than to work with us.

Related: Build or buy AI customer support · Fixed price or time and materials for AI projects · What you own after an AI project · AI consulting and strategy

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.