Sigma Logic AI Lead with AI. Thrive with Innovation.
Buying AI

What we turn down, and why

Six requests we decline, the reasoning for each, and what we propose instead. Published so you can hold us to it, and test anyone else the same way.

On this page 8 sections
  1. Key takeaways
  2. Who this applies to
  3. The six
  4. What a good refusal looks like
  5. The one we get wrong most often
  6. When to ignore a refusal list
  7. Frequently asked questions
  8. Next step

An agency with no refusal rules will build whatever you ask for, which sounds like service and is the single most reliable way to end up with a system nobody can defend. These are the six requests we decline, the reasoning in each case, and what we propose instead. Most of them are declines that cost us the larger engagement, which is the only reason the list is worth anything.

Ask any supplier what they refuse to build. The answer, or the absence of one, is more informative than the case studies.

Key takeaways

  • A refusal list is only credible if the items on it are commercially expensive. Ours are.
  • Four of the six are declined on evidence: no baseline, no owner, no data, no agreed process.
  • Two are declined on principle, and would stay declined at any price.
  • Every decline comes with the smaller thing we would do instead, because a refusal without an alternative is just a lost meeting.
  • Use these as tests on whoever you are considering, including us.

Who this applies to

You are choosing a supplier, or you are one. The commercial framing of the same subject is how to evaluate an AI agency proposal, which is about reading what you are sent. This is the other side of the table, written down so it can be checked.

The six

1. A build with no way to tell whether it worked

If nobody can say what number this is supposed to move, or what that number is today, we will not start the build.

This is the most frequent decline and the least popular. The system will get built, it will produce output, and in six months there will be no way to establish whether it helped, which means no way to defend the spend and no way to justify the next phase. The whole argument is in we shipped without an eval set.

Instead: a short measurement engagement first. Establish the baseline, then decide what to build. It is a smaller invoice and it makes the second one arguable.

2. A system with no named owner after launch

If the answer to “who runs this in six months” is a shrug, or a team rather than a person, we say no.

An AI system is not a deliverable, it is a thing that needs someone to edit content, approve threshold changes, respond to alerts and read the escalation queue. Without a name, the system decays quietly and the failure lands on whoever is nearest a year later. The roles are set out in what an AI roadmap should actually contain.

Instead: we scope the version a real person can own, which is usually narrower, or we help make the case for the role first.

3. Automating a process nobody has agreed on

If three people describe the process differently, automation does not resolve that. It picks one of them and enforces it at speed, and the disagreement resurfaces as a system defect.

Instead: map the process, write it down, get the three people to sign the same page. Then automate. This is often a week, and it is a week that would otherwise be spent arguing about a build.

4. A knowledge assistant over knowledge that is not written down

If the answers live in people’s heads, retrieval has nothing to retrieve, and the model will answer anyway from general knowledge. Confident, plausible, wrong, at scale. This one we decline firmly because the failure lands on customers.

Instead: the documentation project, which is the cheaper half and delivers value with or without an assistant afterwards. The mechanics of what goes wrong are in why retrieval misses the answer that is there.

5. Anything that pretends to be a person

A system that claims to be human when asked, invents a name to build rapport, or is deliberately designed so customers cannot tell. This is a principle rather than a judgement call, and it does not move on budget.

The commercial argument is on the same side anyway: the customers who find out are the ones who were paying enough attention to matter, and they tell people. But we would decline it if the commercial argument pointed the other way.

Instead: disclosure in the first message, which costs nothing and removes the thing people are otherwise testing for.

6. Deceptive or manufactured content

Fake reviews, fabricated case studies, invented statistics, synthetic testimonials, or volume content designed to appear independent. Also declined on principle.

We are strict about this on our own site, which is why there are no client logos and no outcome figures on it. We have not published case studies because we do not yet have client-approved numbers to publish, and inventing them is not an option that exists. That is a visible commercial cost of the rule, on our own home page, and it is the most honest evidence we can offer that the rule is real.

Instead: publish the method rather than the outcome. It is what this entire corpus is, and it is why the first article built on our own measurement is a visibility baseline rather than a client result.

Six requests declined, and on what basis Six requests we decline. Four are declined on evidence: a build with no way to tell whether it worked, a system with no named owner after launch, automating a process nobody has agreed, and a knowledge assistant over knowledge that is not written down. Two are declined on principle: anything that pretends to be a person, and deceptive or manufactured content. The first four move if the evidence turns out to be there. The last two do not move, and would not move on budget. Four declined on evidence, two on principle No way to tell whether it worked evidence No named owner after launch evidence A process nobody has agreed evidence Knowledge that is not written down evidence Anything pretending to be a person PRINCIPLE Deceptive or manufactured content PRINCIPLE The first four move if you show us the evidence is there. The last two do not move, and would not move on budget.
The distinction matters when you are on the other side of it. Ask a supplier which of their refusals are evidence and which are principle, because only one kind is negotiable.

What a good refusal looks like

Three properties, and they are worth applying to anyone.

It is specific. “We only take projects that are a good fit” is not a refusal. “We do not start a build without a baseline number” is.

It has a cost. A refusal that loses nothing is a preference. Ask a supplier for one that cost them a deal, and listen for whether the story has a real shape.

It comes with an alternative. Turning down work and leaving the client where they were is not integrity, it is inconvenience. Every item above has a smaller thing attached.

Three tests for whether a refusal list is real Three tests to apply to anyone’s refusal list. It is specific, because saying you only take projects that are a good fit is not a refusal. It has a cost, so ask for one that lost them a deal. And it comes with an alternative, or it is just an ended meeting rather than integrity. A refusal that costs nothing is a preference, so listen for whether the story that cost them a deal has a real shape. A list of this shape is easy to write, and the items are what make it worth anything. Three tests for anyone’s refusal list, including ours It is specific "a good fit" is not a refusal It has a cost ask for one that lost a deal It comes with an alternative or it is just an ended meeting A refusal that costs nothing is a preference. Listen for whether the story that cost them a deal has a real shape. A list this shape is easy to write. The items make it worth something.
The second test does most of the work. Specificity can be drafted; a lost deal has details that are hard to invent on the spot.

The one we get wrong most often

We are inconsistent about the volume threshold.

Below a few hundred support contacts a month, or a handful of runs a week for an automation, the honest answer is usually to do nothing yet, and we do say it. But we have taken on work just above that line where the arithmetic was thin, because the client was persuasive about growth that had not happened yet. Twice that growth did not arrive, and the client owned a system that could not repay its own maintenance.

The rule we now try to hold: size the case on the volume that exists, not on the plan. If the case only works at three times today’s volume, the honest recommendation is to wait, and to spend the money on getting to three times.

What the no-invented-content rule costs on our own site What is not on this site: client logos, outcome percentages and case studies, which are the usual proof. What is here instead: the method in full, our own measurement, and opinions with their commercial price named, which is what we can evidence. We have no client-approved numbers yet and inventing them is not an option, so the absence is a fact about our stage rather than a claim about our record. It is also the most honest evidence we can offer that the rule is real. What the content rule costs us, visibly What is not on this site Client logos Outcome percentages Case studies THE USUAL PROOF What is here instead The method, in full Our own measurement Opinions with their price WHAT WE CAN EVIDENCE No client-approved numbers yet, and inventing them is not an option. The absence is a fact about our stage, not a claim about our record. It is also the most honest evidence we can offer that the rule is real.
Every agency site claims integrity. This is the version of that claim that can be checked, by looking at what is missing from the page you are on.

When to ignore a refusal list

When it is generic. A list this shape is easy to write and costs nothing unless the items are real. Test them with your own situation and see whether the answer changes.

When your circumstances genuinely differ. Every rule here has conditions. A regulated buyer with a compliance deadline has different constraints from a founder with a hunch, and the same advice does not fit both.

When it is being used to avoid work. Refusals can be laziness wearing principle. The test is whether the alternative offered is real, specific and smaller, or whether the conversation simply ends.

Frequently asked questions

Do you turn down much work?

Enough that it is a real category rather than a marketing position, and the most common decline by far is the first one on the list. Most of those conversations end with a measurement engagement instead of a build, and some end with nothing.

What if we insist?

For the four evidence-based items we will state the concern once, in writing, and then build what you have decided, because it is your business and you may know things we do not. For the two principles, we will not, at any price.

Is this not just an argument for smaller projects?

It is an argument for projects that can be defended afterwards. Several of them get larger later, because a system with a baseline and an owner is one you can justify extending.

How do we use this list on another supplier?

Ask what they refuse to build, then ask for an example that cost them money. The specificity of the second answer tells you whether the first was real.

Would you publish a project that failed?

We would, with the client’s agreement, and we do not have one to publish yet. Until then the honest position is that the absence of case studies on this site is a fact about our stage, not a claim about our record.

Next step

If you want a supplier to tell you when not to build something, that is a reasonable thing to test before you engage one. AI consulting and strategy starts with the question of whether the thing should exist, and sometimes ends there.

Related: How to evaluate an AI agency proposal · We shipped without an eval set · When the answer is not a chatbot · Responsible AI as an engineering decision · AI consulting and strategy

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.