Sigma Logic AI Lead with AI. Thrive with Innovation.
Decision systems

Lead scoring that sales actually trusts

Most scoring models are ignored within a quarter, and the cause is never the algorithm. It is no feedback from closed deals, no explanation, no definition.

On this page 12 sections
  1. Key takeaways
  2. Who this applies to
  3. The three failures, in order of how often they cause the problem
  4. Fit and intent are two questions
  5. Build the feedback loop before the model
  6. Measure the thing that actually moves revenue
  7. The score is not the intervention
  8. The enrichment trap
  9. What we build and what we push back on
  10. When this is not worth it
  11. Frequently asked questions
  12. Next step

A lead score is only worth building if somebody changes what they do because of it. Most do not survive a quarter, and the cause is never the algorithm - it is a score fitted to a definition of “good lead” nobody agreed, never corrected by what actually closed, and delivered without a reason attached. Fix those three and a simple model outperforms a sophisticated one nobody follows.

Key takeaways

  • A score that cannot be explained gets overridden, and every override is a signal you then throw away.
  • Fit and intent are different questions and averaging them into one number hides which is missing.
  • Without a feedback loop from closed-won and closed-lost, the model is frozen at the day it was built.
  • Response time usually moves conversion more than scoring accuracy does, and it is cheaper to fix.
  • Below roughly 300 closed deals a year you do not have the data for a model. Use a rule and say so.

Who this applies to

You have inbound volume that outstrips the sales team’s capacity to work all of it, and you are deciding what to prioritise. Also relevant if you already have a score nobody looks at.

The three failures, in order of how often they cause the problem

Scoring is a decision system, and it fails the way decision systems fail - see better decisions, not just faster ones for the general version. Three specific shapes account for most of it.

Nobody agreed what “good” means. Marketing scores for likelihood to convert to an opportunity. Sales cares about likelihood to close, at a reasonable size, this quarter. Those are different targets and a model can only be fitted to one. When the definition is not settled first, the model is quietly optimising for a metric the people using it do not care about.

The score has no reason attached. “84” tells a rep nothing. “84 - matches your best-fit industry, visited pricing twice this week, company size in range” tells them how to open the call. Explanation is not a nicety here: it is what makes the score usable rather than merely present.

There is no loop back from outcomes. The model was fitted once, on deals that closed before the product changed and before the current segment mix existed. It is now answering last year’s question with total confidence.

Fit and intent are two questions

Collapsing them into a single number is the most common structural mistake, because the two combinations that look identical at 60 need opposite responses.

Low intentHigh intent
High fitRight company, not looking. Nurture, and do not burn a call on it yetWork this now. This is the entire reason the system exists
Low fitIgnore. Both signals agreeThe dangerous quadrant: a student, a competitor, or somebody researching. Looks urgent, converts at almost nothing

A single 0-100 score puts the top-right and bottom-right cells in the same band, and the bottom-right one is where reps lose faith. They work three of them, none convert, and the score stops being read. Reporting two axes costs nothing and prevents that specific loss of trust.

Fit and intent as two separate questions A two by two grid. High fit and low intent: nurture, right company but not looking. High fit and high intent: work this now, the reason the system exists. Low fit and low intent: ignore, both signals agree. Low fit and high intent: the dangerous quadrant, a student, a competitor or a researcher, looks urgent and converts at almost nothing. LOW INTENT HIGH INTENT HIGH FIT LOW FIT Nurture Right company, not looking. Do not burn a call on it yet Work this now The entire reason the system exists Ignore Both signals agree The dangerous quadrant A student, a competitor, a researcher Looks urgent, converts at nothing
A single score collapses the two axes and the bottom-right quadrant is where it goes wrong. High intent from the wrong company is the lead reps learn to distrust, and then they distrust the score.

Build the feedback loop before the model

This is the part that gets deferred and the part that decides whether the thing survives.

Write every score to the CRM record at the time it was given, not just the current score. You need to know what the model said when the rep acted, otherwise you cannot evaluate anything later.

Capture the outcome against it. Closed-won, closed-lost with reason, and disqualified are three different things and lumping them together destroys the signal - a lead disqualified as out-of-territory says nothing about scoring quality.

Capture the override. When a rep works a lead the model ranked low, or skips one it ranked high, that is a labelled example arriving free. This is the highest-value data in the system and almost nobody collects it.

Re-fit on a schedule, quarterly for most businesses, and compare against the previous version on held-out recent deals before switching.

Without those four, you have a model rather than a system, and a model decays without anybody being able to prove it.

Measure the thing that actually moves revenue

Scoring accuracy is a proxy. The number that matters is whether high-scored leads convert better than the ones below them, and by enough to change behaviour.

The honest test: take the leads the model ranked in its top decile last quarter and compare their conversion rate against the rest. If the lift is under about 1.5x, the score is not earning the attention it asks for and the reps are right to ignore it.

And before spending anything on the model, check response time. Across a lot of inbound sales work the gap between a five-minute first response and a five-hour one dwarfs the difference between a good model and a mediocre one. Routing and alerting are cheaper to fix, and they move the same number.

The score is not the intervention

Worth separating, because teams routinely buy the first and needed the second.

A score ranks a queue. It does nothing about what happens to a lead once ranked, and in most businesses the losses are downstream of the ranking: nobody followed up a second time, the nurture track has not been touched in a year, or the routing rule sends everything to whoever is least busy rather than whoever knows the industry.

Rank the queue by all means. Then check whether the top of that queue is actually being worked differently from the bottom, because if it is not, the model has changed nothing and the diagnosis was wrong.

The enrichment trap

Buying third-party firmographic data feels like the way to improve fit scoring, and it does help - with two caveats worth stating before the invoice.

Coverage is uneven and correlated with size. Enrichment is excellent on large, well-known companies and poor on small ones. If your best customers are small, enrichment adds least where you need it most, and the model will learn to prefer large companies because that is where the data is.

Missing data is a signal, and it is usually the wrong one. A model given blank fields will often treat absence as a negative. Encode “unknown” explicitly rather than letting a null read as a low value.

What we build and what we push back on

The first deliverable is the definition of a qualified lead, written down, agreed by whoever runs sales, before any scoring exists. It sounds like a workshop artefact and it is the thing that decides whether anything after it works.

The second is the feedback loop, and we build it before the model rather than after. A rule-based score with a working outcome loop beats a good model without one, because the first improves every quarter and the second decays.

What we push back on hardest: a request for a scoring model when the business closes fewer than a few hundred deals a year. There is not enough signal to fit anything, and the honest answer is a transparent rule that the sales team helped write. That is a smaller engagement than the one being asked for, which is an awkward conversation and the right one - see how to evaluate an AI agency proposal for what an agency that does not say this looks like on paper.

Where we get it wrong ourselves: we have under-weighted how much of the value is in routing rather than scoring. Twice now the measurable win came almost entirely from getting the lead to the right person in minutes rather than hours, and the score contributed far less than the model’s accuracy suggested it should.

When this is not worth it

Low volume. If the team can work every inbound lead properly, prioritisation buys nothing. Spend the effort on response time and follow-up instead.

Fewer than roughly 300 closed deals a year. You cannot fit or validate a model on that - the differences you would act on are inside the noise. Write a rule.

Where the sales process is the constraint. If deals stall for reasons unrelated to lead quality, better prioritisation moves nothing. Find the actual bottleneck first.

When nobody will own it. A score with no owner is not re-fitted, not evaluated, and ignored within two quarters.

Frequently asked questions

Should the score be visible to reps?

Yes, with its reasons. A hidden score cannot be trusted or corrected, and the override data you lose by hiding it is the most useful data the system generates.

How often should we re-fit?

Quarterly for most businesses, and always after a material change to the product, pricing or target segment. Compare against the previous version on recent held-out deals before switching.

Do we need machine learning for this?

Frequently not. If the rule “target industry, right size, visited pricing” captures most of the signal, that is the system - and it has the advantage of being explainable by construction.

What about intent data from third parties?

Useful where it is accurate and expensive where it is not. Treat it as one input to test rather than a reason to rebuild, and measure the lift it produces before renewing.

How do we stop reps ignoring the score?

Involve them in defining a qualified lead, show the reasons alongside the number, and report the top-decile lift openly. A score that has been shown to work gets used; one that arrives by assertion does not.

Next step

If you have inbound volume and no agreed definition of a qualified lead, that definition is the first piece of work and it is not a technical one. The lead generation and management engagement starts there, then builds the outcome loop before the model.

Related: Better decisions, not just faster ones · How to measure whether an AI system works · How to evaluate an AI agency proposal · Lead generation and management

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.