Sigma Logic AI Lead with AI. Thrive with Innovation.
AI visibility

Measuring AI visibility: the method, and our own baseline

Nine buying questions, 69 cited sources, 64 distinct domains. What the answer layer in this niche is made of, and what it says about a site three days old.

On this page 13 sections
  1. Key takeaways
  2. Who this applies to
  3. Why rank tracking does not transfer
  4. The method
  5. What we ran, and what we could not
  6. Result one: the answer layer here is made of supplier content
  7. Result two: whether anyone gets named depends on how the question is phrased
  8. Result three: our own baseline is zero, and the engine is describing a site that no longer exists
  9. The brand-name problem we did not expect to measure
  10. What we changed because of this
  11. When this is not worth measuring
  12. Frequently asked questions
  13. Next step

There is no rank to track, so the unit of measurement is a prompt and the thing you record is who gets cited and who gets named. We ran nine buying questions in this niche on 9 September 2026. Of 69 cited sources, 86 percent were companies selling into the space, and 64 of the 69 were distinct domains. Four of the nine questions named any organisation at all. Our own site appeared in none of them, and the engine was still describing the version of it that we replaced.

This is the first article on this site built on our own measurement rather than on reasoning. The sample is small and the method is repeatable, and both of those statements matter more than the numbers.

Key takeaways

  • The source layer in this niche is overwhelmingly supplier-written, which means there is no gatekeeper to satisfy. Publishing is the way in.
  • It is also extremely fragmented. Sixty-four distinct domains across sixty-nine citation slots, with almost no repetition.
  • Naming is a function of phrasing. “Best” questions named organisations; “how do I choose” and “how much does it cost” named nobody.
  • Being indexed is a precondition nothing else substitutes for, and a relaunch resets it.
  • One engine, one country, one run, nine questions. Treat the direction as informative and the precision as not.

Who this applies to

You want to know whether AI assistants mention you, and you have discovered that the reporting you use for search does not transfer. This is the measurement method, and a worked example with the numbers left in.

The concept, and how generative engine work differs from search, is GEO vs SEO. Why an assistant recommends someone else is why ChatGPT recommends your competitor. This article is about how to put a number on it.

Why rank tracking does not transfer

Three properties break the tools you already own.

There is no position. An answer either mentions you or it does not. There is no eleventh place, so the metric is a rate over a set of prompts rather than an average position.

The same question does not return the same answer. Responses vary between runs and between users. A single check proves nothing, which makes the cadence part of the method rather than an operational detail.

A prompt is not a keyword. People ask assistants longer, more situational questions than they type into a search box, and the same underlying need appears in a dozen phrasings. You are sampling a space, not tracking a list.

What follows from all three: define a fixed prompt set, run it on a schedule, and record what came back. The prompt set is the instrument, and changing it invalidates your history.

The method

Build a prompt set from buying questions, grouped by intent. We used five families: vendor selection, cost, build decision, tool choice, and risk. Twenty to fifty prompts is a workable set. Ours was nine, which is a pilot rather than a study, and it is described here so it can be repeated and enlarged.

Record four things per prompt. Every domain cited. Every organisation named in the answer text, which is not the same list. Whether you appear in either. And the date, engine and country, because all three change the result.

Separate cited from named. A domain in the source list has been read. An organisation in the answer text has been recommended. The second is what a buyer acts on, and the two overlap less than you would expect.

Fix the cadence and keep the history. Monthly is enough for most businesses. The number that matters is the trend across runs, and it only exists if the prompt set stays still.

State the limits in the report. Engine, country, date, number of prompts, and whether it was one run or several. A share-of-voice figure without those five is decoration.

What we ran, and what we could not

Nine buying questions plus one brand query, on 9 September 2026, against a retrieval-augmented answer engine with a US index, one run per question. We also ran three index-freshness tests, described below.

Two things we could not do, stated because they bound the result. ChatGPT and Perplexity both required an account to answer, and we do not create accounts to run a measurement, so neither is represented here. That is a real limitation: this is one engine’s view, and engines differ in what they retrieve and how readily they name suppliers.

Everything below is counted from the returned source lists and answer text, not estimated.

Result one: the answer layer here is made of supplier content

Sixty-nine citation slots across the nine questions, and 64 distinct domains.

What the cited page wasSlotsShare
A company selling into the space, on its own site5986%
A directory, marketplace or review site710%
Media23%
Academic11%

The classification rule is deliberately blunt: anything published by an organisation that sells software or services in this market counts in the first row, whether it is a product vendor or a consultancy. Only five domains appeared more than once in the whole set.

Two conclusions follow, and they point in opposite directions.

The encouraging one. There is no gatekeeper. In a niche where 86 percent of what gets cited is written by suppliers, the entry requirement is publishing something worth retrieving, not persuading a journalist or paying a directory. Almost every domain we saw was somebody’s own blog.

The sobering one. The fragmentation cuts both ways. Sixty-four distinct domains across sixty-nine slots means the field is enormous and each source is doing very little work. Nobody has consolidated this space, so nobody is being consistently rewarded, and a single well-cited article is a smaller asset here than the same article would be in a niche with five recognised sources.

What the cited sources were, across nine buying questions Of 69 cited sources across nine buying questions, 59 or 86 percent were companies selling into the space publishing on their own sites, 7 or 10 percent were directories or marketplaces, 2 or 3 percent were media, and 1 was academic. Sixty-four of the sixty-nine were distinct domains and only five appeared twice. That means there is no gatekeeper to satisfy, and also no consolidated source that anyone is being rewarded for. Measured on 9 September 2026 against one engine with a US index, one run per question. 69 cited sources across 9 buying questions A company selling in the space 59 (86%) A directory or marketplace 7 (10%) Media 2 (3%) Academic 1 (1%) 64 of the 69 were distinct domains. Only five appeared twice. No gatekeeper to satisfy, and no consolidated source being rewarded. Measured 9 September 2026, one engine, US index, one run per question.
The classification is deliberately blunt: any organisation selling software or services here counts in the top row. Almost every cited page was somebody’s own blog.

Result two: whether anyone gets named depends on how the question is phrased

Four of the nine questions named any organisation. Twenty-three distinct organisations were named across the whole set.

The split is not random, and it is the most immediately useful thing we found.

“Best” and “top” phrasing produced names. The two vendor-selection questions using that phrasing returned eight and seven named suppliers respectively. In both cases the names came from listicles and directories in the source list, not from the suppliers’ own pages.

Everything else named nobody. “How to choose an AI consulting firm”, “questions to ask an AI development agency”, “how much does an AI customer support agent cost”, “build or buy”, “why do AI pilots fail” all returned substantive answers built from supplier blogs, and not one of them recommended a supplier. The blogs were used as material and their authors went uncredited in the answer text.

So there are two different games. Being cited is won by publishing the best explanation of a problem, and it earns you influence over what the answer says without earning you a mention. Being named is won by appearing on other people’s lists, and it is a different activity with different work attached. A company doing only the first will see its arguments reflected back in answers that recommend somebody else.

Which questions named an organisation and which named nobody Four of the nine buying questions named any organisation. The two using best or top phrasing named eight and seven suppliers; a tool comparison named three products; a risk question named five organisations. The other five, covering how to choose a consulting firm, what to ask an agency, how much an agent costs, build or buy, and why pilots fail, named nobody at all. The names came from listicles rather than from the suppliers own pages, and five substantive answers were built from supplier blogs whose authors went uncredited. Being cited shapes what the answer says; being named sends the buyer somewhere. Four of nine questions named any organisation "Best AI automation agency" 8 named "Best n8n implementation partner" 7 named "n8n vs Make vs Zapier" 3 named "How accurate are AI agents" 5 named "How to choose an AI consulting firm" nobody "Questions to ask an AI agency" nobody "How much does an AI agent cost" nobody "Build or buy AI support" nobody "Why do AI pilots fail" nobody The names came from listicles, not from the suppliers’ own pages. Five answers were built from supplier blogs that went uncredited. Being cited shapes the answer. Being named sends the buyer somewhere.
Two different games with different work attached. A company doing only the publishing half sees its own arguments in an answer that recommends somebody else.

Result three: our own baseline is zero, and the engine is describing a site that no longer exists

Our site appeared in none of the 69 slots. That is expected for a domain with no authority, and it is not the interesting part.

Three tests, run the same day and each verified against the live site:

A site-restricted search returned six URLs, all from the previous site. Every path returned was a pre-relaunch address: the old post URL, the old service pages, the old about page. None of the sixty-one articles published since the rebuild appeared. All four of those legacy paths return a working 301 redirect to their current equivalent, checked live, so the redirects are correct and the index simply predates them.

The engine returned the old page title. It described the homepage as “AI consulting services for businesses”. The live title is “AI systems that ship, integrate and keep working”. The description it gave of the company was the previous site’s marketing copy, in a register the rebuild deliberately removed.

An exact-phrase search for a sentence that appears verbatim on a live article returned nine results and none of them ours. That is the cleanest available proof of absence: the string exists on the page, the page is served, and the index does not have it.

The honest reading is not that the answer layer is slow. The domain has only been serving the new site properly since 6 September, three days before this measurement. So this is a day-zero baseline, not a verdict. What it establishes is the starting position, and the useful number will be how many weeks it takes for the new corpus to displace the old one. We will publish that when we have it.

There is a general lesson in it for anyone relaunching. On the day you replace a site, the answer layer keeps describing the old one, including copy you retired on purpose, and there is no mechanism to tell it otherwise. Correct redirects do not accelerate this. They only guarantee that the person who follows a stale link arrives somewhere sensible.

Three tests of whether the new site is in the index Three tests run and verified on the same day. A site-restricted search returned six URLs, all pre-relaunch paths, and none of the sixty-one published articles. The homepage title returned was the old one, retired on 6 September. An exact-phrase search for a sentence that is served verbatim on a live page returned nine results and none of them ours. The redirects are correct and the index simply predates them: the domain had been serving the new site for three days when this was run. A relaunch resets this, and nothing you control makes it happen faster. Three tests, run and verified on the same day Site-restricted search 6 URLs, all pre-relaunch paths 0 of 61 articles Homepage title returned "AI consulting services for businesses" retired on 6 Sept Exact phrase from a live page 9 results, none of them ours the string is served The redirects are correct. The index simply predates them. The domain had been serving the new site for three days. A relaunch resets this, and nothing you control makes it happen faster.
This is a day-zero baseline, not a verdict on the engine. The useful number is how many weeks the recovery takes, and that needs a second run.

The brand-name problem we did not expect to measure

A quoted search for our own brand returned six other technology companies with Sigma in the name, across AI consulting, analytics, software services and IT support.

This is a measurable liability rather than a nuisance. When a buyer asks an assistant about you by name, the retrieval has to disambiguate you from six similarly named firms, several with far more authority. Everything that pins an entity down helps here: a consistent legal name, an address, a named author with a real profile, and structured data that says which organisation this is. We had already done most of that; we had not previously thought of it as an AI visibility control, and after this measurement we do.

What we changed because of this

We stopped treating publishing as the whole strategy. The corpus argues well and it is invisible in the two questions that name anybody. Getting onto credible third-party lists is separate work, and it is now on the plan rather than assumed to follow.

We put the vendor’s own partner directory first. Where a tool we work with runs a certified partner listing, that listing appeared in the source set. It is one of the few directory-shaped assets in this niche that is not purely commercial.

We date-stamped the baseline and fixed the prompt set. The nine questions are recorded verbatim so that the next run is comparable. Enlarging the set is allowed; changing it is not, unless we start again.

We are not spending on this until the index catches up. Nothing in an AI visibility programme functions before you are retrievable. That ordering is unglamorous and it is what the measurement says.

When this is not worth measuring

Before you are indexed. Everything above is downstream of retrieval. If an exact-phrase search for your own sentence returns nothing, that is the only number you need this month.

If your buyers do not ask assistants. Some categories are still bought through procurement lists, referrals and trade relationships. Measure the channel that exists rather than the one being written about.

At a cadence you cannot sustain. A monthly run of twenty prompts that actually happens beats a quarterly programme of two hundred that gets skipped, because only one of them produces a trend.

Frequently asked questions

How many prompts do I need for this to mean anything?

More than we used. Nine is a pilot, and we have labelled it as one. Twenty to fifty covers the intent families for most businesses, and the sample size should appear next to every number you report.

Should I run each prompt more than once?

Yes, once the set is established. Answers vary between runs, so a single response is one sample of a distribution. Three runs per prompt per month is a reasonable compromise between cost and stability.

Does being cited help if the answer does not name us?

Yes, and less than it feels. Your argument shapes what the answer says, which is real influence over how a buyer thinks about the problem. It does not send them to you. Both halves are worth having and they are not substitutes.

Can we make an assistant re-read our site?

Not directly, and be sceptical of anyone selling that. What you control is being retrievable, being unambiguous about who you are, and being mentioned in places that are read. The rest is waiting, and measuring while you wait.

Is this the same as an AI visibility tool subscription?

The tools do this at a scale and cadence a person cannot, across engines that require accounts, which is a real advantage. The method is the same, and running it by hand once is a good way to find out whether the subscription would tell you anything you would act on.

Next step

If you want the version of this for your own business, it is a prompt set, one run and a written baseline, and it usually costs less than a month of the tooling. AI search visibility starts there, and if the answer is that you are not indexed yet, we will tell you that and stop.

Related: GEO vs SEO: what actually changes · Why ChatGPT recommends your competitor · Writing content that assistants actually cite · llms.txt and structured data · AI search visibility

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.