On this page 13 sections
- Key takeaways
- Who this applies to
- Why rank tracking does not transfer
- The method
- What we ran, and what we could not
- Result one: the answer layer here is made of supplier content
- Result two: whether anyone gets named depends on how the question is phrased
- Result three: our own baseline is zero, and the engine is describing a site that no longer exists
- The brand-name problem we did not expect to measure
- What we changed because of this
- When this is not worth measuring
- Frequently asked questions
- Next step
There is no rank to track, so the unit of measurement is a prompt and the thing you record is who gets cited and who gets named. We ran nine buying questions in this niche on 9 September 2026. Of 69 cited sources, 86 percent were companies selling into the space, and 64 of the 69 were distinct domains. Four of the nine questions named any organisation at all. Our own site appeared in none of them, and the engine was still describing the version of it that we replaced.
This is the first article on this site built on our own measurement rather than on reasoning. The sample is small and the method is repeatable, and both of those statements matter more than the numbers.
Key takeaways
- The source layer in this niche is overwhelmingly supplier-written, which means there is no gatekeeper to satisfy. Publishing is the way in.
- It is also extremely fragmented. Sixty-four distinct domains across sixty-nine citation slots, with almost no repetition.
- Naming is a function of phrasing. “Best” questions named organisations; “how do I choose” and “how much does it cost” named nobody.
- Being indexed is a precondition nothing else substitutes for, and a relaunch resets it.
- One engine, one country, one run, nine questions. Treat the direction as informative and the precision as not.
Who this applies to
You want to know whether AI assistants mention you, and you have discovered that the reporting you use for search does not transfer. This is the measurement method, and a worked example with the numbers left in.
The concept, and how generative engine work differs from search, is GEO vs SEO. Why an assistant recommends someone else is why ChatGPT recommends your competitor. This article is about how to put a number on it.
Why rank tracking does not transfer
Three properties break the tools you already own.
There is no position. An answer either mentions you or it does not. There is no eleventh place, so the metric is a rate over a set of prompts rather than an average position.
The same question does not return the same answer. Responses vary between runs and between users. A single check proves nothing, which makes the cadence part of the method rather than an operational detail.
A prompt is not a keyword. People ask assistants longer, more situational questions than they type into a search box, and the same underlying need appears in a dozen phrasings. You are sampling a space, not tracking a list.
What follows from all three: define a fixed prompt set, run it on a schedule, and record what came back. The prompt set is the instrument, and changing it invalidates your history.
The method
Build a prompt set from buying questions, grouped by intent. We used five families: vendor selection, cost, build decision, tool choice, and risk. Twenty to fifty prompts is a workable set. Ours was nine, which is a pilot rather than a study, and it is described here so it can be repeated and enlarged.
Record four things per prompt. Every domain cited. Every organisation named in the answer text, which is not the same list. Whether you appear in either. And the date, engine and country, because all three change the result.
Separate cited from named. A domain in the source list has been read. An organisation in the answer text has been recommended. The second is what a buyer acts on, and the two overlap less than you would expect.
Fix the cadence and keep the history. Monthly is enough for most businesses. The number that matters is the trend across runs, and it only exists if the prompt set stays still.
State the limits in the report. Engine, country, date, number of prompts, and whether it was one run or several. A share-of-voice figure without those five is decoration.
What we ran, and what we could not
Nine buying questions plus one brand query, on 9 September 2026, against a retrieval-augmented answer engine with a US index, one run per question. We also ran three index-freshness tests, described below.
Two things we could not do, stated because they bound the result. ChatGPT and Perplexity both required an account to answer, and we do not create accounts to run a measurement, so neither is represented here. That is a real limitation: this is one engine’s view, and engines differ in what they retrieve and how readily they name suppliers.
Everything below is counted from the returned source lists and answer text, not estimated.
Result one: the answer layer here is made of supplier content
Sixty-nine citation slots across the nine questions, and 64 distinct domains.
| What the cited page was | Slots | Share |
|---|---|---|
| A company selling into the space, on its own site | 59 | 86% |
| A directory, marketplace or review site | 7 | 10% |
| Media | 2 | 3% |
| Academic | 1 | 1% |
The classification rule is deliberately blunt: anything published by an organisation that sells software or services in this market counts in the first row, whether it is a product vendor or a consultancy. Only five domains appeared more than once in the whole set.
Two conclusions follow, and they point in opposite directions.
The encouraging one. There is no gatekeeper. In a niche where 86 percent of what gets cited is written by suppliers, the entry requirement is publishing something worth retrieving, not persuading a journalist or paying a directory. Almost every domain we saw was somebody’s own blog.
The sobering one. The fragmentation cuts both ways. Sixty-four distinct domains across sixty-nine slots means the field is enormous and each source is doing very little work. Nobody has consolidated this space, so nobody is being consistently rewarded, and a single well-cited article is a smaller asset here than the same article would be in a niche with five recognised sources.
Result two: whether anyone gets named depends on how the question is phrased
Four of the nine questions named any organisation. Twenty-three distinct organisations were named across the whole set.
The split is not random, and it is the most immediately useful thing we found.
“Best” and “top” phrasing produced names. The two vendor-selection questions using that phrasing returned eight and seven named suppliers respectively. In both cases the names came from listicles and directories in the source list, not from the suppliers’ own pages.
Everything else named nobody. “How to choose an AI consulting firm”, “questions to ask an AI development agency”, “how much does an AI customer support agent cost”, “build or buy”, “why do AI pilots fail” all returned substantive answers built from supplier blogs, and not one of them recommended a supplier. The blogs were used as material and their authors went uncredited in the answer text.
So there are two different games. Being cited is won by publishing the best explanation of a problem, and it earns you influence over what the answer says without earning you a mention. Being named is won by appearing on other people’s lists, and it is a different activity with different work attached. A company doing only the first will see its arguments reflected back in answers that recommend somebody else.
Result three: our own baseline is zero, and the engine is describing a site that no longer exists
Our site appeared in none of the 69 slots. That is expected for a domain with no authority, and it is not the interesting part.
Three tests, run the same day and each verified against the live site:
A site-restricted search returned six URLs, all from the previous site. Every path returned was a pre-relaunch address: the old post URL, the old service pages, the old about page. None of the sixty-one articles published since the rebuild appeared. All four of those legacy paths return a working 301 redirect to their current equivalent, checked live, so the redirects are correct and the index simply predates them.
The engine returned the old page title. It described the homepage as “AI consulting services for businesses”. The live title is “AI systems that ship, integrate and keep working”. The description it gave of the company was the previous site’s marketing copy, in a register the rebuild deliberately removed.
An exact-phrase search for a sentence that appears verbatim on a live article returned nine results and none of them ours. That is the cleanest available proof of absence: the string exists on the page, the page is served, and the index does not have it.
The honest reading is not that the answer layer is slow. The domain has only been serving the new site properly since 6 September, three days before this measurement. So this is a day-zero baseline, not a verdict. What it establishes is the starting position, and the useful number will be how many weeks it takes for the new corpus to displace the old one. We will publish that when we have it.
There is a general lesson in it for anyone relaunching. On the day you replace a site, the answer layer keeps describing the old one, including copy you retired on purpose, and there is no mechanism to tell it otherwise. Correct redirects do not accelerate this. They only guarantee that the person who follows a stale link arrives somewhere sensible.
The brand-name problem we did not expect to measure
A quoted search for our own brand returned six other technology companies with Sigma in the name, across AI consulting, analytics, software services and IT support.
This is a measurable liability rather than a nuisance. When a buyer asks an assistant about you by name, the retrieval has to disambiguate you from six similarly named firms, several with far more authority. Everything that pins an entity down helps here: a consistent legal name, an address, a named author with a real profile, and structured data that says which organisation this is. We had already done most of that; we had not previously thought of it as an AI visibility control, and after this measurement we do.
What we changed because of this
We stopped treating publishing as the whole strategy. The corpus argues well and it is invisible in the two questions that name anybody. Getting onto credible third-party lists is separate work, and it is now on the plan rather than assumed to follow.
We put the vendor’s own partner directory first. Where a tool we work with runs a certified partner listing, that listing appeared in the source set. It is one of the few directory-shaped assets in this niche that is not purely commercial.
We date-stamped the baseline and fixed the prompt set. The nine questions are recorded verbatim so that the next run is comparable. Enlarging the set is allowed; changing it is not, unless we start again.
We are not spending on this until the index catches up. Nothing in an AI visibility programme functions before you are retrievable. That ordering is unglamorous and it is what the measurement says.
When this is not worth measuring
Before you are indexed. Everything above is downstream of retrieval. If an exact-phrase search for your own sentence returns nothing, that is the only number you need this month.
If your buyers do not ask assistants. Some categories are still bought through procurement lists, referrals and trade relationships. Measure the channel that exists rather than the one being written about.
At a cadence you cannot sustain. A monthly run of twenty prompts that actually happens beats a quarterly programme of two hundred that gets skipped, because only one of them produces a trend.
Frequently asked questions
How many prompts do I need for this to mean anything?
More than we used. Nine is a pilot, and we have labelled it as one. Twenty to fifty covers the intent families for most businesses, and the sample size should appear next to every number you report.
Should I run each prompt more than once?
Yes, once the set is established. Answers vary between runs, so a single response is one sample of a distribution. Three runs per prompt per month is a reasonable compromise between cost and stability.
Does being cited help if the answer does not name us?
Yes, and less than it feels. Your argument shapes what the answer says, which is real influence over how a buyer thinks about the problem. It does not send them to you. Both halves are worth having and they are not substitutes.
Can we make an assistant re-read our site?
Not directly, and be sceptical of anyone selling that. What you control is being retrievable, being unambiguous about who you are, and being mentioned in places that are read. The rest is waiting, and measuring while you wait.
Is this the same as an AI visibility tool subscription?
The tools do this at a scale and cadence a person cannot, across engines that require accounts, which is a real advantage. The method is the same, and running it by hand once is a good way to find out whether the subscription would tell you anything you would act on.
Next step
If you want the version of this for your own business, it is a prompt set, one run and a written baseline, and it usually costs less than a month of the tooling. AI search visibility starts there, and if the answer is that you are not indexed yet, we will tell you that and stop.
Related: GEO vs SEO: what actually changes · Why ChatGPT recommends your competitor · Writing content that assistants actually cite · llms.txt and structured data · AI search visibility