Sigma Logic AI Lead with AI. Thrive with Innovation.
AI visibility

Writing content that assistants actually cite

The page structures that survive being retrieved and extracted: answer-first openings, self-contained sections, real numbers, and claims a machine can attribute.

On this page 14 sections
  1. Key takeaways
  2. Who this applies to
  3. Write for the chunk, not the page
  4. The answer-first opening
  5. Specific claims get cited
  6. Headings as questions
  7. Tables are strong, and images are invisible
  8. What does not work
  9. Dates, and the claim that goes stale
  10. The uncomfortable dependency
  11. What we do, and where it costs us
  12. When not to bother
  13. Frequently asked questions
  14. Next step

A page gets cited when a retrieved chunk of it answers the question on its own. That favours pages that state the answer in the first paragraph, keep sections self-contained, and carry specific attributable claims - a number, a threshold, a named trade-off - rather than general advice. Most of this is good writing. The part that is genuinely different is designing for extraction rather than for reading order.

The useful mental model: your page will not be read. A few hundred words of it will be pulled out of context and asked to stand alone.

Key takeaways

  • The retrieved unit is a chunk, not a page. Write sections that survive being lifted out.
  • Answer first, then support. A conclusion in paragraph nine is invisible to extraction.
  • Specific claims get cited; general advice gets paraphrased without attribution.
  • A heading that matches a real question beats a clever one.
  • Nothing here works if the page is not in the retrieved set at all - that is a sourcing problem.

Who this applies to

You are writing or commissioning content and want it to be usable by assistants as well as readers. Assumes you have read why ChatGPT recommends your competitor and understand that being retrieved at all is upstream of this.

Write for the chunk, not the page

Retrieval systems split documents into passages and retrieve the passages that match. Your article is not evaluated as a whole; some portion of it competes on its own.

Three consequences.

Sections must be self-contained. A section that opens “As we saw above, this means…” is unusable when extracted. Repeat the subject rather than referring back. It reads slightly redundantly to someone consuming top to bottom and it is the difference between citable and not.

Front-load each section. Not just the article - every section should state its point in the first sentence, then support it. A section that builds to a conclusion loses that conclusion when only the opening is retrieved.

Keep one idea per section. A section covering three loosely related things matches weakly for all three. Splitting improves both retrieval and readability.

The answer-first opening

The strongest structural change available, and the one most content resists because it feels like giving away the conclusion.

Put a direct, complete answer in the first paragraph - two or three sentences, specific enough to stand alone, including the caveat if there is one. Then spend the rest of the article earning it.

The objection is that readers will leave. In practice they do not, because anyone who needed only the summary was never going to convert, and anyone with the actual problem reads on to find out whether you understand it. What you gain is a paragraph that can be lifted and attributed.

Every article on this site opens that way, including this one. It is a format decision, not a writing style.

Specific claims get cited

The clearest pattern in what gets attributed versus absorbed.

General advice is paraphrased without a source, because it is common knowledge and nothing marks it as yours. Specific claims are cited, because the specificity requires attribution.

Compare:

  • “Evaluation is an important part of AI projects.” Absorbed, uncited, correctly - it is not yours.
  • “Evaluation typically runs 15-20% of an AI build, and a test set of 150-400 stratified cases is enough for most business systems.” Citable. It commits to numbers somebody may want to attribute.

So: commit to numbers, name thresholds, state ranges, take positions. The general version of a sentence is safer and invisible.

This has an obvious integrity condition. Do not invent numbers to be citable. Made-up specificity is worse than vagueness: it is more likely to be repeated and it is wrong. Where a figure is derived rather than measured, say which - this site marks derived pricing as derived and holds articles that would need first-hand data it does not have.

What you publishWhat an extractor gets from itCitable
“Costs vary depending on your requirements”Nothing. No claim to liftNo
“Most support agents land between $12k and $30k to build”A range, attributable to youYes
A figure with the number only in the imageNothing. Images are not readNo
The same figure with the number in the captionThe number and its contextYes
A 900-word section answering four questionsA blurred chunk matching none of them wellRarely
A 150-word section answering one questionA chunk that matches that question closelyYes
“As we discussed above, this is why it matters”A chunk that cannot stand aloneNo
A comparison table with labelled rowsRows lift cleanly as structured factsYes
Why the 150-word section gets cited and the 900-word one does not Left: one 900-word section that answers four questions in one chunk, and matches none of them closely enough to be lifted. Right: four 150-word sections, each answering one question, each marked as cited. ONE 900-WORD SECTION FOUR 150-WORD SECTIONS Answers four questions in one chunk Matches none of them closely enough to be lifted as the answer One question, one answer cited One question, one answer cited One question, one answer cited One question, one answer cited RETRIEVAL SCORES PASSAGES, NOT PAGES. A SECTION IS A PASSAGE.
The retriever scores passages. A section that answers one question is a passage that matches one question, and the four short sections between them cover everything the long one did.

Headings as questions

Assistants retrieve against queries. A heading that matches the shape of a real question is a stronger match than a clever one.

  • “The Evaluation Imperative” - matches nothing anyone types
  • “How to measure whether an AI system works” - matches the query directly

Use the words your reader would use, not internal vocabulary. Then add an FAQ section with genuine questions, because that structure maps almost exactly onto how questions are asked and is disproportionately likely to be the retrieved chunk.

Tables are strong, and images are invisible

Comparison tables extract cleanly and carry a lot of information per token. Where you are comparing options, a table is likely to be the highest-value chunk on the page.

The corollary matters more than people expect: anything you communicate only in an image is invisible to this pipeline. A diagram carrying an argument that appears nowhere in text is decorative as far as retrieval is concerned. Every figure on this site has a caption that states its claim in words, and an aria-label describing what it shows - which serves screen readers first and extraction second.

What does not work

Keyword stuffing. Retrieval is semantic. Repeating a phrase does not increase match strength and reads badly.

Content volume without depth. Forty thin pages are not better than eight good ones. Thin pages match weakly and dilute your site.

Writing for the machine at the reader’s expense. The formats above are largely just clear writing. When a technique makes the page worse for a person, it is not worth doing - assistants are increasingly assessing quality signals humans would recognise, and a page nobody wants to read is a poor bet in any case.

Claiming things you cannot support. An assistant that surfaces your claim exposes it. Unsupported claims that get repeated are a liability with reach.

Dates, and the claim that goes stale

A specific claim is citable, which means it can be repeated for as long as the page exists. That is the benefit and it is also the exposure.

A figure that was accurate when published gets quoted back at you two years later, by a system that has no view on whether it still holds. Nothing in the pipeline checks whether your 2026 pricing range is still your 2028 pricing range.

Three practical habits.

Date the claim, not just the page. “As of 2026” inside the sentence travels with the extracted passage. A dateModified in the page metadata does not, because the chunk is what gets retrieved and the chunk carries no metadata.

Say what the claim is derived from. “Derived from published market rates” and “measured across our own projects” are different claims with different lifespans, and a reader who knows which one they have can judge staleness themselves.

Keep a review cadence on anything with a number in it. Not a full rewrite - a check that the figure still holds, and an update or an explicit note where it does not. Articles with specific claims need this and general advice does not, which is a real ongoing cost of writing the citable version.

The uncomfortable dependency

None of this matters if your page is never retrieved.

Retrieval happens against a search index. If your page does not rank, or your domain has no authority, structural work makes an invisible page marginally better-structured. The order is: be findable, then be extractable, then be citable.

Which means for a new site, third-party presence and conventional search work come first, and this article’s advice is what you apply to the pages once they can actually be reached. Anyone selling extraction optimisation to a site with no visibility is selling a second-floor extension on a house without foundations.

What we do, and where it costs us

Every article on this site follows this structure, and the honest description is that it is 80% ordinary good writing and 20% genuine adaptation.

The 20% that is real: self-contained sections that repeat their subject rather than referring back, an answer-first opening as a hard rule, and stating the claim of every figure in its caption because images do not survive extraction.

Where it costs us: the answer-first rule means the most valuable thing in each article is free, in the first paragraph, before anyone has spoken to us. A more commercial structure would gate the answer behind the argument. We think that trade is correct for a services business - a buyer who reads the summary and leaves was not a buyer - but it is a real trade and worth naming rather than pretending it is costless.

The rule we hold hardest is the integrity one. The temptation this whole discipline creates is to manufacture specificity, because specific claims get cited. That is how bad numbers enter circulation. Where we do not have a figure, the article says so or does not publish - which is why four of the cost articles on this site are written and held back.

When not to bother

When the content is not for discovery. Documentation for existing customers, internal material - structure it for its actual readers.

When you have nothing specific to say. A well-structured page with no substantive claim is still a page with no substantive claim.

Before you are findable. Covered above, and the most common misallocation.

Frequently asked questions

Does answer-first hurt time on page?

Some readers leave sooner. In practice those readers wanted a fact, not a supplier. The people with the underlying problem read further, because the summary demonstrates you understand it.

How long should sections be?

Roughly 150-400 words. Short enough to be one idea, long enough to stand alone when extracted.

Do FAQ sections still help?

Yes, and for retrieval they may be the strongest single structure, because the question-and-answer shape matches how queries arrive. Write genuine questions, not keyword variants of the same one.

Should we add FAQPage schema?

It is cheap and low-risk. Whether it changes assistant behaviour specifically is not well established - see llms.txt and structured data. Do it for search results; treat any assistant benefit as unproven upside.

How do we know if it is working?

A sampled prompt panel run monthly, tracking citation rate. Not a single test - answers vary between runs for the same prompt.

Next step

The structural work is worth doing on the pages you already have before writing new ones. The AI search visibility engagement covers answer-first content engineering with a citation baseline measured first, so there is something to compare against.

Related: Why ChatGPT recommends your competitor · llms.txt and structured data · GEO or SEO: what actually changes · AI search visibility

Let's talk

Got a workflow this applies to?

Describe it in a couple of sentences. We will tell you whether it is worth automating, what we would build, and roughly what it takes.