LUCENTCOMMERCEGET A FREE STORE AUDITFREE AUDIT

AEO · GEO · SEO · AI · 20 MAY 2025 · 8 MIN READ

What AI assistants say about your brand, and why

Assistants do not remember your brand, they retrieve it. Two of the three things they retrieve are not on your site — and Google says there is no markup that changes this.

A search result and the page behind it

An assistant answering “is this brand any good” is doing one of three things: retrieving pages from your site that its crawler was allowed to fetch, retrieving what third parties have written about you, or recalling whatever was in its training data. Only the first is yours to control, and the other two usually carry more weight for a retail brand. There is no markup that shortcuts this. Google’s own documentation on AI features states that “there are no additional technical requirements” and that there is “no special schema.org structured data that you need to add.”

IN SHORT

  • AI answers about a brand are assembled from retrieved sources at answer time, not from a stored profile of your company.
  • Google states that appearing in its AI features requires only that a page be indexed and eligible to show with a snippet, with no additional technical requirements.
  • Google also states you do not need new machine-readable files, AI text files, or special schema.org markup to appear in these features.
  • Crawler access is a real decision: OAI-SearchBot governs indexing for ChatGPT search while GPTBot governs model training, and blocking one is not the same as blocking the other.
  • Shopify themes control crawler rules through robots.txt.liquid, whose default rule groups Shopify updates regularly.
  • For most retail brands the highest-leverage work is off-site — reviews, comparisons and coverage — because that is what there is to retrieve.

There is no profile of your brand

The intuition to discard first is that an assistant holds an opinion about your company that you might edit. It does not. When someone asks whether your products are worth buying, the system assembles an answer from sources it can reach at that moment, plus a general sense of the category absorbed during training. Ask the same question twice and you can get two differently sourced answers, because the retrieval changed.

That makes the practical question much narrower than “how do we optimise for AI”. It is: when something goes looking for a factual statement about our brand, what does it find, and is that statement correct? Three places supply it.

  • Your own pages, as fetched by whichever crawler the assistant uses. Everything on-site you can control lives here.
  • Third-party sources — retailer listings, review sites, comparison articles, forum threads, press. Not yours, and usually more of it than you have.
  • Training data, which is a snapshot: old prices, a discontinued range, a founder who left. You cannot edit it, and it fades as retrieval-grounded answers become the norm.

What Google actually says, and what it kills

There is a whole industry selling special files and bespoke markup for AI visibility. Google’s documentation on AI features in Search is unusually blunt about it: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.”

The eligibility bar it does state is the ordinary one — “a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements” — followed by “there are no additional technical requirements.” And it declines to promise anything: “Just because a page meets all requirements, best practices, and complies with the policies, doesn’t mean that Google will crawl, index, or serve its content. Indexing and serving isn’t guaranteed.”

Read together, that is a specific and useful message. The work is being crawlable, being indexed, and having content clear enough to quote. The work is not a new file in your root directory. If a proposal in front of you is mostly a file and a score, it is selling you the easy half of a problem that does not have an easy half.

None of which makes structured data pointless — we put Product, FAQPage and BreadcrumbList markup on everything, because it earns rich results in classic search and it forces a discipline of stating facts explicitly. Just do not buy it as an AI cheat code, because the people who run the largest AI feature in the world have said in writing that it is not one.

You have already made a crawler decision, probably by accident

Assistants use named crawlers, and each one does a different job. OpenAI documents GPTBot as the crawler for training generative models and OAI-SearchBot as the one that indexes for search results inside ChatGPT. It also documents ChatGPT-User as user-triggered fetching — a page a person asked the assistant to look at — and notes it is “not used for crawling the web in an automatic fashion,” so robots.txt rules may not govern it. Google, separately, offers Google-Extended as the control for limiting AI training and grounding in some of its other systems, distinct from Googlebot.

The distinction matters more than it sounds. Disallowing GPTBot keeps your content out of training. Disallowing OAI-SearchBot keeps you out of the search index the assistant reaches for when a customer asks for a recommendation. Plenty of brands have blanket-blocked both under one instruction to “stop AI scraping us” and taken themselves out of the shop window to protect a copyright position they had not articulated.

Have the conversation properly, with whoever owns brand and whoever owns legal in the room, and decide the two separately. We would generally allow search crawling and treat training as a policy call the business makes on its own terms. Either way, decide it — because the default is a decision too.

How to change it on Shopify

Shopify serves a default robots.txt and it is fine for most stores; Shopify notes that “the default rules are updated regularly to ensure that SEO best practices are always applied.” When you do need your own rules, the mechanism is a robots.txt.liquid template in the theme, which outputs groups of rules using the robots Liquid object. Shopify is specific that it “can’t be a JSON template” and that while you can replace the whole thing with plain text rules, it is “strongly recommended to use the provided Liquid objects whenever possible.”

Follow that recommendation. Iterating over robots.default_groups and appending your own rules means you inherit Shopify’s ongoing updates; pasting a static file means you inherit a snapshot from the day someone pasted it, and nobody revisits a robots.txt. We have seen a store lose a section of its catalogue from search for months because a hand-written file outlived the reason it was written.

For a retail brand, most of the answer is off-site

This is the part clients least want to hear. If an assistant is asked to compare three brands in your category, the sentences it has to work with mostly come from people who are not you — reviews with substance, comparison articles, a forum thread where somebody described the fit, a retailer page with full specifications. A brand with a beautiful site and no third-party presence gives a retrieval system nothing to say beyond its own marketing, and marketing copy is exactly the kind of source these systems weight down.

The uncomfortable conclusion is that the highest-leverage AEO budget for many brands is not a website project at all. It is getting reviewed by people who write in detail, getting your specifications onto the places that list your category properly, and making sure your own factual claims are consistent everywhere they appear — because a contradiction between your site and a retailer’s listing is resolved by a machine picking one, and it may not pick yours.

What this does mean for the site: your facts need to be stated somewhere quotable, in text, in one obvious place. Materials, dimensions, care, compatibility, delivery times, returns terms, sizing. Every one of those we routinely find either absent, trapped in an image, or loaded by an app after the page renders — which for retrieval purposes is the same as absent.

What we would actually change on the site

Ordinary, unglamorous, and it works for human readers too, which is the test a tactic should pass before you adopt it.

  • Answer the question in the first paragraph, self-contained, so it can be lifted without the paragraph before it. Anything that opens with context and reaches the answer in paragraph four is unquotable.
  • Put the facts in text. No specification tables rendered as images, no delivery windows living only in a shipping widget, no key detail injected by an app script.
  • Use tables and numbered steps. Both survive extraction far better than prose describing the same thing.
  • One page per real question. A page that genuinely answers “how do I choose a size” is retrievable; the same information scattered across a product description, an FAQ accordion and a blog post is not.
  • Name things consistently. One product name, one spelling, one model number, on your site and everywhere you supply data. Inconsistency is how a machine ends up describing two products as one.
  • Keep it current, and say when. A page contradicted by a newer source will lose to it, so dated facts and a visible last-updated date are worth more than they look.

The publishing constraint nobody costs in

All of the above is a content programme, and content programmes die on the same thing: publishing a page takes a developer two weeks. If answering a new customer question requires a ticket, the questions go unanswered and the category gets described by whoever did answer them.

This is why a composable section library is the piece of infrastructure that actually moves AEO for most stores. Merchandisers assembling a proper answer page — hero, comparison table, numbered steps, FAQ — on a Tuesday afternoon is worth more over a year than any markup you could add. The tactic list is cheap; the ability to execute it repeatedly is not.

How to tell whether any of it worked

Be honest that this is poorly measurable and refuse to pretend otherwise. There is no rank tracker for an assistant’s answer, and the visibility scores being sold are models of a black box rather than observations of one.

What you can do: sample. Write down the fifteen questions a real buyer in your category asks, run them across the assistants your customers use, and record what comes back and which sources it cited — the same fifteen, monthly, by the same method. That is a crude instrument, but it is the only one with actual signal in it, and it catches the thing that matters most: a factual error about your brand repeating across systems.

Then watch the things you can measure and accept that attribution is partial. Referral traffic from assistants where it identifies itself, branded search volume, and direct traffic that arrives already knowing what it wants. If someone offers you a single AI visibility number, ask how it is derived; if the answer is a proprietary model, you are buying a reassuring chart.

Questions this raises

Do I need llms.txt or special markup for AI search?

Google’s documentation on AI features says no: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.” Other systems publish no requirement for such a file either. Treat it as optional at best.

Is structured data still worth adding?

Yes, but for the right reason. Product, FAQ and breadcrumb markup earns rich results in classic search and makes your facts explicit and machine-readable. It is good practice with a direct payoff — it just is not a documented requirement for appearing in AI answers.

Should I block GPTBot?

That is a business decision about training, not an SEO decision, and it should be made separately from search crawling. Blocking GPTBot limits use of your content for model training; blocking OAI-SearchBot removes you from the index ChatGPT search draws on. Confusing the two is the expensive mistake.

How do I change robots.txt on Shopify?

Add a `robots.txt.liquid` template to the theme. It cannot be a JSON template, and Shopify strongly recommends using the provided Liquid objects rather than pasting plain text, so that you keep inheriting the default rule groups Shopify updates regularly.

Why does an assistant describe my products incorrectly?

Usually because the correct version is not retrievable. The facts may be in an image, injected by an app after render, contradicted by a retailer listing, or only present in training data from two years ago. Put the current facts in plain text on a page that is indexed, and make sure your data is consistent wherever else it appears.

Can I track AI visibility like keyword rankings?

Not reliably. Answers vary by phrasing, user and retrieval, so there is no stable position to track. Sample a fixed set of buyer questions on a schedule, record the answers and the cited sources, and watch branded and direct traffic alongside. Be sceptical of any single proprietary visibility score.

NEXT STEP

Free store audit

A senior Shopify engineer reviews your storefront, theme performance and checkout, then sends a prioritised list of fixes.