Greek Supermarket Chains in ChatGPT Answers

Two identical runs, two clean sessions, same day. Supermarket chains were cited as a source in 2 of 15 answers the first time and 8 of 15 the second. Nothing about the brands changed in between.


We ran a frozen prompt set across Greece’s supermarket sector — fifteen questions a real shopper would ask, from “which supermarket has the best prices” to “how much is Zewa toilet paper” to “what is Lidl’s returns policy.” Every question run twice, on ChatGPT, logged out, fresh session each time, same day. Identical wording. No variation.

We expected the two runs to broadly agree. They did not.

In the first run, a supermarket chain appeared as a cited source in 2 of the 15 answers. In the second, 8. Same questions, same engine, same day, no change to any website in between.

That gap is the finding, and it matters more than any single result inside it.


What changed and what didn’t

The instability was not evenly distributed. It concentrated almost entirely in one type of question.

Buying-decision questions were volatile. “How much does a litre of Delta milk cost” produced no usable answer at all in the first run — the engine said it could not find a reliable current price. In the second, it returned four chain prices with four citations, two of them to chain domains. “Which chain has the best quality-to-price ratio” cited nothing in the first run and cited a chain in the second. Across the product and comparison categories, the answer source flipped in most of the set.

Operational questions were stable. Delivery charges, loyalty mechanics, return policies, delivery thresholds — eight runs across four questions, and in all eight the chain’s own domain was the source, with accurate figures and a direct link to the official page. Not one deviation.

The pattern is not “chains are invisible.” It is that chains are reliably visible on questions where no third party competes, and unpredictably visible everywhere else.

For a brand, the second half of that sentence is the harder problem. Consistent absence is at least a stable target. A source that appears in one session and vanishes in the next cannot be managed by looking at it once.


The chains that weren’t in the room

Citation was not the only thing that moved. So did the list of chains under consideration.

Counting how many of the fifteen answers named each chain:

ChainRun 1Run 2
Sklavenitis109
Lidl1010
My Market1011
AB Vassilopoulos98
Masoutis98
Galaxias15
Kritikos11
Bazaar02
Market In01
Thanopoulos01

The five largest chains are stable. A one-point swing across fifteen questions is noise; they are named whether or not they are cited, and they are named consistently.

Everything below them is not. Galaxias went from a single appearance to five. Three chains appeared only in the second run. The consideration set the engine drew from expanded from six chains to nine, between two identical runs on the same day.

This is a different problem from citation, and for a mid-sized chain it is the more urgent one. A brand that is never cited but always named is at least in the comparison. A brand whose presence in the answer depends on the session is not competing on merit — it is not reliably competing at all.

It also means market leadership functions as a floor. The top five are named because they are the obvious answer to “Greek supermarkets,” regardless of anything on their websites. For everyone else, inclusion appears to be contested each time.

Why this breaks the screenshot

Almost everything published about AI visibility rests on screenshots. Someone runs a query, sees their brand, posts the result. Or doesn’t see it, and posts that instead.

Our own two runs demonstrate why that evidence is close to worthless. Had we stopped after the first, we would have written that Greek supermarkets are systematically excluded from AI answers about their own products — a clean, publishable, wrong conclusion. Had we run only the second, we would have written that they are doing reasonably well. Both articles would have been built on real screenshots of real answers.

The measurement unit for AI visibility is not the answer. It is the distribution of answers across repeated runs.

This is also why the tooling in this space needs reading carefully. A visibility score derived from a single pass per prompt is reporting one sample from a distribution nobody has characterised, which is the distinction between a real AI visibility audit and a screenshot..


What we can say with two runs

Not much, and we would rather say that than overstate it.

Two runs per question establish that variance exists and is large. They do not establish its shape, its frequency, or what drives it. We do not know whether the second run’s higher citation rate reflects a change in the engine, a retrieval difference, or ordinary randomness. We can say that identical inputs produced materially different sourcing within a single day.

What survives with reasonable confidence:

  • Operational information is stable and chain-owned. Eight of eight, both runs. This is the one place a chain can count on being the answer.
  • Third-party platforms are always present. Price-comparison sites, leaflet aggregators, marketplaces and, newly, a state-run price platform appeared across both runs, in every category except operational. Even where a chain was cited, it was rarely cited alone. We measured the same dynamic in Greek banking, where comparison sites dominated across every question type.
  • Self-published comparison claims are discounted. In one run the engine cited a chain’s own basket-comparison claim and then explicitly noted that the chain had produced the comparison itself, declining to treat it as independent evidence. That single observation says more about what earns a citation than any amount of on-page optimisation advice.

That last point deserves emphasis, because it inverts a common recommendation. Publishing your own comparison content does not automatically make you the source on comparison questions. The engine can recognise self-interest, and appears to weight for it.


What a brand should actually do with this

Stop treating a screenshot as a status. Whether you appeared in an AI answer last Tuesday tells you almost nothing. Whether you appeared in 6 of 10 runs of a fixed question set, this month versus last, tells you something.

Separate the questions you own from the questions you contest. Operational information is defensible and currently well-held. Category and comparison questions are contested, unstable, and mostly held by third parties. These require different work and should not sit in the same report.

Look at who else is in the answer, every time. In this sector the intermediary layer is dense — comparison platforms, leaflet aggregators, marketplaces, media, and now a government platform. A chain that is cited alongside four aggregators has a different problem from one cited alone.

Assume the ground moves. A measurement is a timestamp, not a state. The only useful version of this work is repeated.


Methodology

Fifteen questions across six categories of buyer intent: chain selection, promotions, specific product, private label, and operational information. Category-level comparison questions (“which supermarket has the cheapest dairy”) were designed but not yet run, and are excluded from all figures here.

Each question was run twice on ChatGPT, logged out, in a fresh session, with no preceding turn in the conversation. Geolocation remained active and is recorded as a condition rather than suppressed. Both runs took place on the same day.

Each answer was scored for mention (is the chain named), citation (is it used as a source) and self-citation (is the source its own domain). Citation figures and mention counts are reported separately above.

Mention counts are lower bounds. Several answers extended below the captured area, so a chain named only in the uncaptured portion would not be counted. Four of the fifteen questions name a specific chain in the question itself, which inflates that chain’s mention count; those four are the operational questions.

An earlier attempt at a second run was discarded entirely: the questions had been asked sequentially within a single conversation, and the engine referred back to prior turns in its answers. Those runs are excluded and are not reflected in any figure here. Conversation context appeared to increase chain citation substantially, which is itself worth measuring properly — as a designed variable rather than as contamination.

Two runs per question is a small sample. It is enough to establish that variance is large and to disqualify single-run evidence. It is not enough to quantify anything, and we have not tried to.

The next iteration extends the same frozen set across Perplexity and Google AI Overviews, adds the category-comparison questions, and increases repetitions.

If you want to know where you stand

Every figure in this article came from a fixed question set, run twice, logged. Any brand can do this. Most don’t, and the ones that do usually run it once — which, as the numbers above show, tells you very little.

We run this measurement as a service for brands that want the distribution rather than the screenshot, across sectors and engines. If you want to know what your consideration-set position looks like across repeated runs, that is what an AEO / GEO audit measures.

Book an AEO / GEO Audit