There is no "your ChatGPT ranking": 13,500 searches prove it

If someone offers you "your ChatGPT visibility score" as a single number, be suspicious: that number cannot exist. The most ambitious study published so far on AI hotel recommendations proves it, and its conclusions should change both what you measure and who you buy tools from.
The study
Kollective's "Boutique Hotel AI Visibility Index" (interim phase published in July 2026; the full methodology was announced for September) analysed boutique hotels across 100 destinations: 13,500 AI searches, 93,000 hotel mentions and 9,600 distinct properties, spread across ChatGPT, Copilot, Gemini, Google's AI Mode and its AI Overviews. Mandatory disclaimer: Kollective sells AI visibility measurement, so it has skin in the game — yet its findings run, curiously, against the easy sales pitch of its own market.
The four data points that dismantle the "ranking"
The number one changes 45% of the time — on the same day. Repeating the same query on the same platform, the top recommendation changed in 45% of repetitions. We are not talking about weeks apart: we are talking about the same question, hours later.
Platforms agree only 4% of the time. ChatGPT, Copilot, Gemini, AI Mode and AI Overviews name the same hotel as first choice in barely 4% of queries. There is no "AI ranking": there are five different systems with their own sources and criteria.
70% of hotels appear on only one platform. Being well placed on Gemini says nothing about your presence on ChatGPT. Each platform is a separate channel with its own logic.
Even consistency differs between platforms. Gemini repeated its recommendations around 59% of the time; Google's AI Overviews, only 41% — and in one out of four "boutique" or "romantic" queries, Google simply didn't generate an AI summary at all.
The good news for the small players
Hidden in the study is a figure that should encourage any independent hotel: descriptors completely change the playing field. Big chains get only 3% of mentions in "boutique" hotel searches and 12% in "romantic" ones, against 25% in "luxury" searches. Translation: in queries with personality — the ones describing an experience rather than a standard category — independents dominate. A well-defined identity, which in the classic Google ranking was a disadvantage against brand budgets, is an asset in AI answers.
What to measure instead (and how to do it for free)
If point-in-time position is noise, the signal is consistency of appearance: how often you show up in each platform's answers, for your relevant queries, over time. That is exactly the metric the study proposes, and you can approximate it without buying anything:
- Define 5-10 real customer queries ("boutique hotel in [destination]", "romantic hotel near [place]").
- Run them on ChatGPT, Gemini, Copilot and Google every week or fortnight, 2-3 times per platform.
- Record in a spreadsheet whether you appear, in roughly what position, and which sources the answer cites.
- After eight weeks you will have your consistency baseline — and you will know which platform deserves your effort.
It is manual, modest work, but it produces the data point that matters. And when a vendor offers you an automated dashboard, you will know what to ask: not "what's my position?" but "how do you measure consistency across repetitions and platforms?". If the answer is a single number, the 13,500-search study has already told you what that number is worth.

