Research
Can You Trust AI Hotel Answers? We Checked 3,100 Facts
We fact-checked 3,100 facts in 414 AI answers about 46 real hotels. 97% of facts were right, yet 1 in 4 answers contained a wrong claim. The full study.

Everyone is asking whether AI should make the booking. We asked an earlier question: is AI even giving travellers accurate information about hotels? We fact-checked 414 AI answers about 46 real hotels. This is what travellers are being told.
The travel industry has decided that agentic booking is next. OpenAI ships an operator that fills in checkout forms. Booking platforms are racing to publish agent-ready APIs. Every conference panel this year has asked some version of "will AI make the booking?"
We think there is a question that comes before that one: before we let AI book the room, can we trust what it tells travellers about the room?
Millions of travellers already use AI as their front desk: they ask ChatGPT and Google whether a hotel has a pool, whether it takes pets, whether it is on the beach, whether it is even open. If the industry is about to hand this channel the payment step, it is worth knowing how it performs on the information step.
So we tested it.
What we did
We took 46 four- and five-star hotels across five Mediterranean markets: Limassol, Paphos, Athens, Rhodes and Chania. For each hotel we built a ground truth of ten checkable facts, star rating, pools, beachfront, distances, pet policy, adults-only status, spa, parking and opening status, triangulated across five sources: the hotel's own website, its Booking.com profile, Google's hotel data, its Google Business Profile and Tripadvisor. Where the sources disagreed, we resolved the conflict by hand against the hotel's own site.
Then we asked the AI the way a traveller would: one natural question per hotel covering all ten facts, put to ChatGPT (GPT-5) with web search on, ChatGPT with web search off, and Google AI Mode. Three times each, because AI does not always give the same answer twice. That is 414 graded answers, over 3,700 individual fact checks, all captured on the same day. Every raw response is archived.
We also set four traps: hotels that no longer exist as asked. One demolished. One closed for a multi-year rebuild. Two renamed.
The good news first
AI is more accurate than the discourse suggests. Across everything the three engines asserted, 96 to 97 percent of facts were correct. Star ratings: zero errors in 414 answers. Whether the hotel is open: zero errors. Adults-only status: zero errors. Both renamed hotels were correctly identified as their new brands by every engine in every run, and the answers were sometimes fresher than the listings themselves. One Cretan resort that went adults-only this April is still listed as family-friendly on two major platforms; the AI got it right anyway.
If your instinct is to dismiss AI answers as hallucination soup, the data does not support you.
The same data, the way a guest receives it
Now the catch in that 97 percent, and it is easiest to see in raw counts. Across all the answers, we checked just over 3,100 individual facts. The engines got about 100 of them wrong. A hundred mistakes in three thousand facts sounds like a rounding error.
Here is why it is not. Those hundred mistakes did not pile up inside a handful of terrible answers. They spread out, mostly one per answer, across nearly a hundred different answers. And there were only 414 answers in total.
So the same data reads two ways. Count the facts: 97 percent correct. Count the answers a traveller actually received: 96 of 414 contained at least one wrong claim, close to one in four. By engine: 16 percent for ChatGPT with web search, 21 percent for Google AI Mode, and one in three for ChatGPT answering from memory alone.
The guest lives on the second count. Nobody experiences your per-fact error rate. They experience the one sentence that said you take no dogs.
The errors are not random noise either. They run in consistent directions:
- Pets. The single worst fact in the study. Where a hotel's policy is published, AI told travellers "no pets" at hotels that allow them 24 times, and the reverse only twice. AI's default hotel does not take dogs. If yours does, that is a filtered-out booking you never see.
One hotel, one question, five answers. The Leonardo Plaza Cypria Maris in Paphos accepts dogs up to 10 kg; its own website says so, with the fee, and Booking.com agrees. Google's hotel data and Tripadvisor say nothing either way. Go one ring further out on the web and the signal flips: aggregators such as Trivago and Trip.com state flatly that pets are not allowed. We asked the AI nine times. Five answers told travellers pets are not allowed, and one of them cited the hotel's own website while contradicting it. ChatGPT with web search was the only condition that never asserted the wrong policy; it mostly declined to answer. A dog owner never argues with that answer. They just book somewhere else, and the hotel never knows.
- Spas and pools. AI invents them. Fourteen times it promised a spa at a hotel that has none; seven times it added an indoor pool to a hotel with only outdoor ones. These are the errors that become front-desk arguments, because the guest arrives expecting something the AI sold them.
- Beachfront. The industry's fuzziest word, and AI inherits the fuzz: hotels across a coastal road were called beachfront and genuinely beachfront hotels were talked down, in both directions, usually echoing whichever listing the answer leaned on.
Sixteen of the 46 hotels came through all nine of their answers without a single error. Others were wrong in nearly every answer. The difference was not luxury tier or brand. It was the state of their information layer, which brings us to the traps.
The ghost hotel that six platforms will still sell you
The Curium Palace in Limassol was demolished. The site is becoming an office development. We asked all three engines whether it was a good option for a trip.
Google AI Mode said, correctly, that it no longer operates. All three runs.
ChatGPT with web search told us, all three runs, that the hotel is a solid four-star choice and that "availability is being offered today on Google Hotels, Expedia and Hotels.com". It cited a hotel association member list from 2017, a university accommodation page and live OTA listing pages.
Here is the uncomfortable part: ChatGPT was not entirely making it up. We checked. On the day of the study, Expedia, Hotels.com, Kayak, Momondo and Thomas Cook all still carried the Curium Palace as a live four-star hotel: page up, photos up, 223 reviews, a 3.5 score. Run an actual date search, though, and the truth leaks out sideways: "Sold out! Our last room has already been booked." The hotel is not sold out. The hotel is not there. But no platform says so; the listing simply sits in a state that looks, to a machine, like a popular hotel on a busy week.
And beyond the OTAs sits something stranger still: a third-party booking site at curiumpalacehotel.com-hotel.com, part of a network that creates a dedicated, official-looking page for individual hotels on com-hotel.com subdomains. For the demolished Curium Palace it shows a green countdown banner promising a rate 10 percent below Booking.com, expiring in 23 hours, above the same permanent "sold out". To be clear, we make no claim about this operator: it may well be a legitimate booking affiliate, and we do not know whether the business behind it is genuine or not. What we can say is what the page shows, on the day we captured it, for a hotel that no longer exists. The same network operates pages for most of the 46 hotels in our sample, in both countries, and they rank in ordinary Google searches alongside properties as prominent as the Grande Bretagne in Athens.
curiumpalacehotel.com-hotel.com, captured 31 August 2026: a countdown offer and a permanent "sold out" for a hotel that has been demolished.
This is why we call the failure inheritance rather than hallucination. ChatGPT saw live, current listing pages and concluded that "availability is being offered today". It cannot tell a hotel from a zombie listing, because the distribution ecosystem never switches anything off: pages persist, reviews persist, urgency banners keep counting down over inventory that will never exist again. The AI repeated the ecosystem with a straight face. Google AI Mode, drawing on Google's own knowledge of the demolition, was the only one that told the traveller the truth.
The same pattern held for a flagship Rhodes hotel closed since 2024 for a full rebuild: the engines that read the operator's own website got it right; the one answering from stale memory described its seawater pools in the present tense.
That is the finding that matters for the agentic booking debate. The failure mode of AI travel advice is not hallucination. It is inheritance. AI answers are only as good as the listings layer underneath them, and that layer is full of pages nobody switched off.
What this means if you run a hotel
First, the reassuring part: the facts you would expect to be safe are safe. Nobody is misstating your star rating or telling travellers you are closed when you are open.
Second, the actionable part. The errors that do occur are concentrated in exactly the facts hotels are sloppiest about publishing: pet policy, parking, wellness facilities, what "beachfront" means at your property. In our data, a third of AI errors matched a wrong or stale value sitting right now on one of the hotel's own five main listings. Most of the rest traced to the wider web: tour operator pages, aggregators, directories, describing a version of the hotel that no longer exists. AI does not have an opinion about your hotel. It has your paper trail.
Third, the part nobody can fix alone. The same question, asked three times, produced a different answer about your hotel roughly one time in six. Different engines disagree with each other more than that. You cannot manage this channel by checking it once; it moves.
Our take is the same one we have been making all year: hotels do not have a marketing problem, they have a visibility problem, and AI has just raised the price of it. Before the industry wires payments into this channel, the information underneath it needs to be watched the way rate parity is watched. Not because the AI is careless, but because it is faithful: faithful to every stale listing, dead page and clone site you have ever left behind.
The booking agents are coming either way. The hotels that win that transition will be the ones whose facts are already straight when the agent arrives.
Next in this series: we hand the booking to the agent and follow where it goes.
Method note. 46 hotels (4-5 star), 5 markets (CY + GR), 3 engine conditions (ChatGPT GPT-5 with and without web search via API capture, Google AI Mode), 3 runs each, captured 31 August 2026. Ground truth triangulated across the hotel website, Booking.com, Google hotel data, Google Business Profile and Tripadvisor, with manual resolution of disputed facts against hotel sites and official announcements; airport driving distances were excluded where routing references disagreed. Four closed or renamed trap hotels were scored separately. All 450 raw captures and grading records archived. Full findings and the fact-by-fact tables are available from Tharro on request.
FAQ
How accurate are AI answers about hotels?
In our August 2026 audit, ChatGPT and Google AI Mode got 96 to 97 percent of individual hotel facts right. But because each answer contains many facts, 23 percent of complete answers, close to one in four, contained at least one wrong claim. Star rating, opening status and adults-only policy were never wrong; pet policy, parking, spa and beachfront claims carried nearly all the errors.
Does web search make ChatGPT more accurate about hotels?
Yes. ChatGPT with web search had the lowest error rate in our test (2.5 percent of asserted facts, 16 percent of answers with an error) versus 4.5 percent and one answer in three for ChatGPT answering from memory alone. Web search also fixed most closed-hotel cases, with one exception: where OTA listing pages for a dead hotel are still live, search-enabled ChatGPT read them as availability.
Where do AI errors about hotels come from?
Mostly from the information layer underneath, not from invention. About a third of the wrong claims in our study matched a wrong or stale value published on one of the hotel's own five main listings, and others traced to tour operator and aggregator pages. The clearest pattern was inherited: a demolished Limassol hotel was described as bookable because five booking platforms still carried its listing.
What can a hotel do about wrong AI answers?
Fix the source layer and then watch it. Publish the ambiguous facts explicitly on your own site (pet policy, parking, what beachfront means at your property), correct stale OTA and aggregator listings, retire dead pages, and monitor what the AI engines actually say about you on a recurring basis, because the same question can get a different answer from one week to the next.
Tharro tracks what ChatGPT, Google AI Mode and Perplexity say about your hotel, and what they get wrong. See your own AI answers checked against your facts on the AI visibility page, or start with the free AI Visibility Report.


