Research4 min read

Your Review Score Isn't Comparable to Your Competitor's

By Cosmin Costean
LinkedIn
Data card comparing two hotels whose review scores come from different platform mixes, with a gap of zero point zero seven between them

You are 0.07 behind the hotel next door. Both numbers are correct. The comparison is not.

Review benchmarking looks like the most objective thing in hotel marketing. Everyone has a score, the scores are public, and the arithmetic is trivial.

Then you look at what the scores are made of.

The problem

We compared a resort against its two closest competitors on a blended review score. The subject sat 0.07 behind the leader β€” close enough to feel actionable, and the sort of gap that generates a quarterly initiative.

Here is what the blend was actually made of. The subject's score was dominated by one platform, where it held the overwhelming majority of its reviews. The leader's score came mostly from two different platforms, because its presence on the first was thin to the point of being negligible.

So the comparison was: a score built mostly from Platform A, against a score built mostly from Platforms B and C. Same league table, different sports.

Why the platforms are not interchangeable

Guest populations differ. Platforms attract different traveller types, from different countries, booking through different channels. A hotel strong with one nationality can look different on two platforms for reasons that have nothing to do with the hotel.

Scoring behaviour differs. Some platforms' users cluster near the top of the scale; others spread out. A given experience produces different numbers depending on where it is recorded.

Prompting differs. Platforms that solicit reviews aggressively after every stay produce a different sample from platforms where reviewing is self-motivated β€” and self-motivated reviewers skew toward the extremes.

None of this is anyone's fault. It just means a 0.07 gap between two differently-composed blends is not a 0.07 gap in guest satisfaction.

The measurement trap underneath it

There is a second problem that is easier to fix and more embarrassing to find.

Some review data reflects the platform's own total. Some reflects only what your tooling managed to collect. When those two get blended without distinction, a platform where you have a genuine total of thousands and a platform where a scraper collected a couple of hundred are weighted as if they were the same kind of number β€” and the platform with a real total quietly dominates the blend.

The visible symptom is a recommendation like "this platform is thin, only 149 reviews β€” one bad night is visible." It sounds precise and actionable. It may be measuring how deeply the data was collected rather than how many reviews exist.

Ask your provider one question: is this the platform's total, or your sample? If they cannot answer per platform, the depth verdicts are not reliable.

What to do instead

Compare per platform, not on a blend. Your score on Platform A against their score on Platform A is a real comparison. Blended against blended usually isn't.

Show the composition. Any comparison should display each hotel's review count per platform beside the score. The asymmetry becomes obvious in one glance and nobody has to be warned about it.

Flag missing baskets explicitly. If a competitor has effectively no presence on a platform, say so on the screen rather than quietly producing a number.

Compare category scores. Cleanliness, location, staff, value are more robust across platforms than the headline, because they measure something narrower.

Watch your own trend. Your score against your own score last quarter is the one comparison with no basket problem at all.

What this does not mean

It does not mean review benchmarking is useless. Large gaps are real: a hotel at 8.2 against a comp set at 9.1 has a genuine problem no methodology explains away.

It means small gaps are noise, and a strategy built on closing 0.07 is a strategy built on a measurement artefact. Reserve the effort for gaps that survive being looked at properly.

FAQ

How big does a gap need to be before it means something? Depends on volume, but as a rule of thumb, if the gap is smaller than the difference between your platform scores, it is inside the noise.

Should I push guests toward one platform? Only with a commercial reason β€” visibility on a channel that drives bookings for you. Doing it to improve a benchmark optimises the measurement, not the hotel.

Do AI assistants read review scores? They read review content more than scores, and across multiple platforms. Consistency across platforms matters more than the number on any one.

What about review recency? It matters and is usually ignored. A score built mostly on reviews from two years ago describes a hotel that may no longer exist.


Want your review standing compared properly β€” per platform, per segment, against your comp set? See how Tharro does it or book a 30-minute call.