Skip to article

One hotel, five names: why catalogue size lies

Signing another hotel supplier grows the listing count on day one. Whether it grew the catalogue depends entirely on work nobody has scheduled yet: deciding which of those listings are hotels you already had.

Adding a hotel supplier is one of the few procurement decisions that can make a product worse while making its headline number better. The listing count jumps the day the feed connects. Whether the catalogue grew depends on work nobody in that meeting has scheduled: deciding which of the new listings are hotels you already had.

A bedbank calls it "Grand Hyatt Dubai." A second supplier sends "Hyatt Grand, Dubai." A chain feed shouts "GRAND HYATT DXB", abbreviates the road and drops the pin on the city centroid, kilometres from the door.

Three records, one building. Let them reach a traveller as three cards and the extra supply has produced a worse product than one supplier alone. Reviews split three ways. One card reads sold out while its twin has rooms. The customer spends their attention comparing a hotel against itself. Every dashboard counts that demand three times, and the growth chart looks excellent.

Hotel mapping collapses that back down without erasing the differences that are real. It is not a cleanup pass after aggregation. It decides whether aggregation was worth doing.

BEDBANK A · HTL99213 BEDBANK B · AE-3321 CHAIN FEED · GHDXB "Grand Hyatt Dubai" rooftop pin · +971 4 555 0182 "Hyatt Grand, Dubai" pin 40 m off · phone missing "GRAND HYATT DXB" pin on city centroid CANONICAL PROPERTY prop_7f21 one card in search fields merged per authority three rates compete MISSED MATCH → DUPLICATE CARD FALSE MATCH → WRONG BUILDING
Duplicate-listing collapse. Three supplier records converge on one canonical property; the listings stay attached underneath it.

The machinery that does the collapsing has its own essay: How we solved room mapping. This one is the product argument.

Mapping is three problems wearing one word

Ask two teams what mapping means and you get two answers.

LayerQuestionCost of getting it wrong
Property identityAre these listings the same physical property?Duplicate cards or merged neighbouring hotels
Room identityAre these rate plans for the same sellable room?False price comparisons and disappointed guests
Content authorityWhich source should provide each field?Stale names, weak photos or incorrect amenities

The layers interact and cannot share one similarity score. Two hotels can share a resort address and a postcode and still be two hotels. Two room names can differ by one word and by an entire bed. Authority is per field: best photos and best coordinates rarely come from the same source. So when a provider quotes one mapping accuracy figure, ask which layer it describes. It is almost always property identity, the easiest of the three.

The canonical record is not a supplier record

A canonical property is not the best listing promoted to master. It is a separate, stable entity that listings attach to, and it outlives all of them.

Useful evidence includes:

  • Geographic distance and coordinate confidence.
  • Normalised street address and postal code.
  • Telephone numbers and domains.
  • Chain, brand and property identifiers.
  • Name tokens after removing generic words and location suffixes.
  • Neighbourhood and landmark context.
  • Content evidence such as matching photos.

Every one of those lies. Coordinates sit on city centroids. Names go stale the day a property rebrands and the feed does not. Addresses differ by language and by entrance. There is no reliable signal to find, only the discipline of combining unreliable ones and tracking how unreliable each is.

So layer the resolver:

  1. Blocking generates plausible candidates using location, brand or address so every listing is not compared with every property.
  2. Feature scoring evaluates independent evidence instead of raw string similarity alone.
  3. Conflict rules prevent dangerous merges, such as distinct unit numbers or known properties in one complex.
  4. Decision thresholds separate auto-match, human review and new-canonical creation.
  5. Versioned provenance records why the match happened and which model or rule decided it.

High-confidence pairs merge automatically. Ambiguous ones get a queue with a map, side-by-side content and a reversible decision. Record the evidence, not just the score: a score without its inputs is unauditable the moment a supplier changes their feed, which they will, without telling you.

Precision and recall do not cost the same

A missed match produces a duplicate card, which is embarrassing. A false match produces a traveller at the wrong front desk with a confirmation the clerk cannot find. Those two do not belong behind one threshold. Push automatic merges hard toward precision; generate candidates at a much lower one.

Calibrated well, the split is not close. Across OnArrival's graph of 2M+ properties, as of February 2026, around 97% of listings resolve automatically and fewer than 0.1% ever surface as duplicate cards. The rest goes to a human review queue, and every reviewed pair feeds the next calibration round, which is why the review band narrows over time instead of growing with the catalogue.

Measure more than an overall accuracy percentage:

  • Precision of automatic matches.
  • Recall of known duplicate pairs.
  • False merges by geographic density.
  • Duplicate exposure in customer search results.
  • Time to resolve review-queue items.
  • Reopened mappings after supplier or brand changes.

Every one of those is visible from the customer surface, not just from a model score. If your only mapping metric lives in a notebook, you do not know how many duplicate cards a customer saw yesterday.

Room mapping is the sharper edge

Property identity answers which building. Room mapping answers which room inside it, and that is where money and trust actually move. Three strings from one hotel:

  • "Garden View Double"
  • "Standard Garden Facing"
  • "Twin Room, Garden Side"

Any two might be the same inventory. Any two might differ by one double bed versus two singles. Merge carelessly and the cheapest rate looks comparable when it is not. That mistake never surfaces in a dashboard. It surfaces at check-in, to a person who booked in good faith.

So parse attributes before comparing names: beds, occupancy, view, wing, size, smoking policy, accessibility, board basis, refund conditions. Then separate room type from rate plan. A deluxe king carries refundable, non-refundable, breakfast and member rates, and those should compete inside the correct room identity rather than pose as different rooms.

Where the evidence runs out, a conservative "similar rooms" group beats a confident wrong merge. Unknowns are asymmetric: a missing view can merge with a stated view when everything else agrees, but a conflicting bed configuration is a veto, because bedding is what guests feel most when it is wrong. The attribute grammar behind those rules is in the room-mapping deep dive.

Merge fields, not records

The golden-record pattern, where one supplier is crowned master and everything is taken from it, throws away most of what you paid for. A chain-direct feed has the current name and the official amenities. A mapping vendor may have cleaner geography. A wholesaler may hold the only rate you cannot get elsewhere and the worst photography in the catalogue.

Build lineage per field instead:

FieldPossible authority rule
Property nameVerified chain/direct source, then freshest trusted supplier
CoordinatesVerified geocode with confidence and address agreement
PhotosHighest-quality deduplicated set with source rights retained
AmenitiesNormalised taxonomy with per-source evidence
DescriptionAuthoritative, current and language-appropriate source
CancellationAlways rate-specific; never copied across suppliers

The canonical entity ends up looking more like this than like any supplier payload:

{
  "property_id": "prop_7f21",
  "fields": {
    "name":        { "value": "Grand Hyatt Dubai", "source": "chain_direct", "as_of": "2026-07-12" },
    "coordinates": { "value": [25.2262, 55.3325], "source": "verified_geocode", "confidence": 0.98 },
    "amenities":   { "taxonomy": "canonical_v4", "evidence": ["bedbank_a", "chain_direct"] }
  },
  "listings": ["bedbank_a:HTL99213", "bedbank_b:AE-3321", "chain_crs:GHDXB"]
}

Every displayed value must be replaceable without changing the property ID. That is how a rebrand updates content while saved trips, reviews and analytics keep pointing at the same hotel.

Image deduplication is comprehension, not tidiness

Feeds repeat the same lobby in four crops, and eight near-identical lobby shots tell a traveller nothing about whether the bathroom has a step in it. Group near-duplicates on perceptual similarity, sharpness and scene classification, pick one representative, then balance the gallery across exterior, room, bathroom, dining and amenity. Do not hand the ordering to an aesthetic score: an honest photo of the real room view beats a fifth polished lobby, and a photo of an accessible bathroom beats both.

Rate comparison comes last, and only after identity

"Lowest rate" should be a claim about a comparable total, not the smallest number extracted from a feed. Once property and room identity are stable, normalise:

  • Total price for the stay.
  • Included and excluded taxes or fees.
  • Currency and conversion context.
  • Payment timing: now or at property.
  • Cancellation schedule and deadlines.
  • Meal plan and occupancy.
  • Loyalty eligibility or member restrictions.

Then keep supplier fulfilment data behind whichever rate was chosen. The customer sees one hotel and options they can genuinely compare; the order system knows exactly who has to confirm, service and refund the booking.

Mapping does not remove choice. It removes accidental repetition so the meaningful choices become visible.

Change is an event, not an exception

Properties rebrand, split towers, close for refurbishment and move between chains. Suppliers reuse identifiers on properties that no longer exist. A catalogue without history survives none of that.

Store mappings as versioned relationships with provenance. Make merge, unmerge, alias and successor first-class operations rather than database surgery. Before a merge lands, look at what it touches: live bookings, saved items, reviews, analytics.

Then monitor for:

  • Large coordinate moves.
  • Sudden name and brand changes.
  • Conflicting phone or domain evidence.
  • A surge of unmatched listings in one market.
  • Previously mapped rooms drifting in attributes.
  • Customer searches that expose near-identical cards.

Mapping is never finished. It is a maintained knowledge graph wired to live commercial inventory, and the maintenance is the product.

Bring a rigged test set

Do not evaluate a mapping platform with a clean list of famous hotels. Send a rigged one:

  1. One property represented by several suppliers.
  2. Two neighbouring hotels in the same complex.
  3. A recently rebranded property.
  4. Equivalent rooms with inconsistent names.
  5. Similar rooms with different bedding.
  6. Rates that differ in taxes, breakfast and cancellation.

Ask to see the canonical property, the match evidence, the review path, the room grouping, the field provenance and the rate comparison. Then unmerge a pair you deliberately broke, and ask which saved trips, analytics rows and live offers repaired themselves.

That last step is the whole test: it separates a maintained catalogue from a one-way import with a search box on top. A large catalogue is not automatically a coherent one, and the quality lives in identity decisions the traveller never has to notice.