Blog · arXiv Analysis · Modified August 12, 2026 · Last reviewed August 12, 2026

The Hotel List Position Becomes the Booking Clerk

The June 2026 arXiv preprint Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection, by Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, and Asher Ali, provides controlled evidence that hotel-card attributes and display order can change an LLM's selection. It does not audit a live travel product or observe a real booking.

For this essay, an AI booking clerk is the decision pipeline that turns a travel request into an offer the user can accept: candidate retrieval, eligibility filtering, ordering, model selection, explanation, price and term presentation, booking handoff, and confirmation. The paper isolates selection from five already-retrieved cards. Governing the clerk requires evidence at every other stage too.

The Assistant Is the Shop Window

The arXiv record for 2606.16344 [cs.AI] lists one version, submitted June 15, 2026. Its object is ordinary and consequential: a traveler asks an assistant which hotel to book, and the assistant turns a supplied set of properties into one recommendation.

That recommendation allocates visibility. A conventional search page may expose alternatives, filters, ranking labels, advertisements, prices, and snippets at once. A conversational interface can compress that market into one named property and a short reason. The paper calls the system an AI infomediary: an intermediary that receives supplier information, weights it, and presents a choice to a traveler.

The distinction between candidate set and choice is essential. Retrieval decides which hotels are eligible to be considered; ordering decides how those hotels reach the chooser; selection names the preferred hotel; transaction logic decides what is actually offered and charged. A clean audit of selection can identify a model response to order, but it cannot tell us why a property was absent upstream or whether the downstream offer remained accurate.

This makes the study a useful companion to the site's work on recommender systems, AI search and answer engines, agentic commerce, algorithmic transparency, and search-agent recommendation reliability. Its fresh angle is the apparently neutral order of options inside a recommendation prompt.

Current Context

By August 12, 2026, conversational hotel choice had already reached product surfaces rather than remaining only a laboratory scenario. Booking.com announced an LLM-assisted Trip Planner in 2023 that surfaced destinations and properties with pricing and deep links into its reservation flow, and in 2024 described broader availability plus Smart Filter, Property Q&A, and experimental review summaries. In November 2025, Google described AI Mode travel planning with real-time hotel data, Maps information, comparisons, and follow-up questions. Google said direct flight and hotel completion was a future capability in that announcement; it did not say hotel booking was already live. These company posts establish product direction, not current availability in every market, prevalence, accuracy, neutrality, or equivalence to the models in the paper.

Short-term lodging has a specific U.S. price rule. The FTC's Rule on Unfair or Deceptive Fees took effect May 12, 2025 and covers businesses that offer, display, or advertise lodging prices, including third-party platforms and travel agents. If a price is shown, the upfront total must include mandatory charges the business knows and can calculate. Government charges such as taxes and genuinely optional additions may be excluded initially, but their nature and amount, and the final amount of payment, must be disclosed before payment. Cancellation and refund terms remain separate material facts that a trustworthy clerk should carry into the comparison.

Enforcement makes the transaction boundary concrete. In July 2026, the FTC filed a complaint and proposed stipulated order against Hopper. The agency alleges hidden, preselected fees, misleading total prices, and misrepresented limitations of support and price-freeze products; the proposed $35 million order requires judicial approval. Those are allegations, not adjudicated findings, but they show why a good recommendation can still become a bad booking when the handoff changes price, consent, or terms.

European ranking duties are layered and in motion. The Platform-to-Business Regulation currently requires covered intermediaries and search engines to describe main ranking parameters, their relative importance, and how remuneration can affect rank. The European Commission proposed repealing that regulation in November 2025, but the ordinary legislative procedure remained ongoing on this page's review date. Separate EU consumer-law amendments require marketplace information about main ranking parameters and treat undisclosed paid higher placement in search results as a prohibited practice. The Digital Services Act requires covered online platforms using recommender systems to explain their main parameters and user controls; very large platforms and search engines must offer at least one option not based on profiling.

Reputation and commercial influence need their own controls. The FTC's Consumer Reviews and Testimonials Rule, effective October 21, 2024, addresses specified business conduct involving fake or false reviews, sentiment-conditioned incentives, undisclosed insider reviews, review suppression, and fake influence indicators. It does not impose a general duty on a platform merely hosting reviews to investigate every submission. FTC native-advertising and endorsement guidance separately makes the disclosure question depend on the commercial relationship and the net impression on consumers. A model should not turn a commission-weighted or paid candidate list into the voice of a neutral travel adviser.

What the Audit Tests

Baig, Gillani, and Ali use a randomized choice-based conjoint: on each trial an instruction-tuned model must select one of five synthetic hotel cards. Every property is described as a four-star hotel near the city center. The randomized fields are rating (3.9, 4.3, or 4.7), review count (45, 420, or 2,100), most-recent-review age (three days or eleven months), a management-response line, chain or independent status, price ($129, $189, or $249 per night), a Green Key certification line, and display position.

The main estimand is an average marginal component effect, or AMCE: the average change in a card's probability of selection when one randomized attribute level changes, averaged across the other randomized attributes. Randomization supports a causal interpretation inside this task and model panel. An AMCE is not the probability that an individual traveler would book, a quality score for a hotel, or a production conversion rate.

The audit uses three scripted personas, nine prompt paraphrases, and twelve fixed open-weight and proprietary model configurations. The paper reports 3,024 main-arm choice sets per model and 61,459 calls across the main and secondary arms. It says the design, hypotheses, and analysis plan were fixed and hashed before main-arm collection. Seven attribute effects and five persona or signal interactions form the Holm-corrected confirmatory family; the position effect, per-model estimates, rationale comparison, and arm contrasts are preplanned but exploratory.

In the pooled main arm, a 4.7 rating rather than 3.9 increased selection by 31.65 percentage points (95% confidence interval 30.94 to 32.36), while a $249 rather than $129 price decreased it by 30.03 points (95% confidence interval 29.32 to 30.75 points in the negative direction). Green Key certification added 11.61 points and 2,100 rather than 45 reviews added 8.31. The management-response line changed selection by 0.11 points (95% confidence interval −0.51 to 0.73); the paper's equivalence test also placed it within a predeclared ±1.5-point band.

The paper reports a 99.98% parse-success rate and broadly stable attribute patterns across several prompt, temperature, formatting, and retest conditions. One caution survived those checks: the supposedly neutral hotel-name placebo was jointly significant. Counterbalancing protects the randomized attribute estimates from name assignment, but the result is a reminder that even synthetic labels were not perfectly inert.

Position Is Not Empty

List-position bias here means a change in selection probability caused only by a card's ordinal slot, with its hotel attributes randomized independently. Relative to slot one, the pooled model selected slots two through five 2.07, 2.38, 3.59, and 3.68 percentage points less often. Position is therefore content-free in the experiment; a live system's order may instead encode relevance, sponsorship, commission, inventory, or another upstream rule.

The pooled result should not be flattened into “LLMs choose the first item.” The paper reports that most audited models were nearly position-neutral, while one configuration, Gemini 2.0 Flash, showed an approximately 26-point first-position advantage. The pool therefore conceals substantial model heterogeneity. The paper also could not complete its planned decomposition of display-position effects from option-letter token preferences because usable token log-probabilities were unavailable.

The authors convert the first-slot advantage over the average of slots two through five into $11.7 per night, with a reported 95% interval of $8.6 to $14.9. They do this by dividing the position AMCE by a per-dollar price effect linearized over the experiment's $129-to-$249 range after a monotonicity check. This is a model-choice price-equivalent: it is not traveler willingness to pay, a fee charged by a platform, the revenue value of rank, or measured consumer harm.

The governance implication survives those qualifications. A retrieval engine, booking platform, partner feed, advertiser interface, or prompt constructor can set candidate order before the model speaks. If the chooser is order-sensitive, upstream arrangement becomes part of the recommendation. The appropriate response is not supplier folklore about “optimizing for the AI.” It is counterfactual testing, paid-influence disclosure, stable records, and an interface that lets the traveler inspect alternatives.

Reasons Are Not the Weights

The rationale result is mixed, not a finding that explanations are useless. Across models, the rank correlation between how often an attribute appeared in reasons and its revealed AMCE importance ranged from +0.59 to +0.85, with a reported median near +0.69. The reasons broadly tracked behavior, but not completely.

The gaps matter. List position appeared in no more than 0.7% of reasons despite its pooled revealed effect, and review volume was under-mentioned. Brand or affiliation was sometimes over-mentioned relative to its 2% revealed importance share, reaching 55% of reasons for one model. The paper obtained headline mention rates with a keyword dictionary across 35,223 parseable reasons and used a separate generative judge as a cross-check.

The requested JSON placed the hotel choice before the one- or two-sentence reason. That makes the text useful for comparing stated and revealed importance, but it does not expose the model's internal causal process. A fluent rationale is a generated account, not an execution trace. Buyers, regulators, and suppliers therefore need behavioral tests and pipeline records in addition to prose explanations.

What It Does Not Prove

The experiment begins after retrieval. It does not test which hotels enter the candidate set, how a live platform ranks them, or whether a missing property is absent because of geography, data quality, inventory, commission, sponsorship, or a commercial agreement. Nor does it test multi-turn constraint gathering, live availability, checkout, payment, cancellation, or post-booking support.

The cards are synthetic, single-turn, English-language, U.S.-dollar descriptions with fixed star class and central location. They omit photographs, amenities, free-text reviews, room types, taxes, accessibility detail, loyalty benefits, neighborhood context, and correlated real-world attributes. That stylization enables identification; it also means absolute recommendation rates and trade-offs should not be transported directly into a live marketplace.

No travelers participated, no reservations were made, and no supplier exposure, conversion, revenue, welfare, or discrimination outcome was observed. The comparison with prior human hotel-choice literature is an ordering and relative-importance comparison, not a like-for-like human conjoint using the same stimuli. “Over-weighting” or “under-weighting” therefore remains the paper's interpretation against a non-commensurable literature benchmark.

The panel is not a census of products. It uses fixed model versions under a standardized harness, including quantized local models; deployed products can add retrieval, proprietary prompts, tools, safety policies, interface logic, and later model updates. The large effect in one configuration and near-neutral behavior in most others makes versioned, per-system reporting more informative than the pooled label “LLM.”

A production booking assistant adds both factual and safety obligations: it must distinguish verified accessibility features from generated inference, avoid unsupported claims about neighborhood safety, carry hard occupancy and date constraints before soft preferences, expose commercial influence, and revalidate the offer before the user authorizes payment. Those requirements are not tested by this paper.

Failure Modes

Candidate-set omission appears when an assistant offers a polished choice from five hotels while hiding that eligible alternatives were excluded upstream by incomplete data, partner coverage, inventory access, or commercial rules.

Position laundering appears when upstream order is treated as neutral formatting and then reappears as the assistant's confident recommendation.

Reason theater appears when a plausible rationale is presented as a causal explanation even though no behavioral test or trace connects the stated factor to the choice.

Commercial invisibility appears when paid placement, commission, preferred-partner status, loyalty economics, or inventory deals shape retrieval or order but are absent from the answer.

Review-signal laundering appears when rating, volume, recency, or certification is treated as reliable without preserving its source, verification status, coverage period, or manipulation controls.

Total-price blindness appears when the clerk compares nightly rates without known mandatory charges, or reaches payment without prominently showing taxes and other permitted exclusions in the final amount. Cancellation, refund, room, and merchant terms must remain visible even though they are not all governed by the FTC fee rule.

Safety fabrication appears when the model invents or overstates accessibility, occupancy suitability, neighborhood safety, distance, or amenity facts instead of retrieving a dated, attributable property record and flagging uncertainty.

Supplier opacity appears when a hotel cannot distinguish exclusion by retrieval, demotion by ranking, loss at selection, stale listing data, or model order sensitivity. Persona overreach appears when a system infers vulnerability or protected characteristics and silently turns those inferences into ranking features instead of asking for relevant constraints and offering a non-profiled comparison.

Governance Standard

A booking clerk needs two gates and one reconstructable record. The recommendation gate must satisfy hard constraints such as dates, occupancy, user-specified accessibility, availability, and stated budget before optimizing soft preferences. The transaction gate must recheck inventory, total price, room and cancellation terms, merchant of record, and booking authority immediately before explicit user confirmation. Neither gate should treat generated prose as verification.

The recommendation record should preserve the query and consented preferences; candidate-source and exclusion rules; retrieval and ranking versions; the exact ordered shortlist; paid, affiliate, loyalty, and preferred-partner status; listing and review provenance; model, prompt, tools, and configuration; attributes shown and withheld; selected property; user-visible reason; offer timestamp; and final handoff. It should be detailed enough for an incident review without retaining sensitive travel or disability information longer than necessary.

Counterfactual testing should permute card order, neutral identifiers, equivalent wording, price presentation, commercial labels, and persona text while holding hotel facts fixed. Report pooled and per-version effects, confidence intervals, invalid-response handling, drift across updates, and whether a few configurations dominate an average. Run potentially confusing permutations in a controlled evaluation environment, not as an unannounced experiment that degrades live consumer choices.

The traveler should be able to inspect the shortlist, change the ranking basis, distinguish hard facts from summaries, see why sponsored or commission-bearing inventory is present, and select a non-sponsored or non-profiled comparison where required or offered. Before payment, the interface should show the total required lodging price most prominently, disclose permitted exclusions and the final amount at the proper stage, restate cancellation and refund terms, and require confirmation of the exact room and seller.

Supplier recourse should operate at the pipeline stage where the error occurred. A property should be able to correct stale rates, amenities, accessibility details, review aggregation, or affiliation; learn whether it failed retrieval, ranking, or selection; receive notice of material ranking-rule changes where applicable; and request a rerun after correction. It need not receive source code to receive a meaningful reason and repair path.

Ranking safeguards should also cover signal integrity. Review volume is not quality if the reviews are fake, duplicated, suppressed, or drawn from an unexplained window; a sustainability badge is not evidence without issuer and validity data; and “safe,” “accessible,” or “family-friendly” is not a free-form attribute the model may improvise. The clerk should cite the underlying record, state uncertainty, and route unresolved high-impact constraints back to the traveler or property.

The Spiralist rule is this: when a chat assistant recommends one option, the list position has already spoken. Audit the clerk before trusting the booking.

Source Discipline

This article treats the study as arXiv version 1, a 32-page preprint submitted June 15, 2026, not as peer-reviewed or independently replicated evidence. Quantitative claims are explicitly paper-reported results from a controlled harness. This review did not rerun the experiment. The arXiv record showed no later version on August 12, 2026.

The evidentiary boundary inside the paper matters. The attribute hypotheses are confirmatory under the paper's analysis plan; position, per-model effects, rationale comparison, and robustness-arm contrasts are exploratory even though the authors describe them as preplanned. The paper says its plan and design were hashed before main collection, but readers cannot verify that claim from a public preregistration link. Its data-availability statement says code and data are available from the authors on reasonable request, while the methods say a public archive will be created.

Competing-interest context is disclosed rather than hidden: the paper states that Baig and Ali are affiliated with Fandaqah, a travel and hospitality company, and that Fandaqah provided computing and API resources. That does not invalidate the experiment, but it raises the value of public artifacts, independent replication, and field audits that measure traveler and supplier outcomes.

Regulatory and platform sources answer different questions. Company announcements establish what their authors said products could do at the stated date, not independent performance. The FTC fee rule establishes a lodging-price baseline, not an LLM-ranking rule; the Hopper filing contains allegations and a proposed order, not a final adjudication. EU duties depend on service type and jurisdiction, and the proposal to repeal the P2B Regulation remains a proposal while the cited legislative procedure is ongoing. Advertising and review sources establish disclosure and conduct rules within their scope, not that a named assistant violated them.

Internal Spiralism links below are editorial context, not external evidence. Primary sources support factual claims; related pages extend the argument about ranking, agentic commerce, pricing, and recommendation governance.

Sources


Return to Blog