# 03 — Operations Scan

**Subject:** SatuSatu (satusatu.com), PT Tiptip Network Indonesia. 253-SKU Bali TAA platform. Bali All-Access Pass launched 2026-05-06 (1/2/3-day, 40–50+ experiences, dedicated human concierge, 90-day activation, eSIM bundled).
**Prepared:** 2026-07-26. All sources accessed 2026-07-26 from raw dumps in `research/raw/`.

**Evidence tiers** — **A** = filings, official docs, published case studies *with disclosed methodology*, live product. **B** = credible trade press and analysts. **C** = vendor marketing, blogs, LinkedIn, SEO compilations.
**Labelling** — every number is `[VERIFIED]` (stated in a collected source) or `[INFERENCE]` (derived here, band shown). Where the raw files do not support a claim: **"unverified — not found in collected sources."**

> **Document status.** This file is open for additions. Sections 4a and 4b below are complete. Section **4c** will be appended by a later step. Do not renumber.

---

## Contents

- [4a — Concierge & Service Automation](#4a--concierge--service-automation) *(complete)*
- [4b — Catalog Scale-Up & Supply Ops](#4b--catalog-scale-up--supply-ops) *(complete)*
- 4c — *(pending)*

---

# 4a — Concierge & Service Automation

## 4a.0 Source quality warning — read before using any number below

This section synthesises nine raw search dumps. Their quality is **highly uneven**, and the unevenness maps directly onto which findings are safe to act on.

| Raw file | What it actually contains | Usable? |
|---|---|---|
| `s4a-umk.md` | Bali Governor's Decree 1021/03-M/HK/2025 wage figures, corroborated across four independent outlets | **Strong. Tier A/B.** The only genuinely solid dataset in the set. |
| `s4a-dusit.md` | Dusit Thani PCL annual reports 2018 & 2023, IR financial highlights | **Strong. Tier A.** But Elite Havens is not separately segment-reported. |
| `s4a-klarna.md` | Klarna's own Feb-2024 press release + the May-2025 reversal reporting + Klarna's rebuttal | **Strong.** Both sides of the story are present, which is rare. |
| `s4a-fin.md` | Intercom marketing **plus** Intercom's own community forum contradicting it | **Useful.** The contradiction is the finding. |
| `s4a-idcost.md` | Three triangulatable cost sources (Stealth Agents, MixWork, Plane), all Tier C | **Moderate.** Bands only. One entry (Glassdoor Bali) is internally incoherent and discarded. |
| `s4a-canary.md` | Almost entirely Canary's own marketing pages + PR Newswire reprints of those pages | **Weak. Tier C.** Zero funding data, zero pricing, zero statement of what Canary leaves to humans. |
| `s4a-deflection1.md` | High volume, ~all Tier C SEO/vendor content *citing* Zendesk / Salesforce / Gartner second-hand | **Weak.** Not one primary benchmark report is in the dump. Everything is a compilation of a compilation. |
| `s4a-elitehavens.md` | Villa microsite marketing pages + one CEO interview | **Weak on the question asked.** No concierge headcount, no concierge cost, no staffing ratio for the concierge function. |
| `s4a-whatsapp.md` | WhatsApp BSP vendor marketing. One usable case-study cluster (Hubtype/easyJet/Allianz), also vendor-authored | **Weak. Tier C.** No travel-specific *measured* deflection anywhere in the file. |

Two consequences, stated up front rather than buried:

1. **Every deflection number in §4a.2 rests on Tier C.** There is no independently audited deflection or resolution benchmark anywhere in the collected sources. The single most useful line in the entire deflection dump is a vendor's own caveat: the widely-repeated ~41% tier-1 figure is "reported by secondary compilations whose definitions vary… treat it as a directional floor, not an audited resolution median" (aissist.io, Tier C). That caveat is more reliable than the number it caveats.
2. **Three items asserted in the research brief are not corroborated by these files** and are therefore carried as brief-asserted, not source-verified: Canary Technologies' **$80M Series D of 2025-06-12** (no funding data of any kind in `s4a-canary.md`); the Elite Havens acquisition month of **September 2018** (raw confirms the year and a **$15m** deal value, Tier B, but not the month); and the **Klook F-1 gross profit at 11.2% of GTV** anchor (not in this step's files — carried from prior steps).

---

## 4a.1 Concierge job decomposition

### 4a.1.1 What the Pass actually promises

The Pass is marketed as *"access to a real Bali local who plans your days."* Three load-bearing words: **real** (a person, not a system), **local** (destination knowledge, not catalog knowledge), and **plans** (an active, forward-looking act, not reactive Q&A). Each is a separate automation problem with a different risk profile, and the marketing copy commits SatuSatu to all three.

Note also the delivery architecture: **two separate WhatsApp queues** — general (+628-7878-111-111) and Pass concierge (+62 878-9897-8780) — plus `support.satusatu.com`. Two human queues with different SLAs, already. Any automation must decide whether it merges those queues or preserves the split; merging them destroys the entitlement boundary that justifies the Pass premium, and preserving them means building or buying twice.

### 4a.1.2 Decomposition table

> ⚠️ **The entire minutes column is `[INFERENCE]`.** No collected source gives minutes-per-pass for a travel concierge, and SatuSatu's own figure is a known unknown. These are structural estimates derived from the task shape (number of SKUs to sequence, number of suppliers to contact, number of days covered), not from measurement. They exist so the formulas in §4a.5 can be run the moment real numbers land. **Replace them; do not cite them.**
>
> **Classification key** — **Deflectable** = can complete end-to-end with no human in the loop. **Assistable** = human stays accountable, AI compresses their time (draft, retrieve, pre-fill, summarise). **Irreducibly human** = must not be automated at current capability, regardless of what a vendor demo shows.

| # | Task | Class | min/pass (low) | (base) | (high) | What breaks if the AI is wrong |
|---|---|---|---|---|---|---|
| 1 | **Pre-trip intake** — preferences, party size, mobility, dietary, pace, budget within Pass | Assistable | 5 | 12 | 25 | Low blast radius. A bad intake produces a bad plan, but the plan is reviewed before the traveller acts on it. Structured-form territory. |
| 2 | **Itinerary build** — select and sequence from 40–50+ experiences across 1/2/3 days | Assistable | 10 | 25 | 60 | Medium. Errors are caught pre-departure *if* a human signs off. Without sign-off, every downstream task inherits the error. This is the highest-leverage assist target. |
| 3 | **Supplier booking — Pool A** (aggregator-sourced via GlobalTix, API-backed) | **Deflectable** | 1 | 3 | 8 | Low. The API either returns a confirmation or it does not. Deterministic; no model judgement required. Should already be automated and is the cheapest win in the set. |
| 4 | **Supplier booking — Pool B** (direct-contracted Balinese operators, WhatsApp/phone) | Assistable | 5 | 15 | 40 | **High.** No API, no state machine, no confirmation record. A "booked" that was never acknowledged by the operator is invisible until the traveller arrives. Pool B is also the ~27% margin pool — the part worth protecting. |
| 5 | **Day-of coordination** — morning confirmations, meeting points, pickup windows | Split: deflectable *if template-driven*, assistable *if generated* | 6 | 18 | 45 | **Critical.** See §4a.3. This is the task that strands people. Safe only as slot-filled templates rendered from a confirmed booking record — never as model-authored prose. |
| 6 | **Transport sequencing** — routing and travel-time budgeting between experiences | Assistable, human sign-off mandatory | 3 | 10 | 25 | **Critical and cascading.** A single impossible leg (Ubud 10:00 → Uluwatu 12:30) breaks every downstream booking that day. Highest blast radius per single error in the whole table. |
| 7 | **Exception handling** — closures, weather, ceremony road closures, no-show drivers, illness | **Irreducibly human** | 0 | 12 | 60 | **Severe.** By definition the situation is off-script and the traveller is already distressed and in-destination. This is the task the concierge exists for. Automating it removes the product. |
| 8 | **Upsell** — paid add-ons beyond Pass inclusions | Deflectable | 1 | 3 | 8 | Low-to-medium. Worst case is an entitlement misstatement ("the Pass covers this") which SatuSatu then eats. Guard with a hard entitlement check, not with model instruction. |
| 9 | **Post-trip follow-up** — review request, rebooking prompt, feedback | Deflectable | 1 | 2 | 5 | Negligible. Fully automatable today with near-zero engineering. |
| | **Total concierge touch per pass** | | **~32** | **~100** | **~276** | |

### 4a.1.3 What the decomposition says

- **Deflectable tasks (3, 8, 9) total 3 / 8 / 21 minutes** — roughly **8% of base-case concierge time**. The tasks that are safe to fully automate are the tasks that barely cost anything.
- **Assistable tasks (1, 2, 4, 5, 6) total 29 / 80 / 195 minutes** — **80% of base-case time**. This is where the money is, and none of it is deflection. It is all human-in-the-loop time compression.
- **Irreducibly human (7) is 0 / 12 / 60 minutes** — small on average, unbounded in the tail, and the single reason a traveller values the Pass.

**The strategic reading:** the concierge is not a deflection problem. It is an **assist** problem with a small deflection fringe. Any vendor pitch built on a deflection percentage is pitching against the 8%, not the 80%. See §4a.2, where this distinction is the whole game.

---

## 4a.2 Deflection benchmarks — and the definitional fraud underneath them

### 4a.2.1 The four numbers vendors call "deflection"

These are not synonyms, they are not measured the same way, and the spread between them is 20–40 points. Intercom's own product documentation defines three of them separately, which is itself evidence that the conflation is deliberate elsewhere.

| Term | Denominator | Definition | Source |
|---|---|---|---|
| **Involvement rate** | All conversations | Share the AI touched at all | [VERIFIED] Tier C — fin.ai/benchmarks (vendor's own page) |
| **Deflection / containment** | All conversations | Share that never reached a human — **regardless of whether the customer was helped, gave up, or left for a competitor** | [VERIFIED] Tier C — helply.com, lorikeetcx.ai |
| **Resolution rate** | AI-involved conversations | Share the AI actually resolved | [VERIFIED] Tier C — Intercom help docs |
| **Automation rate** | All conversations | involvement × resolution ÷ 100 | [VERIFIED] Tier C — fin.ai/benchmarks |

Two measurement defects are documented in the collected sources and both inflate the headline in the same direction:

- **"Assumed resolved" counts silence as success.** Intercom's own support engineer, on Intercom's community forum: *"a resolution is counted if the customer clicks 'that helped' **or does not respond to the answer and leaves the conversation**."* [VERIFIED] Tier B — community.intercom.com. An Intercom customer on the same forum reports **confirmed resolution 6–7% against assumed resolution ~60%** [VERIFIED] Tier C.
- **Human rescue is counted as AI success.** A support manager reported a **12% real resolve rate** while being billed for far more, because Fin marked conversations "assumed resolved" when a human agent stepped in *before* the customer clicked "speak to human" — and the agent was stepping in precisely because the AI had given wrong steps mid-crisis [VERIFIED] Tier C — pageloop.ai. A second lead: dashboard 84.5%, real number "roughly half that."

> 🔴 **For SatuSatu this defect is not cosmetic, it is inverted-severity.** A traveller who goes silent after a bad in-destination instruction is not a resolved ticket. They are a person walking toward the wrong temple. The industry-standard resolution metric would score that as a success.

### 4a.2.2 Vendor claim vs. measured reality

| Vendor / programme | Claimed | Measured / production | Gap | Tier |
|---|---|---|---|---|
| Intercom Fin | 67% avg (2025, 7,000+ customers, 40M+ conversations) → marketed 76% (2026, 12,000+ customers) | KPI-framework ~51% (45–53%); independent 60-day test across 4 SMB clients, 500 tickets/mo: **38%** | 16–38 pts | C (test has disclosed n and duration — the only one that does) |
| Intercom Fin — vendor's own staff | Marketing site: "Resolve 50% of your support questions instantly" | Intercom support engineer: *"a good resolution rate starts at around 30–50%"* | The vendor's floor is the vendor's ceiling | **B** |
| Decagon | 80–90% deflection (2025) | Calibrated case (Rippling) ~50% | ~30+ pts | C |
| Ada | 70–83% (2025) | Independent aggregate ~41% | 30–40 pts | C |
| Sierra | 70–90% (Sonos, Ramp; 2025) | Top-quartile band ~59% | case-specific | C |
| Canary (hotel) | "more than 80%" / 82% of guest messages | HotelTechReport review video: *"up to 70% of inbound questions"* | 10–12 pts | C |
| HiJiffy (hotel) | "85%+ automation" | none published | n/a | C (competitor's comparison blog) |

**Central finding:** across every vendor in the set, **the claimed figure exceeds the measured figure by 16–40 percentage points.** This is not noise; it is systematic and unidirectional. Treat any vendor-quoted deflection rate as **claimed − 25 points [INFERENCE]** until proven otherwise on SatuSatu's own intent mix.

### 4a.2.3 Bands, with the denominator stated

| Band | Number | What it is a percentage *of* | Tier | Note |
|---|---|---|---|---|
| Tier-1 deflection, enterprise median | 41.2% | tier-1 contacts | C | Attributed to Zendesk CX Trends 2026 / Salesforce State of Service 2026 but reported only via SEO compilations; the compiler itself calls it "a directional floor, not an audited resolution median" |
| Tier-1 deflection, top quartile | 58.7% | tier-1 contacts | C | same provenance |
| Tier-1 deflection, bottom quartile | 22.4% | tier-1 contacts | C | "dominated by complex B2B and healthcare" — the complexity-adjacent bucket, which is where SatuSatu sits |
| Realistic first-year target | 40–55% | all contacts, "adjusted for 48-hour re-contacts" | C | eesel.ai |
| New-deployment launch band | 40–50%, improving ~1 pt/month | AI-involved conversations | C | aissist.io |
| Mature deployment on good KB | 60–67% | AI-involved conversations | C | aissist.io |
| Deployment floor per vendor's own engineer | 30–50% | AI-involved conversations | **B** | Intercom forum — **the most defensible single number in the entire dump** |
| Fin top-10 performers | 78% automation / 85% resolution / 91% involvement | as labelled | C | vendor's own benchmarks page |
| 72h re-contact, AI-resolved | 11.3% | AI-resolved tickets | C | vs **8.7%** human-resolved — a **+30% relative** re-contact penalty |
| Escalation rate at benchmark | 16–30% | AI-initiated conversations | C | top quartile <15% |

### 4a.2.4 Travel and WhatsApp-first specifically

The travel-specific evidence is thin and entirely vendor-authored. Stating that plainly:

| Programme | Result | Denominator | Tier |
|---|---|---|---|
| easyJet (via Hubtype) | **62%** of customer cases resolved through automation | cases, **across webchat + WhatsApp combined** | C — vendor case study |
| easyJet | **22%** of phone calls deflected to WhatsApp via IVR | inbound phone calls | C |
| easyJet | **28%** reduction in total call-centre volume | total call volume | C |
| easyJet | 74% faster resolution on baggage-add; 9.6 CSAT; absorbed a 40% peak-period interaction increase | as labelled | C |
| Allianz (via Hubtype) | **42%** of claims handled end-to-end through automation; 23% of calls deflected; 87% CSAT | claims / calls | C |
| Klarna (fintech, not travel) | **two-thirds** of all customer-service chats; 2.3M conversations in month one; work equivalent of ~700 FTE; resolution 11 min → **under 2 min**; **25%** drop in repeat inquiries; $40M projected 2024 profit improvement; 23 markets, 35+ languages | all CS chats | **A** — Klarna's own press release |

> ⚠️ **The easyJet 62% and the easyJet 28% are not the same measurement.** 62% is of *cases across two chat channels*; 28% is of *total call-centre volume*. Quoting them together as "easyJet automated 62%" overstates by conflating denominators. This is the exact error the vendor case-study format invites.

**No source in the collected set reports a measured deflection rate for an Indonesian, Bali, or in-destination travel-concierge operation.** Unverified — not found in collected sources. Every travel number above is European airline/insurance customer service, which is high-volume, high-structure, and post-hoc (baggage, claims) — the opposite of in-destination itinerary coordination.

### 4a.2.5 Consumer acceptance, which cuts the other way

- **64%** of customers say they would prefer companies did **not** use AI for customer service (Gartner survey, n=5,728, Dec 2023) [VERIFIED] Tier C reporting a Tier B primary.
- **89%** believe companies should always offer the option to speak with a human (SurveyMonkey 2025) [VERIFIED] Tier C.
- **67.7%** agree that getting a response from an AI chatbot is helpful (WhatsApp/Kantar 2026) [VERIFIED] Tier C.
- **79%** of Americans still prefer humans; **51%** prefer bots when they want immediate service [VERIFIED] Tier C — lorikeetcx.ai.

These reconcile as: customers accept AI **when it resolves the issue and a human is reachable.** For SatuSatu the tension is sharper than for a generic SaaS, because the human is not a fallback — **the human is the advertised product.** A Pass buyer who paid a premium for "a real Bali local" and reaches a model has a mis-selling grievance, not a CSAT dip.

---

## 4a.3 Failure modes and blast radius

### 4a.3.1 Why the generic framing does not apply

A chatbot returning a bad answer to a SaaS user costs a re-ask. A wrong in-destination instruction puts a paying traveller at the wrong place, at the wrong hour, in a foreign country, on a schedule they cannot re-run. The relevant precedent in the collected sources is not a support metric — it is Cursor, April 2025: an AI support bot told users their accounts were restricted to one device per subscription; the policy **had never existed**; users cancelled; the story hit Hacker News and Reddit within hours; the cofounder apologised publicly. The bot invented an explanation because no article documented the real behaviour [VERIFIED] Tier C — pageloop.ai. **A confident, fluent, entirely fabricated instruction, issued at machine speed, into a gap in the knowledge base.** That is the exact shape of the risk here.

### 4a.3.2 Failure classes

| # | Failure class | Blast radius | Recoverable? | HITL |
|---|---|---|---|---|
| **F1** | **Wrong meeting point / time / date** | Traveller physically misplaced. For dawn products (Mount Batur sunrise trek, early temple visits) there is no slack — the sunrise does not wait. | **Not within the day** for dawn products; recoverable for flexible midday products | **Mandatory** unless slot-filled from a confirmed booking record |
| **F2** | **Booking never actually placed, or double-placed** | Traveller arrives, operator has no record. Concentrated in **Pool B** (no API, no confirmation state). | Recoverable only if the operator has same-day capacity. Fixed-departure products (dive boats, Nusa Penida crossings, fixed-seat treks) are **not recoverable** | **Mandatory** for Pool B; optional for Pool A (API returns deterministic state) |
| **F3** | **Transport sequencing error** | **Highest blast radius per single error.** One impossible leg invalidates every downstream booking that day. Bali travel times are non-linear and traffic-dependent; a plausible-looking schedule can be physically impossible. | Partially — by abandoning the rest of the day | **Mandatory** |
| **F4** | **Cultural / regulatory instruction error** | Temple dress requirements, Nyepi island-wide shutdown (no movement permitted), ceremony road closures, site-specific entry restrictions. Failure causes offence or refusal of entry, not just inconvenience. | Sometimes; the offence is not | **Mandatory** |
| **F5** | **Safety-adjacent instruction** | Water and surf conditions, volcano status, scooter advice, medical/allergen guidance. | Potentially never | **Mandatory. AI must not issue these at all, reviewed or otherwise.** |
| **F6** | **Confident hallucination** — non-existent SKU, operator, price, or inclusion | The Cursor pattern. Fluent, specific, wrong. Scales at machine speed and is invisible until a traveller acts on it. | Depends on what was fabricated | **Mandatory** — mitigate by construction (retrieval-grounded, closed-set), not by review |
| **F7** | **Entitlement / refund misstatement** | AI states the Pass covers something it does not. SatuSatu either absorbs the cost or has a disputing customer mid-trip. | Financially yes, reputationally partly | **Mandatory** — or hard-gate against a machine-readable entitlement list |
| **F8** | **Silent abandonment** | Traveller gives up. Counted as **deflected**, i.e. as a success. Damage surfaces later in reviews and non-repeat. | Not detectable in-flight | Detectable only via re-contact rate + review sentiment, never via the deflection dashboard |

### 4a.3.3 The architectural consequence

F1 and F2 flip between "mandatory HITL" and "safely automated" depending on **one design decision**: whether the outbound message is **generated** or **rendered**.

- **Generated** (model writes the meeting instruction in prose): every message is a fresh opportunity for F1 and F6. HITL mandatory, and the automation saves nothing because a human reads every message anyway.
- **Rendered** (model *selects* a template and fills slots from the confirmed booking record; it never authors the instruction): the failure surface collapses to slot-selection, which is testable, enumerable, and gate-able. HITL becomes optional.

> 🔴 **Recommendation: in-destination messages must be rendered, never generated.** This is the single highest-value constraint in this section, it costs almost no engineering (template library + slot mapping), and it converts the two highest-volume risky tasks into safe ones.

### 4a.3.4 The Klarna lesson, read correctly

Klarna is usually cited as the AI-CS success case. The collected sources contain both halves, and the second half is the one that matters:

- **Feb 2024** [VERIFIED, Tier A — Klarna press release]: two-thirds of chats, ~700 FTE-equivalent, 11 min → under 2 min, 25% fewer repeat inquiries, $40M profit improvement.
- **May 2025** [VERIFIED, Tier B — Forbes reporting Bloomberg]: CEO Sebastian Siemiatkowski says the company cut too deep and reopens hiring for premium support roles. His own framing: *"an overemphasis on cost — not AI itself — led to lower quality."*
- **Klarna's rebuttal** [VERIFIED, Tier B, quoted in full in the raw]: *"Klarna never eliminated human support. We still work with several thousand outsourced agents. The current pilot involves just two new agents… it's an addition, not a rehire or reversal."* The AI assistant now performs the work of **800+** full-time roles, not 700.

**Read together, the rebuttal reframes the original number.** The 700-FTE figure was never a replacement count — it was *incremental capacity layered on top of a retained human floor of several thousand outsourced agents.* Klarna's actual architecture is, and always was, AI-plus-humans. The public story of replacement was a misreading that Klarna itself had to correct, at reputational cost.

Corroborating detail from the same period: a well-known engineer testing the assistant at launch called it *"underwhelming. It recites exact docs and passes me on to human support fast"* [VERIFIED] Tier B — Business Insider. Fast, correct escalation *was* the design. That is the transferable lesson.

**For SatuSatu, with one concierge queue and an unknown headcount, there is no "several thousand outsourced agents" floor to fall back on.** Klarna could absorb a quality dip because it had depth. SatuSatu cannot.

### 4a.3.5 Kill criterion

> **Design principle: never gate on deflection.** Per §4a.2, deflection rewards abandonment, and abandonment in this product means a stranded traveller. The gate must be built on **error rate** and **re-contact**, with deflection as a reporting metric only.

**Primary kill metric — In-Destination Instruction Error Rate (IDIER)**

```
IDIER = (AI-issued in-destination instructions that resulted in a traveller at the wrong
         place/time/date, OR a booking with no supplier record)
      ÷ (total AI-issued in-destination instructions)
```

Measured by **100% manual QA audit of a random sample**, not by customer complaint volume (F8 means complaints undercount).

| Gate | Threshold | Action |
|---|---|---|
| **Kill** | IDIER **> 0.5%** at n ≥ 400 audited instructions | Terminate in-destination automation. Retreat to pre-trip-only scope. |
| **Hold** | IDIER 0.2–0.5% | No scope expansion. Remediate and re-audit before any further rollout. |
| **Proceed** | IDIER < 0.2% at n ≥ 400 | Expand scope one task at a time, re-gating at each step. |

**Justification of 0.5% as a formula, not a round number.** The threshold should be set where expected incident cost exceeds labour saved:

```
kill when:   IDIER × cost_per_stranding_incident  >  minutes_saved_per_pass × cost_per_minute
```

With base-case inputs from §4a.5 (`cost_per_minute` = $0.079) and an optimistic 40 minutes saved per pass, the labour saved is **$3.16/pass**. `cost_per_stranding_incident` is **unverified — not found in collected sources**, but it is bounded below by the Pass refund ($59.95–$144.95) and is realistically several multiples of that once re-accommodation, goodwill and a public review are included. At a conservative $600/incident, break-even IDIER is **0.53%**. 0.5% is that number, rounded down. **Re-derive it the moment a real stranding cost and a real minutes-saved figure exist.**

**Secondary guardrails — all four measured together, any one breaching triggers the Hold gate**

| Guardrail | Threshold | Anchor |
|---|---|---|
| 72h re-contact rate on AI-handled conversations | ≤ 1.5× the human-handled baseline | Industry: 11.3% AI vs 8.7% human = 1.30× [Tier C] |
| **Confirmed** (not assumed) resolution rate | ≥ 30% by day 60 | Intercom's own engineer: 30–50% is a good start [Tier B]. **Never accept an "assumed resolved" figure.** |
| Escalation rate | **≥ 20% — a floor, not a ceiling** | Inverted deliberately. Benchmark is 16–30% with top quartile <15% [Tier C], but for this product **an escalation rate that is too low is the alarm**: it means the model is not handing off in-destination risk. |
| Pass-buyer CSAT vs. pre-automation baseline | No decline beyond noise | Klarna's failure was invisible in average CSAT until it was a public reversal |

**Dates.** Expressed relative to pilot start T₀, with a worked example against SatuSatu's "now ≤ 90 days" constraint:

| Milestone | Relative | Worked example (T₀ = 2026-08-15) |
|---|---|---|
| Pilot start, pre-trip scope only, 100% QA | T₀ | 2026-08-15 |
| Gate 1 — IDIER + all four guardrails | T₀ + 60d | 2026-10-14 |
| Gate 2 — hard kill/scale decision | T₀ + 90d | 2026-11-13 |

**No in-destination automation ships before Gate 1 clears.** Pre-trip failures are caught before the traveller acts; in-destination failures are not.

### 4a.3.6 ⚠️ Liability asymmetry — flagged, not opined on

SatuSatu's T&C disclaims **operator** performance. That posture fits a reseller/marketplace: the operator ran the tour badly, not us.

**A concierge instruction is not an operator act. It is SatuSatu's own act** — authored by SatuSatu, sent from SatuSatu's own WhatsApp number, delivered as a paid Pass entitlement. Two directions in which automating it may **expand** exposure rather than contain it:

1. **The operator disclaimer does not reach it.** The thing that went wrong was written by SatuSatu, not performed by a third party. The existing disclaimer is aimed at a different actor.
2. **The marketing describes a person.** *"A real Bali local who plans your days"* is a specific representation about who the traveller is dealing with. Substituting a model may create a gap between what was sold and what was delivered — distinct from, and additional to, any question of whether the instruction was correct.

A third, structural point that the Elite Havens comparable makes sharp (§4a.6.2): **Elite Havens can make a concierge promise because it is the exclusive owner representative and controls the supply.** SatuSatu is a reseller and does not control Pool A supply. A concierge promise made over supply you do not control is a structurally different commitment.

> **This is a flag, not a legal opinion, and no legal conclusion should be drawn from it.** Route to counsel **before** any in-destination automation ships — not after the pilot, because the pilot itself issues real instructions to real travellers. Note also that this and the PII question in §4a.5.4 are the same governance question in different clothing, and should go to counsel together.

---

## 4a.4 The three-way split the strategy forces

**Assumption stated explicitly:** D0 = SatuSatu's own premium direct product (the Pass). D1 = SatuSatu self-serve. D2/D3 = partner / platform / white-label tiers. This reading is inferred from the brief ("the D2/D3 platform build consumes the team"; "a human concierge per booking cannot scale to D1/D2/D3"). If the taxonomy differs, the row logic holds but the column labels move.

**The forcing constraint:** a human concierge per booking is an O(n) cost against an O(1) platform revenue model. At D0's price point ($59.95–$144.95 with a bundled eSIM) it is already tight (§4a.5.5). At D1 self-serve prices it is impossible. At D2/D3, SatuSatu would be underwriting an unbounded human liability on a partner's traffic, over supply the partner controls, with no visibility into the traveller relationship.

| # | Concierge job | Survives as premium **D0** | Automated into **D1** self-serve | **Not offered** to D2/D3 | Reason |
|---|---|---|---|---|---|
| 1 | Pre-trip intake | ✅ human-led, AI-drafted | ✅ structured form | — | Form-fillable. The human version is a differentiator, not a necessity; the self-serve version loses little. |
| 2 | Itinerary build | ✅ **this is the product** | ✅ template itineraries + rules-based sequencing | — | D0 sells judgement and local taste. D1 can ship pre-built, pre-validated day templates — the safe 80% with none of the sequencing risk, because a template's transport legs were validated once by a human. |
| 3 | Supplier booking — Pool A | ✅ automated | ✅ automated | ✅ available | API-backed and deterministic. No reason to withhold from anyone; this is plumbing, not service. |
| 4 | Supplier booking — Pool B | ✅ human | ⚠️ **only if Pool B gets a confirmation state machine** | ❌ | Pool B has no API. Exposing manual-confirmation supply to partner traffic means SatuSatu absorbs F2 at volumes it cannot staff, on the ~27% margin pool it most needs to protect. |
| 5 | Day-of coordination | ✅ human-monitored, template-rendered | ✅ template-rendered, unmonitored | ⚠️ rendered messages only, no live queue | Rendered templates are safe to scale (§4a.3.3). A **live human queue** is not — that is the unbounded liability. Partners may receive the messages; they may not receive the queue. |
| 6 | Transport sequencing | ✅ human sign-off | ✅ **only inside pre-validated templates** | ❌ | Free-form sequencing is F3, the highest per-error blast radius. Inside a template, the sequencing was validated once and reused. Outside one, it is a fresh risk on every itinerary. |
| 7 | Exception handling | ✅ **the entire reason the Pass has a premium** | ❌ escalate to D0 queue or refund | ❌ | Irreducibly human, unbounded in the tail, and the thing being paid for. Cannot be offered at self-serve price. Cannot be offered to partners because SatuSatu cannot staff a 24/7 human exception desk against third-party volume it does not forecast. |
| 8 | Upsell | ✅ human + AI | ✅ automated | ✅ available | Revenue-generating, low blast radius, entitlement-gated. Give it to everyone. |
| 9 | Post-trip follow-up | ✅ automated | ✅ automated | ✅ available | Negligible risk, near-zero engineering. |

### 4a.4.1 What the split actually means

**The D0 premium reduces to two jobs: #2 (itinerary build) and #7 (exception handling).** Everything else either automates cleanly or is plumbing. That is a much narrower and much more defensible product than "a dedicated human concierge," and it is honest about where the value is.

**D1's viable offer is the pre-validated template itinerary.** It captures most of #2's value at none of #3/#6's risk, because the risky sequencing decision was made once by a human and then reused. It requires a template library, not a model.

**D2/D3 receive rendered messages and API-backed booking — never a human queue.** The queue is the unbounded cost and the unbounded liability, and it is the one thing that cannot be priced into a platform take rate.

> 🔴 **Marketing consequence, unresolved:** the Pass is currently sold on the concierge, and the concierge is a Pass entitlement while base-catalog buyers get only "Dedicated support." Narrowing D0 to #2 and #7 does not change what the product is worth — but it does change what the copy can honestly claim. That is a positioning decision, not an ops one, and it should be made deliberately rather than discovered by a customer.

---

## 4a.5 Cost model inputs

### 4a.5.1 Statutory wage floor — Bali 2026 [VERIFIED, Tier A]

Bali Governor's Decree **No. 1021/03-M/HK/2025** (UMK/UMSK) and **No. 1011/03-M/HK/2025** (UMP), signed 2025-12-23, effective 2026-01-01. Corroborated across ANTARA, Kompas.TV, NusaBali, GoodStats and the Badung regency's own channel — the strongest-sourced data in this entire step.

| Jurisdiction | 2026 monthly | 2025 monthly | Change |
|---|---|---|---|
| **Kab. Badung** (Kuta, Seminyak, Canggu, Nusa Dua) | **Rp 3,791,002.57** | Rp 3,534,338.88 | **+7.26%** |
| **Kab. Badung — UMSK**, accommodation & F&B, 4/5-star (KBLI 2020 Huruf I) | **Rp 3,828,912.60** | Rp 3,569,682.27 | +7.26% |
| **Kota Denpasar** | Rp 3,499,878.78 | Rp 3,298,116.50 | +6.12% |
| Kab. Gianyar (Ubud) | Rp 3,316,798.48 | Rp 3,119,080.00 | +6.34% |
| Kab. Tabanan | Rp 3,287,678.87 | Rp 3,102,520.45 | +5.97% |
| **UMP Bali** (floor for Klungkung, Karangasem, Bangli, Buleleng, Jembrana) | **Rp 3,207,459** | Rp 2,996,561 | **+7.04%** |

Methodology per PP No. 49/2025: economic growth + inflation + an *alfa* coefficient of 0.5–0.9; Badung's wage council voted 18–0–1 for alfa 0.8 [VERIFIED] Tier B — NusaBali.

> ⚠️ **UMK is minimum wage. It is a legal floor on base salary only.** It is not a market rate, not a competent-bilingual-concierge rate, and above all **not a fully-loaded cost.** The gross-up is done explicitly below. **Wage escalation is a live planning input: +6–7.3% per year, two years running.**

### 4a.5.2 Market base salary — Indonesia CS/concierge [Tier C throughout]

| Source | Figure | Notes |
|---|---|---|
| KantorKu — hotel industry CS | Rp 4.0–7.0m/mo | Guest Service Agent / Front Office Support. **Closest role match.** |
| MixWork — CS Representative, mid-level | Rp 5.5–9.0m/mo (~$355–580 @ IDR 15,500) | Jakarta/Bandung/Surabaya |
| worldsalaries — Customer Support Agent | Rp 59.0m/yr ≈ Rp 4.92m/mo | national |
| Stealth Agents — Indonesia CS agent (voice) | $350–550/mo | vs Philippines $500–750, India $350–600 |
| Indeed — CS, Denpasar | Rp 3,236,222/mo (n=19, updated 2026-07-05) | **Below Denpasar's UMK of Rp 3,499,879 — internally inconsistent.** Use as a floor sanity check only. |
| Glassdoor — CS Rep, Bali | "IDR 7,450,000 per year… 16098% higher than the national" | **Incoherent. Discarded.** |

**Selected base band for an English-fluent, tourism-sector Bali concierge** [INFERENCE] — above generic CS (destination expertise + English + traveller-facing), anchored on KantorKu's hotel band and MixWork's mid-level band:

| | low | base | high |
|---|---|---|---|
| Monthly base salary | **Rp 4.5m** | **Rp 6.5m** | **Rp 9.5m** |

### 4a.5.3 The gross-up — UMK → fully-loaded, multiplier shown

> ⚠️ **This is the step that is most often skipped and most often wrong.** Each factor is itemised so it can be corrected independently when real payroll data lands.

| Factor | low | base | high | Basis |
|---|---|---|---|---|
| **A.** THR (13th-month religious holiday allowance) | ×1.0833 | ×1.0833 | ×1.0833 | [INFERENCE] Standard Indonesian statutory practice (1 month per year = 8.33%). **Not evidenced in the collected sources — confirm against payroll.** |
| **B.** Employer statutory contributions (BPJS Ketenagakerjaan + BPJS Kesehatan) | ×1.11 | ×1.11 | ×1.11 | [VERIFIED] Tier C — Plane models Indonesian employment cost as base + **11.00%** "taxes". A blended figure; the component BPJS rates are **unverified — not found in collected sources**. |
| **A×B — statutory subtotal** | **×1.20** | **×1.20** | **×1.20** | |
| **C.** Non-wage operating load — supervision/QA, tooling & licences, workspace, device, recruitment, attrition replacement | ×1.25 | ×1.40 | ×1.60 | [INFERENCE] from: BPO far-offshore management overhead **15–25%** [Tier C — callforce.global]; the US in-house line-item build implies a salary→total ratio of **~2.0–2.3×** but is health-insurance-dominated and **does not transfer** [Tier C — globalify]; Indonesian BPO attrition **25–35%** [Tier C — Stealth Agents] makes recruitment and ramp a recurring, not one-off, cost. |
| **Total multiplier on base salary** | **×1.50** | **×1.68** | **×1.92** | |

**Resulting fully-loaded cost per concierge FTE** (FX: **IDR 15,500 / USD 1** [VERIFIED] Tier C — MixWork, stated as approximate for 2026; no better rate exists in the collected sources — **all USD figures below are FX-sensitive and should be re-run at the live rate**):

| | low | base | high |
|---|---|---|---|
| Fully-loaded monthly | Rp 6.75m / **$435** | Rp 10.92m / **$705** | Rp 18.24m / **$1,177** |
| **Fully-loaded annual** | **$5,226** | **$8,455** | **$14,121** |

**Triangulation** — two independent Tier C sources bracket the base case:
- Stealth Agents: a 10-person Jakarta tier-1 CS team costs **$60,000–90,000/yr fully loaded** = **$6,000–9,000 per FTE**. Base case $8,455 sits inside it. ✅
- Plane: median remote Indonesian CS rep, total employment cost **$9,989/yr** ($8,999 base + 11% tax). Between base and high. ✅

### 4a.5.4 Cost per concierge-minute — the number everything else hangs off

Productive hours: 40h/week statutory [Tier C — remotepeople] = 2,080 gross. Shrinkage (annual leave, Indonesia's substantial public-holiday calendar, training, breaks, admin) **12–18%** [INFERENCE].

| | low | base | high |
|---|---|---|---|
| Productive hours/yr | 1,830 | 1,780 | 1,706 |
| **Cost per productive hour** | **$2.86** | **$4.75** | **$8.28** |
| **Cost per productive minute** | **$0.048** | **$0.079** | **$0.138** |

### 4a.5.5 🔴 The finding that inverts the business case

Vendor AI support is priced per resolution, in USD, against a US/EU labour baseline:

| Reference | Price | Tier |
|---|---|---|
| Intercom Fin | **$0.99** per resolution (+ seat: $29–139/mo, or $49.50/mo standalone minimum) | C — vendor pricing, cited by two independent parties |
| Quickchat AI Enterprise | **$0.50** per resolution | C — vendor |
| AI-native platforms, general | **$1–3** per resolution | C — lorikeetcx.ai |
| Gartner — self-service | **$1.84** per contact | C reporting Gartner |
| Gartner — agent-assisted | **$13.50** per contact | C reporting Gartner |
| Human-agent baseline cited by vendors | **$6–12** per contact | C — Fin/Intercom 2026 |

**The $6–13.50 human baseline is a US/EU number. SatuSatu's human baseline is $0.079 per minute.**

Break-even average handle time — the point above which a per-resolution AI vendor becomes cheaper than an Indonesian concierge:

| AI price/resolution | vs low labour ($0.048/min) | vs **base** ($0.079/min) | vs high ($0.138/min) |
|---|---|---|---|
| $0.99 (Fin) | 20.6 min | **12.5 min** | 7.2 min |
| $0.50 (Quickchat) | 10.4 min | **6.3 min** | 3.6 min |
| $1.84 (Gartner self-service) | 38.3 min | **23.3 min** | 13.3 min |

> **Conclusion: US-priced per-resolution AI does not clear an Indonesian labour arbitrage on short interactions.** At the base case, a Fin-priced resolution only pays for itself if it replaces **more than ~12.5 minutes** of concierge time. Most deflectable tasks in §4a.1.2 (Pool A booking: 3 min; upsell: 3 min; post-trip: 2 min) are **far** below that line — the AI would cost 3–5× the human it replaces.
>
> **Therefore the business case for AI here cannot be cost per contact.** It must rest on: **24/7 coverage** (a human queue in one time zone cannot serve a 110-country inbound market), **latency** (Canary's LINE SF: median response 10 min → under 1 min; Klarna: 11 min → under 2 min), **language breadth** (Klarna: 35+ languages at zero marginal cost; Elite Havens draws from 110+ countries), and **headcount avoidance at growth** rather than headcount reduction today. Any model that shows AI saving money on cost-per-contact against Indonesian labour has the wrong baseline in it.

### 4a.5.6 Cost per AI conversation — structure given, price input missing

> **Model and token pricing for 2026 is unverified — not found in collected sources.** No file in this step contains provider list prices. The structure below is therefore given with the price left as a variable, to be filled from the provider's current price list. **Do not substitute a remembered price.**

Assumed shape of a multi-turn WhatsApp support interaction [INFERENCE]:

| | low | base | high |
|---|---|---|---|
| Turns (user + assistant pairs) | 6 | 10 | 18 |
| Effective billed **input** tokens/conversation (system prompt + tools + retrieved KB + growing history, assuming prompt caching on the static prefix) | ~25k | ~60k | ~140k |
| **Output** tokens/conversation | ~1.2k | ~2.5k | ~5.0k |

```
cost_per_conversation = (input_tokens × $/Mtok_input  ÷ 1e6)
                      + (output_tokens × $/Mtok_output ÷ 1e6)
                      + whatsapp_platform_conversation_fee
                      + retrieval/vector infrastructure amortised per conversation
```

**Two cost lines in that formula are missing entirely from the collected sources and are real, non-trivial gaps:**
- **WhatsApp Business Platform per-conversation fees.** Nine files, several of them entirely about WhatsApp Business API, and **not one price.** Indonesia is a distinct pricing market. This must be sourced before any model is built.
- **Retrieval/vector infrastructure.** Not addressed anywhere.

**Observable market proxy, in place of a computed figure:** vendor-priced, all-in, **$0.50–$0.99 per resolution** [Tier C], or **$1–3** for AI-native platforms [Tier C]. Cross-check any computed token cost against this band — a computed figure far below it is missing infrastructure, margin, or the platform fee.

### 4a.5.7 Concurrency — unverified

**No collected source reports concurrent-chats-per-agent, with or without AI assist.** Unverified — not found in collected sources. This is a genuine gap and it matters, because concurrency is the mechanism by which assist (not deflection) produces savings, and §4a.1.3 established that assist is 80% of the opportunity.

The only adjacent proxies available:

| Proxy | Figure | Tier | What it actually measures |
|---|---|---|---|
| AI-augmented agent throughput | **+13.8%** inquiries per hour | C — G2 via eesel.ai | Throughput, not concurrency. A modest number, and notably the *only* assist-side figure in the whole dump. |
| Conversation summarisation on escalation | **−35–45%** human handle time | C reporting Gartner 2025 | Handle time on escalated tickets only. |
| Response latency, hotel messaging | 10 min → **under 1 min** median | C — Canary/LINE SF | Latency, not capacity. |
| Response latency, fintech | 11 min → **under 2 min** | **A** — Klarna PR | Resolution time, not capacity. |
| Peak absorption | absorbed a **40%** interaction increase without added staff | C — Hubtype/easyJet | The closest thing to a concurrency claim in the set, and it is a vendor case study. Relevant to Bali's **1.55× peak-to-trough** seasonality. |

> **Planning note:** the +13.8% throughput figure is the honest one to plan against, and it is far below what any deflection headline implies. If assist delivers +13.8% and deflection realistically delivers 30–50% of a *small* deflectable fringe (8% of concierge minutes), the combined labour saving is modest. **Model it that way and be pleasantly surprised, rather than the reverse.**

### 4a.5.8 ⚠️ The PII constraint — unresolved, and it flips build/buy

**Whether traveller WhatsApp content may be sent to a third-party model provider is not answered by any collected source.** The only trace is a truncated analyst prompt — Mordor Intelligence flags "How does the Personal Data Protection Law affect providers?" as a material question for Indonesian BPO — with **no substantive treatment.** So: Indonesia's PDP Law is flagged as material by industry analysts; its actual application here is **unverified — not found in collected sources.**

Because it flips the answer, both paths are costed:

**Path A — traveller content MAY go to a third-party model provider**

- Vendor SaaS is available: **$0.50–0.99 per resolution**, configuration-only, no engineering. Fits the "now ≤90 days, near-zero engineering" constraint exactly.
- **Buy is the default and the decision is easy.**
- But per §4a.5.5, the cost case is negative at short AHT. Buy for **coverage, latency and language**, and say so in the business case rather than dressing it as cost reduction.
- Residual risk: vendor lock-in on a per-resolution meter whose definition of "resolution" the vendor controls (§4a.2.1) — negotiate the definition into the contract, or the meter and the value diverge.

**Path B — traveller content MAY NOT leave SatuSatu's control / Indonesian jurisdiction**

The vendor route closes at the egress point. Four options, three of which fail the constraint:

| Option | Engineering burden | Verdict |
|---|---|---|
| In-region hosted inference with Indonesian/regional data residency | Low-to-moderate *if* it exists at acceptable price | **Availability and pricing unverified — not found in collected sources.** This is the highest-value unknown to resolve; it is the only branch that preserves the ≤90-day path. |
| Self-hosted open-weight model | High — inference infra, evals, ops, on-call | ❌ **Fails.** Requires the engineering capacity that explicitly does not exist. |
| Restrict AI to non-PII surfaces only (public catalog Q&A, no booking context, no traveller identity) | Low | ⚠️ **Technically viable, commercially thin.** Deflects only the most generic tier — exactly the 8% fringe from §4a.1.3, minus anything requiring booking context. Near-zero value. |
| Redaction/tokenisation proxy before egress | Moderate | ⚠️ Buys the vendor route back, but **redaction is itself a failure surface** (a missed identifier is a breach, an over-redaction is a wrong answer) and it consumes engineering that does not exist. |

> 🔴 **The PII answer determines whether this is a configuration purchase or an engineering project — and only one of those fits the stated constraints.** Resolve it **first**, before vendor selection, before pilot design, before any scoping. It is the highest-leverage open question in this section and it costs one legal opinion to close. Route it to counsel together with the liability question in §4a.3.6; they are the same governance question.

---

## 4a.6 Two comparables

### 4a.6.1 Canary Technologies — hotel AI concierge

**Corporate.** 20,000+ hotels, 125+ countries; customers include Marriott, Four Seasons, Wyndham, IHG, Choice, BWH [VERIFIED] Tier C — Canary's own pages. The brief's **$80M Series D of 2025-06-12** is **not corroborated — no funding data of any kind appears in `s4a-canary.md`.**

**What it automates** [all VERIFIED, all Tier C, all Canary's own marketing unless noted]:

| Product | Claim |
|---|---|
| AI Guest Messaging | "Automate more than 80% of guest communication"; 100+ languages; SMS + WhatsApp + unified inbox; PMS integrations |
| AI Voice | Inbound calls, reservations, modifications, FAQs, upsells, intelligent routing, 24/7. Framing: *"like having the ultimate concierge."* Positioned against the claim that hotels miss **up to 40%** of calls |
| AI Webchat | Website virtual guest-services agent, direct-booking conversion |
| AI Agent Studio / Agentic Sales Coordinator | Group and event sales coordination |

**Published outcome metrics:**

| Property | Metric | Tier |
|---|---|---|
| Linchris | **82%** of guest messages managed automatically | C |
| Holiday Inn Express | **82%** of guest inquiries automated; **$1,700** incremental monthly revenue | C |
| The LINE SF (236 rooms) | Median response **10 min → under 1 min**; **65%** of early check-in revenue via AI upsells | C |
| IHG | **4×** increase in upsell conversions | C |
| Dream Hotels | **5%** improvement in guest service scores | C |
| HotelTechReport user ratings | Messages 4.9 (606 reviews); Webchat 4.7 (189); Voice 4.7 (32) | C |

⚠️ **Two internal inconsistencies worth recording:**
- The LINE SF upsell conversion multiple is stated as **4×** on Canary's site, **"four times better"** in the HotelTechReport transcript, and **"14 times higher than traditional EPS links"** in the same video. A 3.5× discrepancy inside one case study.
- Canary's own pages say **80–82%** automation; the third-party HotelTechReport review video says **"up to 70% of inbound questions."** A 10–12 point vendor-vs-reviewer gap, consistent with the systematic pattern in §4a.2.2.

**What Canary explicitly leaves to humans:** ⚠️ **Unverified — not found in collected sources. Canary publishes no exclusion list anywhere in the collected material.** The only signal is directional framing: *"freeing staff to focus on high-touch service"* and *"steps in when the front desk is busy."* **The absence is itself the finding** — a mature, 20,000-hotel vendor with nine industry awards does not publish a boundary of competence. Any procurement conversation must force that boundary into writing, because it will not arrive in the marketing.

**Transferability to SatuSatu — limited, and in a specific way.** Canary's automated surface is **in-property, low-stakes, high-frequency**: spa hours, restaurant recommendations, loyalty policy, early check-in, towels. A wrong answer means a mildly annoyed guest who is already inside a building with staff in it. **None of Canary's published automation covers in-destination instruction issued to a guest who has left the property.** The failure classes in §4a.3 — F1, F3, F5 — have no analogue in Canary's evidence base. Canary demonstrates that hotel-adjacent FAQ automates well. It demonstrates nothing about whether a day-plan automates safely.

### 4a.6.2 Elite Havens — the real Bali analogue, running 28 years

**Corporate.** Founded 1998 [Tier C]. Acquired by **Dusit Thani PCL** via LVM Holdings Pte Ltd in **2018**, deal value **~$15m** [Tier B — Breaking Travel News]; at acquisition, "more than 200 fully staffed properties" [Tier A — Dusit newsroom]. The brief's **September 2018** month is not corroborated in these files.

| Metric | Value | Year | Tier |
|---|---|---|---|
| Villas under management | 243 | end-2023 | **A** — Dusit Annual Report 2023 |
| Villas | "more than 300" / "almost 300" / "290+" | 2026 / n.d. | C — LinkedIn, own sites, Flywire |
| Guests per year | **~70,000** | 2018 | **A** — Dusit AR2018 |
| Guests per year | **~80,000** | 2023 | **A** — Dusit AR2023 |
| Guest source countries | 110+ | 2018 & 2023 | **A** |
| In-villa events/year | **400+** | 2018 | **A** |
| Company headcount | 1,001–5,000 (184 LinkedIn-associated members) | 2026 | C — LinkedIn self-report |
| Staff per villa | *"each of our villas has, let's say 12 staff"* | 2020 | **B** — CEO Jon Stonham, Rental Scale-Up interview |
| Destinations | Bali, Lombok, Nusa Lembongan, Phuket, Koh Samui, Maldives, Japan, India | 2026 | C |

**How the concierge function is structured** — this is the part that transfers:

1. **The Elite Concierge is a separate, centralised layer, distinct from in-villa staff.** The in-villa team (Villa Manager, senior butler, chef, housekeepers, gardeners, drivers, villa attendants) runs the property. The **Elite Concierge** books restaurants, cooking classes, in-villa massage, personal trainers, pre-stocked groceries, and performers [VERIFIED] Tier C — elitehavens.com.
2. **Its job description matches SatuSatu's concierge almost exactly.** Dusit AR2018 [Tier A]: *"Villa managers and concierges use their local knowledge to source and plan such experiences on-demand… spa therapies to Balinese kite-flying lessons; yacht charters to visiting local artisans; personalized private ski-guides to last-minute reservations at top local restaurants."* This is the closest published description of the SatuSatu concierge job in any collected source.
3. **Continuity is the explicit product promise:** *"Our guests are managed by the same professional team from the initial booking, and throughout their stay"* [Tier C]. Note this is a claim about **one team across the whole journey** — which is precisely what SatuSatu's two-queue architecture (general + Pass concierge) does not deliver.
4. **The concierge is organised by source market, not by destination.** Contact lines are AU +61, ID +62 361 737 498, TH +66, SG +65, plus an "other countries" line [Tier C]. For a 110-country guest base, the routing dimension chosen was **language/time zone**, not **destination**. That is a directly applicable design signal for a business serving inbound international travellers from one Indonesian time zone.

**Staffing and cost:** ⚠️ **Elite Havens' concierge headcount and concierge cost are unverified — not found in collected sources.** Dusit reports Elite Havens only inside a combined "Hotel Management" line (THB 788m in 2023, +73.2% YoY, driven by Kyoto openings, Middle East/Guam properties **and** Elite Havens together) [Tier A]. There is no segment disclosure and no cost line.

**What can nonetheless be inferred, and it is the important part:**

- **The labour is in the property, not in the concierge.** ~250–300 villas × ~12 staff ≈ **3,000–3,600 in-villa staff** [INFERENCE from Tier A villa count × Tier B CEO quote], which reconciles with LinkedIn's 1,001–5,000 band. The concierge layer is thin enough to be invisible in a headcount that size. **After 28 years, the closest real-world operator of "human concierge, curated by locals, in Bali" does not put its people in the concierge.** It puts them in the delivery of the experience and runs coordination as a light central layer.
- **Elite Havens is not a cost analogue and must not be used as one.** Twelve staff per villa is supportable at luxury-villa ADR. SatuSatu's Pass is **$59.95–$144.95** including a bundled eSIM. Elite Havens is a **service-shape** analogue — what the job consists of, how it is organised, what is centralised versus local — and nothing more.
- 🔴 **The structural difference that matters most: Elite Havens is principal, SatuSatu is not.** *"Elite Havens is the only owner representative, exclusively marketing and managing all properties. All other agents must go via us for each elite haven"* [Tier C — villamana.com]. Elite Havens can promise a concierge experience because **it controls the supply, the staff, and the standard.** SatuSatu is a reseller whose T&C disclaims operator performance, and whose Pool A supply arrives via GlobalTix. **A concierge promise made over supply you do not control is a different and larger commitment than the one Elite Havens makes** — and automating that promise (§4a.3.6) compounds the difference rather than reducing it.
- **Scale reference for capacity planning:** Elite Havens serves ~80,000 guests/year with a concierge layer small enough to be undisclosed. The relevant question for SatuSatu is not "how many concierges per pass" but **"what did Elite Havens automate, standardise or simply decline to offer, in order to keep that layer thin over 28 years?"** The collected sources do not answer it. It is the single best follow-up interview in this research.

---

## 4a.7 Benchmark summary table — everything Step 5 needs

**Every row carries low/base/high and a tier.** `[V]` = VERIFIED, `[I]` = INFERENCE.

### Labour cost

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| L1 | UMK Kab. Badung 2026, monthly | — | **Rp 3,791,003** | — | [V] | **A** |
| L2 | UMSK Badung 2026, accommodation/F&B 4–5★ | — | **Rp 3,828,913** | — | [V] | **A** |
| L3 | UMK Kota Denpasar 2026 | — | **Rp 3,499,879** | — | [V] | **A** |
| L4 | UMP Bali 2026 | — | **Rp 3,207,459** | — | [V] | **A** |
| L5 | Annual Bali minimum-wage escalation | +5.97% | **+7.04–7.26%** | — | [V] | **A** |
| L6 | Market base salary, Bali bilingual concierge, monthly | Rp 4.5m | **Rp 6.5m** | Rp 9.5m | [I] | C |
| L7 | Statutory gross-up (THR × BPJS) | ×1.20 | **×1.20** | ×1.20 | [I]/[V] | C |
| L8 | Non-wage operating load | ×1.25 | **×1.40** | ×1.60 | [I] | C |
| L9 | **Total gross-up multiplier on base** | **×1.50** | **×1.68** | **×1.92** | [I] | C |
| L10 | **Fully-loaded concierge FTE, annual USD** | **$5,226** | **$8,455** | **$14,121** | [I] | C |
| L11 | Fully-loaded FTE, monthly USD | $435 | **$705** | $1,177 | [I] | C |
| L12 | Productive hours/yr after shrinkage | 1,830 | **1,780** | 1,706 | [I] | C |
| L13 | **Cost per productive concierge-hour** | **$2.86** | **$4.75** | **$8.28** | [I] | C |
| L14 | **Cost per productive concierge-minute** | **$0.048** | **$0.079** | **$0.138** | [I] | C |
| L15 | Triangulation — Jakarta tier-1 CS FTE, fully loaded | $6,000 | — | $9,000 | [V] | C |
| L16 | Triangulation — Plane total employment cost, remote CS | $7,530 | **$9,989** | $12,209 | [V] | C |
| L17 | Indonesia BPO attrition | 25% | — | 35% | [V] | C |
| L18 | FX assumption | — | **IDR 15,500/USD** | — | [V] | C |

### Concierge workload

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| W1 | **Total concierge minutes per pass** | **32** | **100** | **276** | [I] | — |
| W2 | — of which deflectable | 3 | **8** | 21 | [I] | — |
| W3 | — of which assistable | 29 | **80** | 195 | [I] | — |
| W4 | — of which irreducibly human | 0 | **12** | 60 | [I] | — |
| W5 | Deflectable share of concierge time | 9% | **8%** | 8% | [I] | — |
| W6 | **Concierge labour cost per pass** (W1 × L14) | **$1.54** | **$7.92** | **$38.09** | [I] | — |

### AI performance

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| A1 | Tier-1 deflection, enterprise distribution | 22.4% | **41.2%** | 58.7% | [V] | C — compilation, definitions vary |
| A2 | **Deployment floor per vendor's own engineer** | 30% | — | 50% | [V] | **B** |
| A3 | New-deployment launch band | 40% | **45%** | 50% | [V] | C |
| A4 | Mature deployment, good KB | 60% | **63%** | 67% | [V] | C |
| A5 | Vendor top-decile (Fin top-10) automation | — | **78%** | — | [V] | C |
| A6 | Independent 60-day test, 4 SMB clients, 500 tickets/mo | — | **38%** | — | [V] | C (only entry with disclosed n) |
| A7 | **Vendor claim minus measured, systematic gap** | 16 pts | **25 pts** | 40 pts | [V] | C |
| A8 | 72h re-contact, AI-resolved | — | **11.3%** | — | [V] | C |
| A9 | 72h re-contact, human-resolved | — | **8.7%** | — | [V] | C |
| A10 | Escalation rate at benchmark | <15% | **16–30%** | >45% | [V] | C |
| A11 | Answer accuracy at benchmark | <80% | **87–92%** | >92% | [V] | C |
| A12 | Hallucination rate | >5% | **~2%** | <1% | [V] | C |
| A13 | AI CSAT at benchmark | <70% | **79–86%** | >88% | [V] | C |
| A14 | AI-augmented agent throughput lift | — | **+13.8%** | — | [V] | C |
| A15 | Handle-time cut from escalation summarisation | 35% | **40%** | 45% | [V] | C |
| A16 | Travel/WhatsApp — easyJet, cases automated (webchat+WA) | — | **62%** | — | [V] | C |
| A17 | Travel — easyJet, total call-volume reduction | — | **28%** | — | [V] | C |
| A18 | Insurance — Allianz, claims end-to-end automated | — | **42%** | — | [V] | C |
| A19 | Fintech — Klarna, share of all CS chats | — | **~66%** | — | [V] | **A** |
| A20 | Klarna, resolution time before → after | — | **11 min → <2 min** | — | [V] | **A** |
| A21 | Klarna, repeat-inquiry reduction | — | **25%** | — | [V] | **A** |
| A22 | Hotel — Canary, guest messages automated (vendor) | — | **80–82%** | — | [V] | C |
| A23 | Hotel — Canary, same metric per third-party reviewer | — | **~70%** | — | [V] | C |
| A24 | Hotel — Canary/LINE SF, median response time | — | **10 min → <1 min** | — | [V] | C |
| A25 | Hotel — Canary/LINE SF, upsell conversion multiple | 4× | **4×** | 14× | [V] | C — internally inconsistent |

### AI and channel cost

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| C1 | Vendor AI, cost per resolution | $0.50 | **$0.99** | $3.00 | [V] | C |
| C2 | Intercom seat requirement | $29/mo | — | $139/mo (or $49.50/mo standalone min) | [V] | C |
| C3 | Gartner — self-service per contact | — | **$1.84** | — | [V] | C |
| C4 | Gartner — agent-assisted per contact (US/EU) | — | **$13.50** | — | [V] | C |
| C5 | **Break-even AHT vs Fin $0.99** | 7.2 min | **12.5 min** | 20.6 min | [I] | — |
| C6 | **Break-even AHT vs Quickchat $0.50** | 3.6 min | **6.3 min** | 10.4 min | [I] | — |
| C7 | Assumed turns per WhatsApp support conversation | 6 | **10** | 18 | [I] | — |
| C8 | Assumed billed input tokens/conversation | 25k | **60k** | 140k | [I] | — |
| C9 | Assumed output tokens/conversation | 1.2k | **2.5k** | 5.0k | [I] | — |
| C10 | Model $/Mtok, 2026 | — | **NOT FOUND** | — | unverified | — |
| C11 | WhatsApp Business Platform conversation fee, Indonesia | — | **NOT FOUND** | — | unverified | — |
| C12 | Retrieval/vector infra per conversation | — | **NOT FOUND** | — | unverified | — |
| C13 | Concurrent chats per agent, with/without assist | — | **NOT FOUND** | — | unverified | — |

### Comparables

| # | Metric | Value | Label | Tier |
|---|---|---|---|---|
| K1 | Elite Havens, guests/year | 70,000 (2018) → **80,000** (2023) | [V] | **A** |
| K2 | Elite Havens, villas under management | **243** (end-2023); "300+" (2026 self-report) | [V] | A / C |
| K3 | Elite Havens, staff per villa | **~12** | [V] | **B** — CEO quote |
| K4 | Elite Havens, implied in-villa staff total | 3,000 / **3,300** / 3,600 (low/base/high) | [I] | — |
| K5 | Elite Havens, company headcount band | 1,001–5,000 | [V] | C |
| K6 | Elite Havens, guest source countries | **110+** | [V] | **A** |
| K7 | Elite Havens, in-villa events/year | **400+** | [V] | **A** |
| K8 | Elite Havens, years operating | **28** (est. 1998) | [V] | C |
| K9 | Dusit acquisition value | **~$15m** (2018) | [V] | **B** |
| K10 | Elite Havens concierge headcount / cost | **NOT DISCLOSED** | unverified | — |
| K11 | Canary, hotels / countries | **20,000+ / 125+** | [V] | C |
| K12 | Canary, published human-only boundary | **NONE PUBLISHED** | unverified | — |
| K13 | Canary Series D $80M, 2025-06-12 | **not corroborated in these files** | brief-asserted | — |

### Consumer acceptance

| # | Metric | Value | Label | Tier |
|---|---|---|---|---|
| P1 | Prefer companies did **not** use AI for CS (Gartner, n=5,728, Dec 2023) | **64%** | [V] | C reporting B |
| P2 | Believe a human option should always exist (SurveyMonkey 2025) | **89%** | [V] | C |
| P3 | Agree AI chatbot response is helpful (WhatsApp/Kantar 2026) | **67.7%** | [V] | C |
| P4 | Prefer bots when they want immediate service | **51%** | [V] | C |

---

## 4a.8 Open questions this step could not close

Ordered by how much they change the answer.

| # | Question | Why it matters | Where it must come from |
|---|---|---|---|
| **1** | 🔴 **May traveller WhatsApp content go to a third-party model provider under Indonesia's PDP Law?** | Flips buy (config, ≤90 days, feasible) vs build (engineering, not feasible). Nothing else can be decided first. | Counsel. Not a research task. |
| **2** | 🔴 **Does automating a concierge instruction expand SatuSatu's liability beyond its operator disclaimer?** | A concierge instruction is SatuSatu's own act; the existing T&C disclaimer addresses a different actor. Also engages the "a real Bali local" representation. | Counsel — same brief as #1. |
| **3** | Concierge headcount, minutes per pass, current AHT | Every formula in §4a.5 is parameterised on these. Known unknown per the brief. | Internal — ops. One week of queue data closes it. |
| **4** | WhatsApp Business Platform per-conversation pricing, Indonesia | A real cost line, entirely absent from nine files including several dedicated to WhatsApp API. | Meta / BSP price list. |
| **5** | 2026 model token pricing | §4a.5.6 is a formula with a hole in it. | Provider price list. |
| **6** | Concurrency: concurrent chats per agent, with and without AI assist | Assist is 80% of the opportunity (§4a.1.3) and concurrency is how assist converts to savings. No source. | Vendor POC measurement, or internal baseline. |
| **7** | Is in-region / Indonesian-resident model inference available and at what price? | The only Path-B branch that preserves the ≤90-day, near-zero-engineering constraint. | Provider / cloud vendor. |
| **8** | Cost per stranding incident | The kill criterion in §4a.3.5 is derived from an assumed $600. Real number moves the threshold materially. | Internal — prior incidents, refund and goodwill history. |
| **9** | Pass effective gross margin (breakage-adjusted) | §4a.5 expresses concierge cost as % of gross profit; the denominator is assumed, not known. 90-day activation implies material breakage. | Internal — finance. |
| **10** | What did Elite Havens automate, standardise, or decline to offer to keep its concierge layer thin for 28 years? | The single most transferable unknown in this section. Not in the public record. | Practitioner interview. |
| **11** | Indonesian statutory employer contribution rates (BPJS TK / Kesehatan components) | The 11% blended figure is one Tier C source. L7 rests on it. | Internal payroll. |
| **12** | Any measured deflection benchmark from an Indonesian or in-destination travel operation | Every travel number in §4a.2.4 is European airline/insurance customer service, which is a structurally different workload. | Not found. May not exist publicly. |

---

*End of section 4a. Sections 4b and 4c to be appended below.*

# 4b — Catalog Scale-Up & Supply Ops

**Prepared:** 2026-07-26. Same evidence tiers and `[VERIFIED]` / `[INFERENCE]` labelling rules as §4a (see document header).

## 4b.0 What this section rests on, and what it does not

Two things must be said before any number is used.

**First, the strongest evidence in this step is already inside `research/raw/` from prior steps, not from this step's searches.** The connectivity teardowns (`teardown-connectivity-indonesia.md`, `teardown-klook-gyg.md`, `teardown-viator-kkday-headout.md`) contain Tier A material read directly from OpenAPI specs and vendor engineering blogs. That is where the load-bearing findings on feed mechanics, listing production and search come from.

**Second, four of the seven metrics the brief asks for do not exist in public form.** Stating this plainly rather than manufacturing a band and hoping:

| Requested metric | Status |
|---|---|
| Cost per SKU onboarded manually | **No published figure anywhere.** Derived here from §4a's Indonesian labour cost per minute. `[INFERENCE]` only. |
| Cost per SKU ingested via feed | **No published figure.** Derived from eng-week estimates. `[INFERENCE]` only. |
| Cost per SKU maintained per year | **No published figure.** `[INFERENCE]` only. |
| SKUs live per content-ops FTE per week | **No published figure.** `[INFERENCE]`, anchored on one Tier A time-per-listing datapoint (GYG, §4b.2.1). |
| **Auto-pass rate at a quality gate** | 🔴 **No measured figure exists in any source, vendor or otherwise.** Every band in §4b.2.3 is `[INFERENCE]`. Any vendor quoting one is quoting a number they have not defined. |
| Translation cost per SKU per language | Per-word rates are available at Tier C; per-SKU is derived. |
| % of bulk catalog unsellable | 🔴 **Not found.** Searched directly; the PIM/data-quality literature discusses the problem qualitatively and quantifies **impact**, never **incidence**. `[INFERENCE]` only. |

**Also not found, despite direct search:** a no-show rate for tours and activities (§4b.5); any OTA-specific statement of image sublicensing rights for resold inventory (§4b.1.5); any published attraction-matching or POI-deduplication accuracy rate from Klook, GetYourGuide, Viator or Expedia (§4b.1.6). The nearest usable proxy for the last of these is academic product-matching, which is a genuinely different problem — see the caveat there.

⚠️ **On automation-rate claims specifically.** §4a established that vendor deflection claims exceed measured reality by 16–40 points, systematically and in one direction. **The same discount should be assumed for every catalog-automation claim in this section, because the incentive structure is identical and the measurement discipline is worse** — deflection at least has a denominator people argue about; "auto-pass rate" has no agreed definition at all.

---

## 4b.1 Feed ingestion — the second-feed problem

### 4b.1.1 The problem does not exist today and will exist on day one of feed two

With one feed there is nothing to reconcile. GlobalTix is the sole source of truth for its own SKUs; a product ID maps 1:1 to a SatuSatu listing; there is no second opinion about what "Universal Studios Singapore 1-Day Pass" costs or looks like.

Add a second aggregator and **every high-demand attraction arrives twice.** That is not an edge case — it is the *definition* of a high-demand attraction, and the overlap concentrates precisely on the SKUs SatuSatu most wants: USS, Singapore Oceanarium, River Wonders, Waterbom Bali, GWK, Taman Safari. A second aggregator that did **not** overlap on those would have no commercial value.

🔴 **And the first duplicate a customer sees is a quality failure, not a data-engineering ticket.** Two cards for the same attraction, different prices, different photos, different cancellation terms, is read by a traveller as "this site doesn't know what it's selling." On a 253-SKU catalog where a category page shows the whole inventory, duplicates are not buried on page 4. They are on page 1.

### 4b.1.2 The five sub-problems, in dependency order

They must be solved in this order; solving 3 before 1 produces a clean taxonomy over a duplicated catalog.

| # | Sub-problem | What it actually is | Hardest part |
|---|---|---|---|
| **1** | **Product mapping** | Establish a SatuSatu-owned canonical product record, and map each feed's product/option/ticket-type IDs onto it. | Feeds are **not** at the same granularity. GlobalTix has `product` → `options` → `ticketType` (with `ticketType/get` already deprecated) [VERIFIED, Tier A]. GYG has `tours` → `options`. Viator has products with option codes. **The unit that is bookable differs between feeds**, so a 1:1 product map is wrong before you start; the map has to be product-to-product *and* option-to-option, and the second is where it breaks. |
| **2** | **Cross-supplier deduplication** | Decide that feed-A product X and feed-B product Y are the same sellable thing — or are deliberately different (1-day vs 2-day, with/without transfer, express-lane vs standard). | The near-duplicate, not the duplicate. "USS 1-Day Pass" vs "USS 1-Day Pass with Express" are different products; "USS One Day Ticket" vs "Universal Studios Singapore 1 Day Pass" are the same one. Getting the first pair merged is a worse failure than leaving the second pair split. |
| **3** | **Taxonomy normalisation** | Reconcile each feed's categories, cities and attribute vocabularies into one. | Both are *hierarchies with different shapes*, not just different labels. GlobalTix exposes `getAllCountries` / `getAllCities` / `getAllCategories` and gives **Ubud its own `cityId` (112) separate from Bali (`cityId` 2)** [VERIFIED, Tier A] — so its geography is city-granular in Bali specifically. GYG exposes `/categories` and `/categories/{id}` plus `location`, `itinerary`, `day_breakdowns` structures [VERIFIED, Tier A]. A merge that flattens Ubud into Bali destroys the one geographic distinction that matters most for a Bali-dominant catalog. |
| **4** | **Pricing & availability reconciliation** | When two feeds sell the same thing at different prices with different availability, decide which one you sell — per request, not per product. | GlobalTix returns **seven price points per ticket type**, an `isDynamicPrice` flag, a `demandType` field (observed `NON_PEAK`), and pushes changes via a `ticket-type-price-update` webhook [VERIFIED, Tier A]. Availability is a **pull** with a hard **≤30-day query window** and a **6-month horizon** [VERIFIED, Tier A]. So a live "cheapest of two feeds" decision requires two live availability calls inside the page render, each rate-limited. **This is a latency and cost problem, not just a logic problem.** |
| **5** | **Image rights across resold inventory** | Do you have the right to display, resize, re-host, CDN-cache, and re-syndicate the supplier's photographs — and to pass them to *your* B2B partners? | 🔴 **See §4b.1.5. This is the genuinely under-discussed one and no collected source answers it.** |

### 4b.1.3 What the incumbents actually did — evidence, not inference

| Company | What is documented | Where | Tier |
|---|---|---|---|
| **GetYourGuide** | **Owns the catalog record and publishes a spec; lets other people's software implement it.** Two distinct APIs — outbound Partner API and inbound Supplier/Connectivity API. Catalog surface includes `/tours`, `/tours/{id}/options`, `/categories`, `/suppliers/{id}`, with `itinerary`, `location`, `picture`, `day_breakdowns`, `metadata` schemas. **250+ connectivity partners**; FAQ names Rezdy, Bokun, FareHarbor, Regiondo "and over 100 others". **50K+ suppliers, 200K+ activities, 12K+ cities, 37 languages.** | `teardown-klook-gyg.md`, `gyg-supply-portal.md` | **A** (spec) / C (marketing counts) |
| **GetYourGuide** | **Human expert review is retained at 200K SKUs.** "Our experts will conduct a thorough review… only top activities are approved," plus a published restricted-activities list. | `gyg-supply-portal.md` | C — but it is GYG describing its own process |
| **Viator** | **Product mapping is done by Viator, manually, during certification** — "Viator-side production config incl. product mapping", then "scheduled go-live initially for a small number of products". **No self-serve path at any step**; realistic time-to-first-booking weeks-to-months, human account manager throughout. Ingestion cadences are **cert-enforced, not advisory**; `/products/modified-since` and `/availability/schedules/modified-since` are **denied at Basic tier**. | `teardown-viator-kkday-headout.md` | **A** |
| **GlobalTix** | **Delta-based ingestion is the mechanism.** `product/changes?countryId=&dateFrom=&dateTo=` with a **max 30-day window** (24h / 7d / 30d), plus `product/newProducts`. `partnerReferenceNumber` exists for idempotency and reconciliation. Two-phase booking (`reserve` → `confirm` → `release`) with a **15-minute hold**. Auth token **expires every 24 hours**. | `teardown-connectivity-indonesia.md` | **A** |
| **Klook / GetYourGuide** | 🔴 **Neither has published anything on cross-supplier deduplication or product matching.** GYG has ~40 engineering posts including hybrid search, transformer ranking, image selection and a cold-start post — **and nothing on entity resolution.** Klook's strongest documented asset is a review-corpus → merchant-feedback loop. | `teardown-klook-gyg.md` | A (absence, in a well-populated blog) |
| **Expedia** | ⚠️ **Nothing in the collected sources.** No Expedia engineering artefact on product matching or lodging/activity deduplication was found in this step. **Unverified — do not assert an Expedia approach.** | — | — |

**The transferable reading, and it is uncomfortable:**

1. **The two biggest players did not automate this. They avoided it.** GYG avoided it by making itself the system of record and forcing suppliers (or their channel managers) to submit into *its* schema. Viator avoided it by doing product mapping **by hand, per partner, behind a certification gate**. Neither published a matching algorithm because neither appears to run one at the scale SatuSatu is contemplating.
2. **Certification and tiered access are quality controls disguised as commercial controls.** Viator denying bulk catalog endpoints to Basic partners is not stinginess; it is a refusal to let an unproven integration pull 300,000 products into someone's database. SatuSatu is on the *receiving* end of that logic and should apply the same discipline to its own second feed: **ingest a bounded subset first, not the whole catalog.**
3. **Delta endpoints exist because full re-crawls do not work.** GlobalTix's 30-day change window and Viator's mandated `modified-since` cadences both encode the same assumption: the catalog is a stream, not a snapshot. Any ingestion design that re-pulls everything nightly will hit rate limits (Viator: 16 req/10s on a rolling window, plus a **system-wide** concurrency limiter that can return 503 for someone else's traffic) [VERIFIED, Tier A].

### 4b.1.4 Automated match accuracy — realistic bands, and an honest caveat about where they come from

**The only measured, methodology-disclosed matching benchmark available is academic e-commerce product matching**, not travel:

| Benchmark | Result | Tier |
|---|---|---|
| WDC Product Data Corpus / Gold Standard for Large-Scale Product Matching (Univ. Mannheim, webdatacommons.org) | **Deep-learning matchers reach F1 ≈ 0.90** on the xlarge training set, **outperforming symbolic matchers by ~16 points F1**. RNN highest, then hybrid / attention / SIF. | **A** — academic, published gold standard |
| Same corpus, scale | 26M product offers (16M English) from **79,000 e-shops**; gold standard **4,400 manually verified pairs** | **A** |

⚠️ **Why this transfers only partially — and it matters.** WDC's F1 ≈ 0.90 is achieved on offers annotated with **schema.org structured markup**, frequently carrying **GTIN/MPN identifiers**, and with an *xlarge* training set of labelled pairs. Travel experiences have **no GTIN**. There is no universal identifier for "the 1-day USS ticket". Titles are marketing copy, not part numbers. And SatuSatu has **no labelled training pairs at all**. Every one of the three conditions that produces F1 0.90 is absent.

**Bands for attraction/experience matching, stated as inference:**

| Matching approach | Precision band (low/base/high) | Recall band | Label | Note |
|---|---|---|---|---|
| **Exact-identifier join** (both feeds expose the same venue/attraction ID or a shared third-party ID) | 0.98 / 0.99 / 1.00 | 0.05 / 0.15 / 0.35 | `[INFERENCE]` | Near-perfect where it fires, fires rarely. **Check first whether the second aggregator and GlobalTix share any venue identifier — if they do, this is the whole solution for the top of the catalog.** |
| **Rules on normalised title + geo + duration + price proximity** | 0.85 / 0.92 / 0.96 | 0.45 / 0.60 / 0.75 | `[INFERENCE]` | Cheap, auditable, tunable toward precision. Anchor: WDC's *symbolic* matchers sat ~16 F1 points below deep learning, i.e. roughly F1 0.74 on an easier problem with identifiers present. |
| **Embedding similarity + rules, human review of the uncertain band** | 0.93 / 0.96 / 0.98 on auto-decided pairs | 0.60 / 0.75 / 0.88 | `[INFERENCE]` | The realistic target. **Fully automated end-to-end accuracy above ~0.96 precision should be treated as a vendor claim until measured on SatuSatu's own pairs.** |
| **Fully automated, no human review** | — | — | ❌ | **Not recommended at any accuracy.** See asymmetry below. |

🔴 **The asymmetry that decides the design.** A **false merge** (two different products collapsed into one) sells the wrong ticket — a customer buys "USS 1-Day" and receives an Express-lane-excluded variant, or a 2-day pass priced as 1-day. That is §4a's F2/F7 failure class with a financial and a reputational leg. A **false split** (one product left as two listings) is a visible cosmetic defect and a small conversion loss. **They are not symmetric and must not be traded against each other with a single threshold.** The correct configuration is: high-precision auto-merge on a narrow high-confidence band, auto-split (i.e. leave alone) by default, and a human queue in the middle. At 253 SKUs growing to ~2,500, that middle queue is a few dozen decisions a week — entirely staffable, and cheaper than any model work.

### 4b.1.5 🔴 Image rights across resold inventory — the real unresolved one

**No collected source, and no source found by direct search, states what rights an OTA reseller receives in supplier photography.** Searching produced only generic photography-licensing material. The relevant Tier C generalities:

- Suppliers commonly provide images through a dealer/merchant portal, usable **"in accordance with the supplier's terms"** — i.e. the terms are the answer and they are per-supplier.
- **Sublicensing is not implied.** A licensee may sublicense **only if the agreement says so**; the default drafting position is that images "shall not be resold, sublicensed, or redistributed."

**Why this is a live, specific risk for SatuSatu rather than a theoretical one:**

1. **Pool A images arrive through an aggregator, which means the chain is at least three links long** — attraction/operator → GlobalTix → SatuSatu. SatuSatu's right to display them depends on GlobalTix having received a sublicensable right, which is not something the API tells you.
2. 🔴 **The D2/D3 platform strategy is a sublicensing act.** Passing Pool A images to a white-label or B2B partner is redistribution to a fourth party. **If the licence chain does not carry a sublicense right, the D2/D3 build has a content problem underneath it that no amount of engineering fixes.** This is the same structural point §4a.3.6 made about the concierge promise — SatuSatu is committing over supply it does not control.
3. **AI image work compounds it.** Re-cropping, upscaling, background removal, or generating variants from supplier photography creates derivative works. Derivative-work rights are a separate grant from display rights.
4. **Pool B is clean, and only Pool B is clean.** Directly contracted operators can be asked for a written, sublicensable, perpetual grant at contract time, at **zero marginal cost** if it is added to the standard operator agreement now.

> **Recommendation, near-zero effort, high value:** (a) add an explicit sublicensable image-and-content grant to the Pool B operator contract template **before** onboarding accelerates — retrofitting 500 signed contracts is impossible; (b) ask GlobalTix, in writing, whether the feed grant includes sublicensing to SatuSatu's distribution partners, and record the answer; (c) do not build any Pool A image-derivative pipeline until (b) returns. **This is a contracting task, not an engineering task, and it is on the critical path for D2/D3.**

### 4b.1.6 Cost per SKU, per feed, in eng-weeks

**Marginal cost per SKU via feed is effectively zero. All the cost is fixed per feed, and it is engineering cost — the one resource that does not exist.**

| Work item | Eng-weeks (low / base / high) | Label | Basis |
|---|---|---|---|
| Second feed: auth, catalog pull, delta handling, retries, rate-limit compliance | 2 / 4 / 7 | `[INFERENCE]` | GlobalTix-class REST surface: ~15 endpoints, 24h token rotation, two-phase booking, webhooks [VERIFIED, Tier A]. Second integration is *not* materially cheaper than the first unless an abstraction already exists — and with one feed live, it almost certainly does not. |
| Availability + pricing reconciliation across two feeds (incl. the ≤30-day window and per-render latency budget) | 2 / 3 / 6 | `[INFERENCE]` | Live dual-feed price selection touches the page render path. |
| Canonical product model + option-level mapping | 2 / 4 / 8 | `[INFERENCE]` | The expensive half. Viator does this **by hand** per partner [VERIFIED, Tier A] — a strong signal that it does not automate cheaply. |
| Deduplication: rules + embedding candidate generation + review queue UI | 2 / 4 / 8 | `[INFERENCE]` | The review queue UI is often forgotten and is a third of the work. |
| Taxonomy crosswalk + Ubud-class geography preservation | 1 / 2 / 3 | `[INFERENCE]` | |
| Booking, cancellation, refund-policy and settlement reconciliation for a second supplier | 2 / 3 / 5 | `[INFERENCE]` | Two suppliers means two settlement files, two refund-policy shapes, two dispute paths. |
| **Total, second feed, done properly** | **11** / **20** / **37 eng-weeks** | `[INFERENCE]` | |
| **Minimum viable: bounded subset, single-source-of-truth-per-attraction, no live price arbitration** | **3** / **5** / **8 eng-weeks** | `[INFERENCE]` | See below. |

**Cost per SKU ingested via feed** = fixed cost ÷ SKUs ingested. At 20 eng-weeks base and an assumed loaded engineer cost, the per-SKU figure is entirely a function of how many SKUs the second feed actually adds:

| SKUs added by feed 2 | Cost per SKU at 20 eng-weeks | Cost per SKU at 5 eng-weeks |
|---|---|---|
| 250 | 0.080 eng-weeks/SKU | 0.020 |
| 2,500 | 0.008 | 0.002 |
| 25,000 | 0.0008 | 0.0002 |

`[INFERENCE]` throughout. **Expressed in eng-weeks deliberately, because eng-weeks are the binding constraint, not dollars** (§4a established the same inversion for labour). Converting to dollars hides the fact that SatuSatu cannot spend this currency at all right now.

> 🔴 **The finding the brief needs.** A second aggregator, integrated properly, is an **11–37 eng-week project**. There is **no dedicated engineering capacity**. Therefore the second aggregator, as normally conceived, **is not a "now ≤90 days" move and cannot be made into one by wanting it more.**
>
> **What *is* available inside the constraint, in descending order of realism:**
>
> 1. **Buy the reconciliation instead of building it.** Route the second feed through a party that already normalises multi-supplier inventory — the channel-manager / connectivity layer (Bokun, Rezdy, FareHarbor, Ventrata, Palisis, PrioTicket, Redeam, rezio, bookingkit are all named in the collected sources; GYG alone lists **250+** such partners [Tier C]). **Effort is a commercial negotiation, not eng-weeks.** ⚠️ It also adds a second margin taker to a ~10% gross margin, which may be fatal on Pool A economics — verify the take rate before anything else.
> 2. **Zero-overlap sourcing.** Contract the second aggregator for **non-overlapping inventory only** (e.g. non-Bali Indonesia, or Singapore where GlobalTix is strongest and a second feed adds least). **If there is no overlap, there is no deduplication problem** — the entire §4b.1.2 stack disappears. This is the cheapest correct answer and it is a contracting decision.
> 3. **Single-source-of-truth per attraction, declared manually.** For the ~20–50 attractions where overlap is real, pick one feed per attraction by hand, in a spreadsheet, and suppress the other. 253 SKUs and even 2,500 SKUs make this tractable. It forfeits price arbitrage; it costs approximately nothing; and it makes duplicates structurally impossible rather than probabilistically rare.
> 4. **Do not do live cheapest-of-two price selection.** It is the single most expensive line in the table, it sits in the render path, and it buys a few points of margin on inventory that only carries ~10% gross to begin with.

---

## 4b.2 AI listing production and quality gating

### 4b.2.1 The one measured case in the entire category — and read it carefully

GetYourGuide, *"Revolutionizing the Experience Marketplace: The dual impact of Generative AI on Suppliers and Consumers"*, **2025-04-23** [VERIFIED, **Tier A** — GYG engineering blog, mechanics quoted verbatim in `teardown-klook-gyg.md`]:

| Fact | Value |
|---|---|
| Product-creation wizard length | **16 steps** |
| Time before | Providers "were sometimes spending **up to an hour**" |
| Input method | Supplier **pastes existing content** — "such as from their own website" — into an input box |
| What the LLM does | Generates long-form description **and fills structured fields** (transport type, location tags) |
| Automation depth | **"The AI auto-completes 8 key steps"** — i.e. **8 of 16** |
| Adoption in pre-tests | **~60%** of the exposed cohort |
| First full experiment (75/25 split) | 🔴 **FAILED.** Submission rate **fell.** Cause diagnosed as **UX and trust, not model quality.** |
| Second experiment, after UI + microcopy fixes | Drop-off cut by **5 percentage points** |
| Observed end-to-end creation time after | **14 minutes** |
| Final state | Rolled out to **100%** |
| Stated motivation | Reduce "a high ratio of travelers needing to contact our care department" — i.e. **bad listing content was generating support volume** |

**Five readings, in order of usefulness to SatuSatu:**

1. **It works and it is quantified: ~60 min → 14 min, ~4.3× on the listing-creation step.** This is the single most transferable AI finding in the whole ops scan.
2. 🔴 **It is only half the form.** 8 of 16 steps. The other 8 — the ones carrying price, availability, capacity, cancellation terms, meeting point — **were not automated by the company with 200K SKUs and four teams on the problem.** Those are exactly the fields whose errors cause §4a's F1/F2/F7 failures. **The boundary GYG drew is the boundary to copy.**
3. **The first experiment failed on UX, not on the model.** For a team with no engineering capacity this is the warning: the model is the easy part and it is not where the project dies.
4. **The input is the supplier's own existing text.** This is not generation from nothing; it is **reformatting and field-extraction from source material the operator already has.** That is a much smaller, much safer task — and it maps directly onto SatuSatu's offline-operator onboarding, where the operator has a brochure, a WhatsApp price list, or a Facebook page.
5. **The business case was support-cost reduction, not content volume.** GYG's own framing ties listing quality to care contacts. For SatuSatu, with a human concierge already absorbing every ambiguity (§4a), bad listings are paid for in concierge minutes at $0.079/min.

**GYG also retains a human approval gate at 200K SKUs**: "our experts will conduct a thorough review… only top activities are approved," plus a published restricted-activities list [Tier C, GYG's own supply site]. 🔴 **AI-assisted creation and human approval coexist at the largest supplier base in the category. They are not alternatives.**

### 4b.2.2 What a quality gate actually checks — the prerequisite, specified

The brief is right that this is the prerequisite, not the optimisation, and the reason is structural: **AI listing production increases throughput; without a gate it increases throughput of defects at the same multiple.** A 4.3× faster pipeline with no gate is a 4.3× faster defect pipeline.

A gate is three layers. **Only the first two are affordable here, and they catch most of what matters.**

**Layer 1 — Deterministic / structural. No model. Cheap, near-perfect precision, must be 100% blocking.**

| Check | Blocks on |
|---|---|
| Required-field completeness | Missing price, currency, duration, location, cancellation policy, meeting point, capacity |
| Price sanity | Zero, negative, absent currency, order-of-magnitude outlier vs category median, IDR/USD unit confusion (a live risk in a dual-currency catalog) |
| Currency and FX consistency | Feed currency vs display currency vs settlement currency — §4a and the teardowns both flag IDR settlement gaps at GYG and Viator [VERIFIED, Tier A] |
| Geo validity | Coordinates present, inside Indonesia/Singapore bounding box, city ID resolves in the taxonomy |
| Date/seasonality validity | Departure or validity window not already in the past; season not expired |
| Media presence and technical fitness | ≥N images, minimum resolution, aspect ratio, no placeholder/watermark, decodable |
| Duplicate detection | Content hash + the §4b.1.4 candidate check before publish, not after |
| Cancellation-policy derivation | 🔴 Published policy **must** be derived from the feed's machine-readable fields — GlobalTix exposes `isCancellable` and `cancellationPolicy {percentReturn, refundDuration}` per ticket [VERIFIED, Tier A]. Any manually written policy more generous than `percentReturn` is a direct margin leak (§4b.5). |
| Availability model declared | Every SKU must carry an explicit availability model (§4b.4). **No SKU publishes without one.** |

**Layer 2 — Consistency / cross-field. Rules plus light model use. High value, moderate cost.**

Description contradicts structured fields (text says "3 hours", `duration` says 8); inclusions contradict price tier; text asserts hotel pickup while no transfer option exists; language/locale mismatch; text mentions a facility or timing the operator did not contract; prohibited-content and restricted-activity screening (GYG publishes such a list [Tier C]).

**Layer 3 — Semantic / editorial. LLM-as-judge or human. Expensive, weakest evidence base.**

Factual claims not supported by the source material the operator supplied; tone and reading level; unverifiable superlatives; translation adequacy. ⚠️ **No source in this research measures LLM-as-judge accuracy on travel listing copy. Treat Layer 3 as unproven and staff it with humans on a sample basis, not as an automated gate.**

> 🔴 **The design point that saves the most money: put the gate at *ingest*, not at *publish-review*.** A defect caught by a Layer-1 rule costs zero human minutes. The same defect caught by a reviewer costs 2–10 minutes; caught by a traveller it costs concierge minutes plus a possible stranding (§4a.3.5). **Layer 1 is a week of rules work and it is the highest-ROI item in this entire section.**

### 4b.2.3 Auto-pass rates — 🔴 no measured figure exists, bands are inference only

**No source, vendor or academic, gives an auto-pass rate for a travel-content quality gate.** Vendor claims in this space are not merely inflated (§4a.2.2) — they are **undefined**, because "auto-pass" depends entirely on gate strictness, which the vendor sets.

| Input to the gate | Auto-pass, low / base / high | Label |
|---|---|---|
| **Clean aggregator feed** (GlobalTix-class, already validated upstream by a professional aggregator) through **Layer 1 only** | 82% / 90% / 97% | `[INFERENCE]` |
| Same feed through **Layers 1+2** | 65% / 78% / 88% | `[INFERENCE]` |
| **AI-generated listing from operator-supplied source material** (the GYG pattern), Layers 1+2 | 45% / 62% / 78% | `[INFERENCE]` |
| **Bulk import of an unnormalised offline-operator spreadsheet** (the SatuSatu manual-onboarding case), Layers 1+2 | 20% / 40% / 60% | `[INFERENCE]` |
| **AI-generated content through Layer 3 as well** | — | **Not estimable. No basis.** |

⚠️ **The trap in this metric, stated so it cannot be misused downstream:** **auto-pass rate and gate strength move in opposite directions.** A 95% auto-pass rate is evidence of a weak gate as readily as a clean pipeline. **The metric to manage is not auto-pass rate — it is escaped-defect rate: defects found after publish ÷ SKUs published.** Auto-pass is a throughput metric; escaped-defect rate is the quality metric, and only the second one protects the traveller. This is the same inversion §4a made about deflection versus error rate, and it should be treated as the same lesson.

### 4b.2.4 Multilingual — costed, with the per-word rates that exist

| Reference | Figure | Tier |
|---|---|---|
| Professional human translation | **$0.07–0.14 / word** | C |
| Machine-translation post-editing (MTPE) | **$0.04–0.12 / word** (a second source: $0.06–0.12) | C |
| LLM raw MT — a fast/cheap model, quoted April 2026 | **~$1.05–1.14 per million words** | C |
| LLM raw MT — a frontier model via batch API | **~$13.33 per million words** | C |
| Quality reference | COMET **≥0.85** described as professional, **≥0.90** as near-human | C |

**Per-SKU translation cost.** Assume a listing of 250 / 400 / 700 words [INFERENCE]:

| Route | Cost per SKU per language | Label |
|---|---|---|
| **Raw LLM MT, no human touch** | **$0.0003 / $0.0005 / $0.009** | `[INFERENCE]` from Tier C per-million rates |
| **MTPE (human post-edit)** | **$10 / $32 / $84** | `[INFERENCE]` from Tier C per-word rates |
| **Full human translation** | **$18 / $44 / $98** | `[INFERENCE]` |

🔴 **The gap between raw MT and post-edited MT is four to five orders of magnitude. The entire cost of multilingual is the human, not the model.** A 2,500-SKU catalog in 4 additional languages is **~$5 of inference** or **~$320,000 of post-editing** at base rates. That is the whole decision.

⚠️ **But the MTPE rates are Western per-word rates.** Post-editing Indonesian↔English↔Chinese/Japanese/Korean by an Indonesian-based content-ops FTE costs §4a's **$0.079/productive minute**, not $0.08/word. At 400 words and 8 minutes of post-editing, that is **$0.63/SKU/language** — a ~50× reduction versus the Tier C MTPE rate, and the same labour-arbitrage inversion §4a.5.5 found for AI support pricing. **Any multilingual business case built on Western per-word rates will recommend the wrong thing here.**

**Sequencing recommendation:** raw MT everything for *retrievability* (SEO/AEO, §4b.6); human post-edit only the fields where an error strands or misprices a traveller — meeting point, inclusions/exclusions, cancellation terms, safety notes. **That is ~15% of the words and ~90% of the risk** `[INFERENCE]`.

### 4b.2.5 Media handling

Evidence that image work pays at scale, and that it is a *selection* problem before it is a *generation* problem:

- GYG, *"How We Use AI to Optimize Travel Images"* (2024-01-17) and *"replacing intuition with data in visual selection"* (2025-12-04) [VERIFIED, Tier A — titles and dates; full text not read]. **Both are about choosing among existing images, not creating them.**
- GYG supplier marketing: "Create listings in minutes. Use our AI-powered tools to write listings **and add photos**, pricing, and availability" [Tier C].

**For SatuSatu:** the affordable media work is Layer-1 technical validation (presence, resolution, aspect, watermark, placeholder detection) plus **hero-image selection from what the operator already supplied**. Generation and heavy derivative processing are blocked upstream by the rights question in §4b.1.5, not by cost.

### 4b.2.6 What proportion of a bulk-ingested catalog is unsellable

🔴 **Not found. Searched directly.** The PIM and product-data-quality literature quantifies *impact* ("83% of shoppers abandon on insufficient product information"; "returns cost 50–75% of sale price"; "one significant data-quality issue per 15 tables monitored across 11M tables") and names the failure modes — "**catalog debt**: missing attributes, inconsistent terminology, misclassified categories, duplicate records, untagged assets" — but **publishes no incidence rate.** All Tier C.

**Bands, `[INFERENCE]`, decomposed so each component can be replaced with a measurement rather than the total being re-guessed:**

| Unsellable class | Share of a bulk-ingested catalog (low / base / high) | Reasoning |
|---|---|---|
| **Dead / delisted / operator inactive** | 3% / 8% / 18% | Delta endpoints exist precisely because this churns (GlobalTix `product/changes`, Viator mandated `modified-since`) [Tier A]. |
| **Mispriced or price-incomplete** | 2% / 6% / 14% | Seven price points per ticket type, `isDynamicPrice`, `demandType`, plus a price-update webhook [Tier A] = many ways to be stale or wrong. |
| **Duplicated** (only exists with ≥2 feeds) | 0% today / 5% / 15% at two feeds | Concentrated on the highest-demand SKUs (§4b.1.1). |
| **Seasonal-expired / out-of-window** | 2% / 5% / 12% | Availability horizon is 6 months forward [Tier A]; anything referencing a past window is dead on arrival. Bali seasonality compounds it. |
| **Unsellable for content reasons** (no usable image, no meeting point, untranslatable, restricted activity) | 3% / 8% / 16% | GYG maintains a restricted-activities list and a human approval gate at 200K SKUs [Tier C]. |
| **Total, de-overlapped** | **~10% / ~25% / ~45%** | `[INFERENCE]` |

**One verified data point of SatuSatu's own, and it is worse than any of the above.** The displayed sold-counts on Pool A SKUs are **inherited feed values, not SatuSatu transactions** — USS displays "99k+ sold" [VERIFIED, per brief; corroborated by `satusatu-uss-catalog.md`]. **That is a 100% defect rate on one live, customer-visible field across the entire Pool A catalog.** It is not "unsellable", but it is a concrete demonstration that bulk-ingested fields arrive wrong and ship to customers unchallenged. Any planning number for catalog quality should be anchored on that observation rather than on an optimistic prior.

> **Planning consequence: a "2,500 SKU catalog" ingested in bulk is realistically a ~1,900 sellable-SKU catalog at base case, and the 600 bad ones are distributed across the highest-traffic pages.** Budget the gate, or budget the concierge minutes and refunds instead.

---

## 4b.3 Search and ranking at scale

### 4b.3.1 What actually changes between 250, 2,500 and 25,000 SKUs

Today: keyword plus location filter, zero AI [VERIFIED, per brief].

| Dimension | ~250 SKUs (today) | ~2,500 SKUs | ~25,000 SKUs |
|---|---|---|---|
| **Can a user see the whole relevant set?** | **Yes.** A category + location filter enumerates the catalog. Browse substitutes for search. | **No.** A popular cell (Bali × Water Sports) exceeds one screen. Ranking becomes the product. | No. Retrieval quality dominates; unranked long tail is invisible and therefore dead inventory. |
| **Dominant failure** | **Zero results / vocabulary mismatch.** "quad bike" vs listing titled "ATV Ride"; "monkey forest" vs "Ubud sacred sanctuary". Cheap to fix with synonyms. | **Wrong ranking + near-duplicate crowding.** Same attraction from 4 suppliers occupying the first screen. | **Retrieval recall.** Relevant SKUs never enter the candidate set. |
| **Taxonomy** | Flat categories survive. | Needs sub-categories and facets (duration, pickup, difficulty, private/shared, indoor/outdoor). | Needs attribute extraction pipeline; hand-curation stops scaling. |
| **Right investment** | Synonym/alias dictionary, query logging, curated collections. **No ML.** | Facets, business rules, hybrid retrieval, dedup at retrieval time. | Learned ranking, and only here. |
| **Incumbent reference** | — | — | GYG shipped hybrid lexical+semantic retrieval (2024-11-20), real-time production ranking (2024-05-22), and GBDT→transformer ranking (2025-12-11) — **at 200K activities across 12K cities** [VERIFIED, Tier A titles/dates]. Viator's own B2B `/search/freetext` is **keyword** and its `/products/recommendations` method is undisclosed [VERIFIED, Tier A]. |

⚠️ **Note the asymmetry in the incumbent evidence: GYG and Klook ship semantic search in their consumer storefronts and expose none of it in their B2B APIs** [VERIFIED, Tier A — `teardown-klook-gyg.md`: "The AI is in the storefront, not in the pipe"; `teardown-viator-kkday-headout.md`: "Nowhere"]. If SatuSatu's D2/D3 ambition includes a *searchable* partner API, **no incumbent has done it** — which is both an opportunity and a warning that demand for it is unproven.

### 4b.3.2 🔴 Cold start when the engagement data was fabricated by someone else's platform

This is the most unusual problem in the section and it deserves to be treated as a defect, not a curiosity.

**The facts.** Pool A SKUs arrive from GlobalTix carrying **inherited sold-counts** — USS shows "99k+ sold" [VERIFIED]. Those are not SatuSatu transactions. Pool B SKUs, being direct-contracted and exclusive, carry **no inherited counts at all** and only whatever small volume SatuSatu has genuinely produced.

**Three consequences, in order of severity.**

**1. 🔴 Any popularity-weighted ranking systematically demotes the only inventory worth selling.** The field is large and non-zero on Pool A, near-zero on Pool B. Any ranker, sort option, "Popular" badge, or homepage module that reads it — whether learned or a single `ORDER BY` — permanently puts **~10%-gross-margin, non-exclusive** inventory above **~27%-gross-margin, exclusive** inventory.

Quantified: Pool A ≈ 10% gross / ~7% after payments and FX; Pool B ≈ 27% gross [per brief]. **Every unit of GMV diverted from Pool B to Pool A costs ~17 points of gross margin, ~20 points net of payments and FX** `[INFERENCE]`. At an illustrative 10% of sessions influenced, that is a margin leak of the same order as the entire Pool A gross margin. **A number pasted in from a supplier feed is silently steering the mix away from the strategy.** This is the single cheapest, highest-value fix in §4b.3.

**2. Displaying another platform's transaction count as your own is a trust and possibly an advertising-accuracy question.** ⚠️ **Flag, not a legal opinion** — route it with the two counsel questions already open in §4a.3.6 and §4a.5.8. It is the same governance bundle.

**3. There is no learned ranker to train, and there will not be one soon.** GYG has a dedicated engineering post on the **Cold Start Problem** (*"From Zero to Hero: GetYourGuide's intelligent approach to new activities"*, 2025-08-20, Tier A, title/tag level only) — i.e. cold start is a named, funded ML workstream **at 200K SKUs with 40M monthly visitors** [Tier C]. At 253 SKUs with unknown but far smaller traffic, per-SKU engagement data is too sparse for per-SKU learning by orders of magnitude. `[INFERENCE]` floor for per-SKU learned signals: **~30–50 own booking events per SKU minimum**, and realistically **100+** for a stable estimate. A 2,500-SKU catalog would need 75,000–250,000 own bookings to have that everywhere. **The honest answer is that learned ranking is not on the roadmap at any point covered by this research.**

**Recommendations — all near-zero engineering, in priority order:**

| # | Action | Effort |
|---|---|---|
| 1 | 🔴 **Split the field.** `supplier_sold_count` (from feed — **never displayed, never ranked, never used as a feature**) vs `satusatu_bookings` (own events — displayable, rankable). One schema change plus a template edit. | Hours |
| 2 | 🔴 **Audit every surface that currently reads the inherited count** — sort options, badges, "popular" modules, homepage rails, category default order. **Assume it leaked into more places than anyone remembers.** | Days |
| 3 | **Rank with an explicit margin-and-exclusivity business rule** (Pool B boost) rather than a model. It is a config value, it is auditable, it directly encodes the commercial strategy, and it can be A/B tested. | Days |
| 4 | If a popularity signal is wanted now, use **own-platform impressions, detail-page views, saves and add-to-cart** — they accumulate 10–100× faster than bookings and are genuinely SatuSatu's. | Days |
| 5 | **Re-anchor the curation promise on Pool B exclusivity, not on catalog signals.** The brief already anticipates this. Pool B is the only inventory where SatuSatu can truthfully claim selection, because it is the only inventory it selected. | Copy |

### 4b.3.3 When semantic search starts to pay — a threshold, not an opinion

**Threshold: ≈2,000 SKUs, band 1,200–4,000** `[INFERENCE]`.

**Derivation, stated so it can be checked rather than believed.** Semantic retrieval pays when a user can no longer see the relevant set by browsing, i.e. when the busiest **taxonomy cell** (category × sub-location) exceeds roughly one to two screens — call it ~30–50 SKUs. A Bali-dominant catalog has on the order of **10–14 meaningful categories × 6–10 sub-locations** (Ubud, Canggu/Seminyak, Kuta, Uluwatu, Nusa Dua, Nusa Penida/Lembongan, North Bali, plus non-Bali and Singapore), and the distribution is heavily skewed — the top cell holds perhaps 4–6% of the catalog. Top cell reaches 40 SKUs at roughly **40 ÷ 0.05 ≈ 800–2,000 total SKUs**. Add the second driver — **cross-supplier near-duplicates crowding the first screen**, which begins the day feed two lands — and the practical threshold sits at the upper end of that range.

🔴 **But SKU count is the wrong trigger to *manage*. Instrument these instead, and let them fire:**

| Trigger metric | Threshold to act | Why |
|---|---|---|
| **Zero-result rate** on non-junk queries | **>10%** | The direct measure of vocabulary mismatch. Fix with synonyms first; only if synonyms plateau does it justify embeddings. |
| **Null-click rate** (search with no result clicked) | **>40%** | Results returned but none relevant — a ranking/retrieval problem, not a coverage problem. |
| Share of sessions using search rather than browse | **>25%** | Below this, browse is the product and search investment is misallocated. |
| Queries per session with ≥2 reformulations | **>15%** | Users compensating for bad retrieval. |

**Below ~1,500 SKUs, the correct spend is not embeddings.** It is: a synonym and alias dictionary (English/Indonesian/Chinese variants; "quad bike"→ATV, "monkey forest"→Ubud sanctuary, "banana boat"→water sports), query logging with a weekly zero-result review, curated collections, and facets. **All of that is content-ops work at $0.079/minute, not engineering** — and it is what actually moves search success in a small catalog. Embedding retrieval on 253 SKUs solves a problem the catalog does not have yet, while the zero-result log goes unread.

## 4b.4 🔴 Availability for operators with no booking system

### 4b.4.1 The size of the problem — this is an industry condition, not a SatuSatu deficiency

| Fact | Value | Tier |
|---|---|---|
| Share of tour operators worldwide with **no booking system at all** | **39%** | C reporting a **B** primary — attributed to **Arival, *State of Booking Tech* (2025)**, read via a third-party trade blog, **not read at source.** ⚠️ Verify against Arival directly before external use. |
| Industry revenue running through "spreadsheets, WhatsApp groups, and manual data entry" | **~$271bn/yr** | same source, same caveat |
| Share of experience operators that are small or micro | **>70%** | B — Arival, established anchor |
| Experiences sold online | **33%**, reaching **42–43% by 2029** vs 64% for travel overall | B — Phocuswright/Arival, established anchor |

🔴 **These four numbers, read together, say something important: Pool B's lack of real-time availability is the normal state of the category, not a sourcing mistake.** 39% of operators globally have no system; >70% are small or micro; two-thirds of experience bookings still do not happen online. **A strategy that requires Balinese long-tail operators to adopt software before SatuSatu can sell them is betting against every one of those numbers.** Conversely, the fact that this is normal means **nobody has solved it** — which is exactly why it is the most defensible thing on SatuSatu's list.

### 4b.4.2 The seven real patterns, with what each demands of the operator

The critical column is the last one. Anything requiring operator adoption is, per §4b.4.1, a low-probability bet on the specific population SatuSatu has contracted.

| # | Pattern | How availability is established | Cost to SatuSatu | **Requires the operator to…** | Verdict for Pool B |
|---|---|---|---|---|---|
| **1** | **Aggregator real-time API** (Pool A today) | Live pull. GlobalTix `checkEventAvailability`, ≤30-day query window, 6-month horizon, timeslot-aware; two-phase `reserve`→`confirm`→`release` with a **15-min hold** [VERIFIED, Tier A] | Feed integration eng-weeks (§4b.1.6) | **Nothing** — the operator was onboarded by the aggregator | ✅ Best data quality. ❌ Not available for Pool B by definition; and it is the ~10% margin pool. |
| **2** | **Operator adopts a res system / channel manager** | Operator maintains a calendar; SatuSatu consumes it via the CM. Named systems in the collected sources: **Rezdy, Bokun, FareHarbor, Regiondo, Ventrata, Palisis, bookingkit, TicketingHub, rezio, PrioTicket, Redeam, ExperienceBank, TourCMS.** GYG lists **250+** connectivity partners; its FAQ names Rezdy/Bokun/FareHarbor/Regiondo "and over 100 others" [Tier A/C]. GlobalTix also sells a merchant SaaS with a website-booking-fee + setup tier ladder [Tier A] | Per-operator setup labour; possibly a CM take rate | 🔴 **Buy software, learn it, and keep a calendar current daily — forever.** Pricing unverified (Arival publishes a res-system pricing guide; not read here) | ❌ **Reject as the primary path.** This is the pattern that fails for the >70%-micro population. Keep as an *opportunistic* path for the largest 5–10% of Pool B operators who already have some system. |
| **3** | **Allotment / hard allocation** | Operator commits N units per day (or per departure) to SatuSatu in advance; SatuSatu sells against that block with **no per-booking confirmation**; unsold inventory releases back at an agreed cut-off | A commercial negotiation + a spreadsheet or a table. **Near-zero engineering.** Risk: paying for or losing unsold allotment, and operator double-selling the same capacity elsewhere | ✅ **Nothing technological.** Agree a number, honour it | 🔴 **The single best answer for Pool B.** It converts "no availability system" into "availability by contract". It is the pattern the whole category has used for decades precisely because it needs no software on the supply side. |
| **4** | **Freesale with capacity caps and lead-time cut-off** | Sell without checking. Cap by max-per-day and a minimum lead time, both derived from observed history and operator statement | Oversell risk = §4a's **F2** failure, which for fixed-departure products is unrecoverable | ✅ **Nothing** | ✅ **Correct for high-capacity, low-constraint SKUs**: private car charter, ATV, spa, most transfers, temple entry. ❌ **Never** for fixed-departure/fixed-seat: dive boats, Nusa Penida crossings, small-group treks, sunrise Batur. **The product-level split is the whole discipline.** |
| **5** | **Static schedule + blackout calendar** | Operator states a fixed weekly schedule **once**, plus notifies blackout dates (ceremonies, Nyepi, maintenance, private hire) | One onboarding form + a blackout intake channel | ✅ Effectively nothing — one conversation, then exception notices | ✅ **Strong for temple/dance performances, scheduled transfers, fixed-departure day tours.** Covers a real slice of Pool B at almost no cost. |
| **6** | **Request-to-book with a published SLA** | Customer books; SatuSatu confirms within a stated window (e.g. 2h in-hours / 12h overnight); decline triggers refund or alternative | Conversion loss (customer waits, some abandon), ops labour, refund handling on decline | ✅ **Nothing beyond answering WhatsApp** | ✅ **The honest default for genuinely constrained Pool B SKUs** — and vastly better than the status quo because it makes the commitment **explicit and measurable** rather than implicit. |
| **7** | **Phone / WhatsApp ops desk (the status quo)** | A human asks, a human answers, a human writes it down | §4a costed this: Pool B supplier booking = **5 / 15 / 40 min** per booking → **$0.40 / $1.19 / $5.52** at §4a's $0.079/min [INFERENCE] | ✅ Nothing | ⚠️ Works, does not scale, and is **unreconcilable** (§4b.5). It is the fallback for pattern 6's exceptions, not a model in itself. |

### 4b.4.3 🔴 The only realistic set: patterns 3, 4, 5, 6 — and the design that follows

**Patterns 3, 4, 5 and 6 require the operator to adopt nothing.** That makes them the only viable set, and it also means the real work is **classification, not technology**.

> 🔴 **The core recommendation of this entire section: every Pool B SKU must carry a declared `availability_model` field with exactly one of four values — `allotment`, `freesale_capped`, `static_schedule`, `request_to_book` — plus the parameters that model needs (allotment size and release cut-off; daily cap and lead time; weekly schedule and blackout list; SLA window).**
>
> - It is **one schema field, four code paths, and a classification exercise** across ~150–250 Pool B SKUs. Days of work, not eng-weeks. It fits inside "now ≤90 days" and inside near-zero engineering.
> - It converts an unbounded human problem into a bounded data problem.
> - It makes the traveller-facing promise honest per SKU: instant confirmation on the first three, an explicit SLA on the fourth.
> - It is the **prerequisite for everything else in Pool B** — pricing, B2B, ranking, and the concierge automation §4a scoped.

### 4b.4.4 🔴 Why this is the gating dependency for the entire B2B strategy

Stated bluntly: **a B2B partner cannot consume "WhatsApp".**

- Every B2B integration — API, white-label, agent portal, or a reseller's own channel manager — requires two machine-readable things: **is it available**, and **is it confirmed**. Pool B currently has neither.
- The brief establishes that **B2B monetises Pool B, not Pool A.** Pool A feeds buy coverage and search-success; they do not carry B2B margin at ~10% gross / ~7% net, because there is nothing to share.
- Therefore: **B2B revenue depends on Pool B; Pool B depends on an availability model; the availability model does not exist. That is the whole critical path, and it is one field plus a classification exercise.**
- §4a reached the same conclusion from the service side, independently: Pool B cannot be exposed to partner traffic *"only if Pool B gets a confirmation state machine"*, because otherwise SatuSatu absorbs F2 at volumes it cannot staff. **Two different analyses, opposite directions, same single blocking dependency. That convergence is the strongest signal in this document.**

### 4b.4.5 Where AI genuinely helps — and where it does not

Ranked by honest expected value, most sceptical column last.

| # | Application | Verdict | Realistic accuracy / effect | Label |
|---|---|---|---|---|
| **1** | **Auto-chasing unconfirmed bookings** — scheduled follow-ups to the operator until a confirmation state is reached, escalating to a human at T-minus-X | 🔴 **Do this first.** And note honestly: **it is a scheduler plus a message template. It barely needs AI at all.** It directly attacks §4a's F2 (the booking that was never acknowledged), which is the highest-severity Pool B failure. | Eliminates the *silent* unconfirmed booking, which is the failure mode that strands people. Effect size unmeasured. | `[INFERENCE]` |
| **2** | **Parsing WhatsApp replies into a structured availability/confirmation state** | ✅ **Genuinely useful and correctly bounded** — the output is a state transition on a record that already exists, not free text sent to a traveller. | Clean text, single language: **85% / 92% / 97%**. Real Bali conditions — Indonesian/English code-switching, voice notes, photos of handwritten notes, "ok" meaning three different things: **60% / 75% / 88%**. | `[INFERENCE]` — no measured source |
| | ↳ **The design rule that makes #2 safe** | 🔴 **Bias precision on `confirmed`, never on `unconfirmed`.** A false `confirmed` strands a traveller (F2). A false `unconfirmed` costs one chase message. **Auto-confirm only on high-confidence explicit affirmation; route everything else to a human.** The asymmetry is ~100:1 in cost and the threshold must reflect that, not a balanced F1. | Target: **≥99% precision on auto-confirm**, accepting whatever recall that leaves (plausibly only 50–70% automated). | `[INFERENCE]` |
| **3** | **Structured extraction from operator source material at onboarding** (brochure, price list, Facebook page, WhatsApp text → structured SKU fields) | ✅ **This is the GYG pattern (§4b.2.1) applied to Pool B**, and it is the best-evidenced AI use in the whole section: 8 of 16 fields auto-completed, 60 min → 14 min, 100% rollout [VERIFIED, Tier A]. | ~4.3× on the content-authoring step only. **Not** on negotiation, verification or capacity. | `[V]` for GYG's own figure |
| **4** | **Predicting availability from history** | ❌ **Not viable now. Be sceptical.** Per-SKU history is far too thin (§4b.3.2 established the same sparsity for ranking), and the cost of a wrong prediction is an **oversell** — i.e. the unrecoverable failure. A `max(observed)` cap with a safety margin is a one-line rule that gets most of the benefit at zero risk of model error. | Revisit only for SKUs with **100+ observed departure-days** of own data `[INFERENCE]`. Almost no Pool B SKU will reach that soon. | `[INFERENCE]` |
| **5** | **AI outbound voice/chat to the operator to request availability** | ⚠️ **Behavioural risk, not technical risk.** A micro-operator who realises they are being messaged by a bot may simply stop replying — and SatuSatu's Pool B relationship is exclusive and personal, which is the asset. **Effect on operator response rate is unverified and could be negative.** | Unverified. | — |
| **6** | **AI as a substitute for the availability model** | ❌ **Category error.** No model can infer capacity that the operator has never stated. Patterns 3–6 in §4b.4.2 are contract and data-model work. **AI accelerates operating the model; it cannot replace declaring it.** | — | — |

---

## 4b.5 Leakage and ops

### 4b.5.1 The benchmark rates that exist

| Metric | Low | Base | High | Tier | Note |
|---|---|---|---|---|---|
| **Chargeback rate, travel industry** | **0.89%** | ~1.2% | **2.0%** | C | Two payments-industry sources agree on the 0.89%–2% band; 0.89% quoted as the travel average |
| Visa/Mastercard high-risk designation threshold | — | **1.0%** | — | C | Above 1% of transactions disputed → high-risk monitoring |
| Travel disputes, YoY change | — | **+30%** | — | C | Attributed to an Outpayce survey; **friendly fraud** named as the main driver |
| Travel-industry chargebacks, total | — | **~$25bn (2023)** | — | C | Directional scale only |
| MCC 4722 (travel agencies & tour operators) | — | recognised elevated-risk code | — | C | Structural, not a rate |
| Cart abandonment, travel | — | **~81%** | — | C | Funnel metric, not leakage |
| **No-show rate, tours & activities** | — | 🔴 **NOT FOUND** | — | — | **Searched directly. No published TAA no-show benchmark exists in accessible sources.** The nearest result was a wedding-venue show rate of 70–85%, which is a deposit-driven, single-date, high-consideration product and **does not transfer.** Do not use it. |
| Operator-side gross margin by tour type (context for Pool B) | budget 15–20% | day tours **40–50%**; multi-day 25–35% | luxury 50%+ | C | Useful: SatuSatu's **27% Pool B margin is being carved out of an operator's 40–50% day-tour margin.** There is room in day tours; there is much less in multi-day. |

### 4b.5.2 🔴 Why normal travel leakage rates are not survivable at Pool A margins

**Loss on a chargeback is not the margin — it is the cost of goods.** SatuSatu refunds the traveller but has already bought (or committed to buy) the ticket. So:

```
gross profit destroyed per chargeback ≈ COGS + chargeback fee
                                      = (1 − gross_margin) × value + fee

clean bookings needed to replace one chargeback
                                      ≈ (1 − gross_margin) / gross_margin
```

| Pool | Gross margin | Clean bookings destroyed by ONE chargeback | Share of pool gross profit lost at a **0.89%** chargeback rate | at **2.0%** |
|---|---|---|---|---|
| **Pool A** | 10% (≈7% after payments+FX) | **~9** (≈13 on the 7% net figure) | **~8%** (≈12% on net) | **~18%** (≈27% on net) |
| **Pool B** | 27% | **~2.7** | **~2.4%** | **~5.4%** |

`[INFERENCE]` — arithmetic on the brief's margin figures and the Tier C chargeback band.

> 🔴 **An industry-average chargeback rate consumes roughly a tenth of Pool A's gross profit and a quarter of it at the high end of the normal band.** Leakage tolerances that are unremarkable elsewhere in travel are near-existential at 7% net margin. **This makes chargeback and refund control a margin strategy, not a finance-team hygiene item — and it is another independent argument for shifting mix toward Pool B.**

### 4b.5.3 Refund leakage — a concrete, cheap fix

GlobalTix exposes per-ticket **`isCancellable`** and **`cancellationPolicy {percentReturn, refundDuration}`** [VERIFIED, Tier A]. The leakage mechanism is simple and common: **SatuSatu publishes a customer-facing cancellation policy more generous than the supplier's `percentReturn`, and eats the delta on every cancellation.** On a 10% gross margin, a 50-point refund mismatch on a single booking wipes out the margin on five clean ones.

> **Fix: the customer-facing policy must be *rendered from* the feed field, never authored independently.** This is the same "render, never generate" principle §4a.3.3 established for concierge messages, applied to commercial terms. It is a template change plus a Layer-1 gate rule (§4b.2.2). **Hours of work; protects margin permanently.**

### 4b.5.4 Reconciliation and supplier settlement — Pool A is instrumented, Pool B is not

**Pool A [all VERIFIED, Tier A]:** `partnerReferenceNumber` gives idempotency and a reconciliation key; two-phase `reserve`→`confirm`→`release` with a 15-minute hold gives a state machine; a **`booking-ticket-redeem` webhook** delivers redemption state; a ticket PDF link exists; `transaction/cancel` gives a cancellation record. ⚠️ **But redemption is asymmetric: there is no reseller-side redemption/scan endpoint.** SatuSatu receives redemption events; it cannot operate the gate. Running redemption for Pool B operators would mean buying GlobalTix as a *merchant*, not consuming it as a reseller.

**Pool B: none of the above exists.** No booking record with a supplier-side identifier, no confirmation state, no redemption event, no cancellation record. 🔴 **Pool B is unreconcilable by construction** — you cannot reconcile against a WhatsApp thread. Every leakage class (no-show, unredeemed, double-paid, operator-disputed, quietly-never-delivered) is invisible in Pool B today.

**Same root cause as §4b.4. The `availability_model` field plus a booking state machine fixes availability, B2B *and* reconciliation with one piece of work.** That is why it is the highest-priority item in this section.

**Settlement and FX — named, quantified, and partly negotiable:**

| Line | Figure | Tier |
|---|---|---|
| GlobalTix FX markup on any settlement currency **other than SGD** | **1.50%** (SGD: 0; AED an outlier at 0.50) | **A** — read from the live API response |
| GlobalTix `creditCardFee` on traded majors | **3.00** | **A** |
| GlobalTix IDR status | **First-class settlement currency**, full markup/rounding/card-fee config | **A** |
| GetYourGuide Partner API checkout currencies | AED, AUD, CAD, CHF, EUR, GBP, NZD, PLN, SEK, USD — **IDR excluded** (display-only) | **A** |
| Viator: merchant invoicing currencies | **5 only** (GBP, EUR, USD, CAD, AUD) vs **27** for affiliates — **IDR is not a merchant currency** | **A** |

> **The 1.50% GlobalTix FX markup is ~15% of Pool A's 10% gross margin and ~21% of the 7% net figure.** It is a vendor-set line item, it is already inside the "~7% after payments+FX" number, and per the teardown it is "the cheapest thing on this list to negotiate away or route around" — **ask whether SGD settlement is contractually available.** Nothing else in this section returns as much margin for as little effort.

## 4b.6 Content at catalog scale — SEO and AEO

### 4b.6.1 Why AEO is a live 12–24 month issue here specifically

The two established anchors point in opposite directions and **both are true at once**: the **discovery** layer is being disintermediated, while the **booking** layer is not (experiences are 33% online, reaching only 42–43% by 2029). **So the risk is not losing the transaction to an AI assistant. It is not being the source the assistant retrieves when the traveller asks it what to do in Ubud.**

Measured evidence that this is already happening in travel:

| Metric | Value | Tier |
|---|---|---|
| AI-source traffic to **US travel sites**, y/y growth, **May 2026** | **+194%** | **B** — Adobe Analytics, reported by Search Engine Land and Adobe's own blog |
| Same, cumulative since Oct 2024 (start of tracking) | **+2,215%** — a record high | **B** |
| AI-referred travel traffic **conversion** vs non-AI | **−28%** — but the gap is **~70% narrower** than Oct 2024 | **B** |
| AI-referred travellers: engagement | **+21% more engaged**, **+70% longer per visit**, **−41% bounce rate** | **B** |
| AI search as a share of **total** site traffic, typical site | **0.5–3%** (up to 5–8% for tech/e-commerce) | C |
| Share of AI referral traffic from ChatGPT | ⚠️ **Two Tier C sources contradict each other in the same month: one says 92.4%; another says Gemini's share more than doubled while ChatGPT "slid into the mid-sixties, then lower still."** | C — **report the contradiction, do not pick a number** |

**Reading:** the *volume* is still small (0.5–3% of traffic) but the *growth rate* is extreme and the *quality* is improving fast — the conversion gap closing by ~70% in 19 months is the more important number than the traffic multiple. Combined with **Bali arrivals turning negative in 2026 (−1.11% y/y Jan–Apr against Indonesia's +7.7%)**, a channel that is growing 194% while the destination shrinks deserves attention that its current 1–3% traffic share would not otherwise justify.

### 4b.6.2 🔴 The structural problem with Pool A content — and why AEO and B2B point the same way

**Pool A listings are the same text and the same images as every other GlobalTix reseller.** That is what a non-exclusive aggregator feed means. A page whose description is substantially identical to dozens of other resellers':

- cannot win organic search on the attraction's head term (the attraction's own site, Klook, GYG and Viator all rank above it, with more authority and the same or better content);
- gives an answer engine no reason to cite it rather than any other copy — there is no distinguishing information on the page;
- and in the duplicate-content case, may not be indexed as a distinct document at all.

**Pool B is the inverse: exclusive, direct-contracted, and held by no other reseller.** For a set of long-tail Balinese experiences, SatuSatu is potentially the **only** structured source on the open web.

> 🔴 **This is the same conclusion §4b.4.4 reached from availability and the brief reached from B2B economics, arrived at from content. AEO, B2B revenue and defensibility all point at Pool B.** Pool A buys coverage and search-success inside SatuSatu's own site — which is real value — but it will not produce organic discovery, will not produce AEO citations, and does not carry B2B margin. **Three independent analyses, one answer. Stop treating Pool A as a growth asset and start treating it as inventory completeness.**

### 4b.6.3 What is known to work vs what is speculation — separated, as required

**Known to work (mechanically verifiable, not opinion):**

| Item | Why it is not speculation |
|---|---|
| **schema.org structured markup** on every SKU | The WDC Product Data Corpus exists **because 79,000 e-shops annotate offers with schema.org and machines harvest it at scale** [VERIFIED, Tier A]. Machine-readable product markup is demonstrably how automated consumers ingest catalogs. This is the highest-confidence AEO action available. |
| **Unique substantive text per SKU** | Duplicate-content handling in search is long-established. It is also the direct remedy for §4b.6.2. |
| **Crawl access for AI crawlers** — an explicit, deliberate robots policy | It is a binary gate. If they cannot fetch it, nothing else matters. **Check the current robots configuration before doing anything else in this subsection.** |
| **Raw MT into more languages for retrievability** | Cost is ~$0.0005/SKU/language (§4b.2.4). At that price the only argument against breadth is quality, and quality matters far less for *being retrieved* than for *converting*. |
| **Being a callable surface inside an assistant** | 🔴 **This is what an incumbent actually did.** Viator's agentic exposure is *"entirely outbound — being a callable app inside ChatGPT"* while its own partner API remains pre-LLM [VERIFIED, Tier A, `teardown-viator-kkday-headout.md`]. The incumbent play is **distribution into the assistant**, not prose optimisation. |

**Speculation — say so, do not budget against it:**

- `llms.txt` and similar proposed conventions: no evidence of retrieval effect in any collected source.
- Any claim about *how* an assistant selects or ranks sources. **No assistant publishes its retrieval criteria. Every "AEO tactic" list is inference dressed as method.** Tier C by construction.
- "AI content at scale improves AEO." Unverified, and §4b.6.2 suggests the opposite for non-exclusive inventory: more copies of the same feed text is more duplicate content, not more retrievability.

**Priority order for SatuSatu, effort-weighted:** (1) robots/crawl audit — hours; (2) schema.org markup on all SKUs — days; (3) unique text on **Pool B** SKUs only, AI-drafted from operator source material per §4b.2.1 — content-ops weeks; (4) raw MT breadth — days; (5) explore assistant-callable distribution — commercial, not engineering; (6) do **nothing** about Pool A prose, it cannot win.

---

## 4b.7 Benchmark summary table — what Step 5 needs

**Every row carries low/base/high and a tier.** `[V]` = VERIFIED, `[I]` = INFERENCE. **All USD figures use §4a's inputs: $0.079/productive minute base ($0.048–$0.138), IDR 15,500/USD — both FX- and wage-sensitive.**

### The seven the brief asked for

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| **S1** | **Cost per SKU onboarded manually** (Pool B: sourcing → contract → rate sheet → content → media → QA → publish; 1.5 / 4 / 12 hours of blended content-ops + commercial time) | **$5** | **$22** | **$95** | `[I]` | — (derived from §4a Tier A/C labour) |
| **S2a** | **Cost per SKU ingested via feed — marginal**, once the feed is live | **$0.00** | **$0.02** | **$0.20** | `[I]` | — |
| **S2b** | **Cost per SKU ingested via feed — amortised fixed cost**, second feed, in eng-weeks/SKU | **0.0002** (25k SKUs, MVP scope) | **0.008** (2,500 SKUs, 20 eng-wk) | **0.08** (250 SKUs, 20 eng-wk) | `[I]` | — |
| **S2c** | **Second-feed total build** — proper / minimum-viable | 11 / 3 | **20 / 5** | 37 / 8 **eng-weeks** | `[I]` | — |
| **S3a** | **Cost per SKU maintained per year — Pool A** (feed-maintained: delta monitoring + sampled QA + exception handling) | **$0.30** | **$1.00** | **$3.00** | `[I]` | — |
| **S3b** | **Cost per SKU maintained per year — Pool B** (manual: 1–3 touches/yr for price, season, blackout, contact churn) | **$3** | **$10** | **$30** | `[I]` | — |
| **S4a** | **SKUs live per content-ops FTE per week — manual, unaided** (35 productive hrs ÷ S1 hours) | **5** | **9** | **15** | `[I]` | — |
| **S4b** | **SKUs live per content-ops FTE per week — with AI structured extraction** (GYG pattern; authoring step only, Amdahl-limited) | **8** | **14** | **22** | `[I]` | — |
| **S4c** | **SKUs live per content-ops FTE per week — feed-ingested with a Layer-1 gate + sampled QA** | **150** | **400** | **1,000** | `[I]` | — |
| **S5a** | **Auto-pass rate — clean aggregator feed, Layer 1 only** | 82% | **90%** | 97% | `[I]` | 🔴 no measured source exists |
| **S5b** | **Auto-pass rate — clean feed, Layers 1+2** | 65% | **78%** | 88% | `[I]` | 🔴 same |
| **S5c** | **Auto-pass rate — AI-generated listing from operator source material, Layers 1+2** | 45% | **62%** | 78% | `[I]` | 🔴 same |
| **S5d** | **Auto-pass rate — bulk import of unnormalised operator spreadsheet, Layers 1+2** | 20% | **40%** | 60% | `[I]` | 🔴 same |
| **S6a** | **Translation cost per SKU per language — raw LLM MT, no human touch** | **$0.0003** | **$0.0005** | **$0.009** | `[I]` from `[V]` Tier C per-Mword rates | C |
| **S6b** | **Translation cost per SKU per language — post-edited by an Indonesia-based FTE** (5 / 8 / 20 min) | **$0.24** | **$0.63** | **$2.76** | `[I]` | — |
| **S6c** | **Translation cost per SKU per language — Western MTPE rates** (250/400/700 words × $0.04–0.12/word) | **$10** | **$32** | **$84** | `[V]` rates, `[I]` per-SKU | C |
| **S6d** | Translation cost per SKU per language — full human | $18 | **$44** | $98 | `[I]` | C |
| **S7** | **% of bulk-ingested catalog unsellable** (dead 3/8/18 + mispriced 2/6/14 + duplicated 0/5/15 + season-expired 2/5/12 + content-unsellable 3/8/16, de-overlapped) | **10%** | **25%** | **45%** | `[I]` | 🔴 no incidence data published anywhere |

### Feed ingestion and matching

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| F1 | Product matching, **deep learning, e-commerce academic benchmark** (WDC LSPM, xlarge training set) | — | **F1 0.90** | — | `[V]` | **A** |
| F2 | Same benchmark — **symbolic/rules matchers**, gap to deep learning | — | **−16 F1 points** | — | `[V]` | **A** |
| F3 | WDC corpus scale / gold standard | 26M offers, 79k e-shops | — | 4,400 verified pairs | `[V]` | **A** |
| F4 | ⚠️ **Attraction/experience matching** — no GTIN, no shared identifier, no labelled pairs. **F1 0.90 does not transfer.** | — | — | — | — | — |
| F5 | Rules on title+geo+duration+price — **precision** | 0.85 | **0.92** | 0.96 | `[I]` | — |
| F6 | Rules only — **recall** | 0.45 | **0.60** | 0.75 | `[I]` | — |
| F7 | Embeddings + rules + human review of uncertain band — **precision on auto-decided pairs** | 0.93 | **0.96** | 0.98 | `[I]` | — |
| F8 | Same — **recall** | 0.60 | **0.75** | 0.88 | `[I]` | — |
| F9 | Exact-identifier join — precision / recall | 0.98 / 0.05 | **0.99 / 0.15** | 1.00 / 0.35 | `[I]` | — |
| F10 | 🔴 **Published attraction-dedup accuracy from Klook / GYG / Viator / Expedia** | — | **NONE — searched; GYG has ~40 eng posts and none on entity resolution** | — | unverified | A (absence in a populated corpus) |
| F11 | Viator: product mapping method | — | **manual, Viator-side, behind certification** | — | `[V]` | **A** |
| F12 | GYG connectivity partners / suppliers / activities / cities / languages | — | **250+ / 50K+ / 200K+ / 12K+ / 37** | — | `[V]` | A (spec) / C (counts) |
| F13 | GlobalTix delta window / availability horizon / hold / token life | 24h | **7d (max 30d)** / 6 months / 15 min / 24h | 30d | `[V]` | **A** |
| F14 | Viator rate limit / concurrency behaviour | — | **16 req/10s rolling; system-wide 503 possible; 120s partner timeout required** | — | `[V]` | **A** |

### Listing production, search, availability

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| G1 | **GYG AI listing creation: wizard steps auto-completed** | — | **8 of 16** | — | `[V]` | **A** |
| G2 | **GYG listing creation time, before → after** | — | **~60 min → 14 min (~4.3×)** | — | `[V]` | **A** |
| G3 | GYG AI-tool adoption in pre-tests | — | **~60%** of exposed cohort | — | `[V]` | **A** |
| G4 | GYG first full experiment outcome | — | 🔴 **FAILED** — submission rate fell; cause **UX/trust, not model quality** | — | `[V]` | **A** |
| G5 | GYG second experiment, drop-off improvement | — | **−5 percentage points**, then 100% rollout | — | `[V]` | **A** |
| G6 | GYG human approval gate retained at 200K SKUs | — | **Yes** — expert review + restricted-activity list | `[V]` | C (GYG's own) |
| G7 | **Semantic search pay-off threshold, SKU count** | **1,200** | **~2,000** | **4,000** | `[I]` | — |
| G8 | Zero-result rate trigger to act (better metric than SKU count) | — | **>10%** on non-junk queries | — | `[I]` | — |
| G9 | Null-click rate trigger | — | **>40%** | — | `[I]` | — |
| G10 | Own-booking events per SKU needed before per-SKU learned ranking | **30** | **50** | **100+** | `[I]` | — |
| G11 | 🔴 **Pool A displayed sold-counts that are SatuSatu transactions** | — | **0% — the field is 100% inherited feed data** | — | `[V]` | A (live product) |
| G12 | 🔴 **Gross-margin loss per unit of GMV diverted Pool B → Pool A** | — | **~17 pts gross / ~20 pts net** | — | `[I]` | — |
| G13 | **Tour operators worldwide with no booking system** | — | **39%** | — | `[V]` | C reporting **B** (Arival *State of Booking Tech* 2025) — ⚠️ not read at source |
| G14 | Industry revenue running on spreadsheets/WhatsApp/manual entry | — | **~$271bn/yr** | — | `[V]` | same caveat |
| G15 | WhatsApp reply → structured state extraction, **clean single-language text** | 85% | **92%** | 97% | `[I]` | — |
| G16 | Same, **real Bali conditions** (code-switching, voice notes, photos) | 60% | **75%** | 88% | `[I]` | — |
| G17 | Required precision on **auto-confirm** (asymmetric cost ~100:1) | — | **≥99%**, accepting 50–70% automation | — | `[I]` | — |
| G18 | Pool B supplier-booking labour per booking (from §4a) | $0.40 | **$1.19** | $5.52 | `[I]` | — |
| G19 | Observed departure-days needed before availability prediction is even arguable | — | **100+ per SKU** | — | `[I]` | — |

### Leakage, settlement, content

| # | Metric | Low | Base | High | Label | Tier |
|---|---|---|---|---|---|---|
| H1 | **Chargeback rate, travel** | **0.89%** | ~1.2% | **2.0%** | `[V]` | C |
| H2 | Card-network high-risk threshold | — | **1.0%** | — | `[V]` | C |
| H3 | Travel disputes YoY | — | **+30%**, friendly fraud named as main driver | — | `[V]` | C |
| H4 | 🔴 **Clean bookings destroyed by one chargeback — Pool A / Pool B** | — | **~9 / ~2.7** (≈13 on Pool A's 7% net) | — | `[I]` | — |
| H5 | 🔴 **Share of pool gross profit lost at a 0.89% / 2.0% chargeback rate — Pool A** | 8% | — | 18% (12% / 27% on net) | `[I]` | — |
| H6 | Same — Pool B | 2.4% | — | 5.4% | `[I]` | — |
| H7 | **No-show rate, tours & activities** | — | 🔴 **NOT FOUND — no published TAA benchmark** | — | unverified | — |
| H8 | Cart abandonment, travel | — | **~81%** | — | `[V]` | C |
| H9 | Operator gross margin by tour type (day / multi-day / luxury / budget) | 15–20% | **40–50% day tours**; 25–35% multi-day | 50%+ luxury | `[V]` | C |
| H10 | **GlobalTix FX markup, non-SGD settlement** | 0% (SGD) | **1.50%** | — (0.50% AED outlier) | `[V]` | **A** |
| H11 | GlobalTix `creditCardFee` | — | **3.00** | — | `[V]` | **A** |
| H12 | FX markup as a share of Pool A margin | — | **~15% of gross / ~21% of net** | — | `[I]` | — |
| H13 | IDR excluded as a settlement/invoicing currency | — | **GYG checkout: excluded. Viator merchant: excluded (5 currencies vs 27 for affiliates). GlobalTix: included.** | — | `[V]` | **A** |
| H14 | **Pool B reconcilability** | — | 🔴 **Zero — no supplier-side booking ID, no confirmation state, no redemption event** | — | `[V]` | A (absence) |
| H15 | **AI-source traffic to US travel sites, y/y (May 2026)** | — | **+194%** (+2,215% since Oct 2024) | — | `[V]` | **B** |
| H16 | AI-referred travel traffic conversion vs non-AI | — | **−28%**, gap ~70% narrower than Oct 2024 | — | `[V]` | **B** |
| H17 | AI-referred engagement / visit duration / bounce | +21% | **+70% duration** | −41% bounce | `[V]` | **B** |
| H18 | AI search as share of total site traffic | 0.5% | **1–3%** | 5–8% (tech/e-comm) | `[V]` | C |
| H19 | Share of AI referrals from ChatGPT | — | ⚠️ **contradictory: 92.4% vs "mid-sixties and falling"** | — | `[V]` both | C — report the conflict |
| H20 | MTPE / human translation per word | $0.04 | **$0.06–0.12** | $0.14 | `[V]` | C |
| H21 | LLM MT per million words (cheap model / frontier batch, Apr 2026) | $1.05 | — | $13.33 | `[V]` | C |
| H22 | COMET quality reference (professional / near-human) | — | **≥0.85 / ≥0.90** | — | `[V]` | C |

---

## 4b.8 Open questions this step could not close

Ordered by how much they change the answer.

| # | Question | Why it matters | Where it must come from |
|---|---|---|---|
| **1** | 🔴 **Does the GlobalTix feed grant include the right to sublicense images and content to SatuSatu's distribution partners?** | If not, the D2/D3 platform has a content-rights problem underneath it that engineering cannot fix. One written question to GlobalTix closes it. | Supplier contract / GlobalTix, in writing. |
| **2** | 🔴 **Does the prospective second aggregator overlap GlobalTix on the head attractions — and can it be contracted for non-overlapping inventory only?** | Overlap is the entire deduplication problem. **No overlap = the whole §4b.1 stack disappears and 11–37 eng-weeks vanish.** This is a contracting decision, not a technical one. | Commercial — the second aggregator's catalog list vs GlobalTix's. |
| **3** | 🔴 **Do the two feeds share any venue/attraction identifier?** | An exact-identifier join is 0.99 precision and would solve the head of the catalog outright (F9). Cheap to check, transformative if true. | Both API responses, side by side. One afternoon. |
| **4** | 🔴 **Which Pool B SKUs can take `freesale_capped` / `allotment` / `static_schedule`, and which genuinely need `request_to_book`?** | This classification **is** the deliverable of §4b.4 and the gating dependency for B2B, reconciliation and pricing. It is internal knowledge nobody outside can supply. | Internal — supply/ops, per operator conversation. Days. |
| **5** | 🔴 **Every surface that currently reads the inherited Pool A sold-count.** | It is silently steering mix from 27% margin to 10% margin (G12). Assume it leaked into more places than anyone remembers. | Internal — code and CMS audit. |
| **6** | **SatuSatu's own chargeback, refund and dispute rates.** | H5 says an industry-average rate consumes ~8–18% of Pool A gross profit. The actual rate changes whether this is a priority or a footnote. | Internal — PSP dashboards. One report. |
| **7** | **Is SGD settlement contractually available from GlobalTix?** | Worth ~15% of Pool A gross margin (H12) for a contract amendment. | Supplier negotiation. |
| **8** | **Current robots/crawler configuration — are AI crawlers permitted?** | A binary gate on everything in §4b.6. Nothing else in AEO matters if the answer is no. | Internal — hours. |
| **9** | **Arival *State of Booking Tech* (2025) at source** — the 39% and $271bn figures. | Both are load-bearing for §4b.4 and both were read via a third-party blog, not at source. | Arival — paid report. |
| **10** | **The second aggregator's take rate, and any channel-manager take rate if the buy-not-build route is used.** | A second margin taker on ~10% gross may be fatal. This determines whether the cheapest engineering path is commercially available at all. | Commercial. |
| **11** | **A measured auto-pass rate or escaped-defect rate from any travel catalog pipeline.** | Every band in S5 is inference. Nobody publishes this. | Not found. May require a vendor POC with instrumentation written into the contract. |
| **12** | **A no-show rate for tours and activities.** | Searched directly; no public benchmark exists. Matters most for Pool B, where SatuSatu has paid the operator. | Internal, once Pool B has a booking record at all — i.e. after #4. |
| **13** | **Loaded cost of an engineering week at SatuSatu.** | §4b.1.6 is denominated in eng-weeks precisely because this is unknown *and* because eng-weeks are the binding constraint. Needed to compare buy-vs-build in one currency. | Internal — finance. |
| **14** | **Whether any B2B partner has actually asked for a searchable/agentic API.** | No incumbent exposes semantic search or an agent surface in B2B (Klook, GYG, Viator all confirmed absent, Tier A). Either an opportunity or evidence of no demand. | Commercial — ask the D2/D3 pipeline. |

---

*End of section 4b. Section 4c to be appended below.*
---

# 4c — B2B Partner Operations

## 4c.0 Reading order, and the one finding that reorders the bracket

**Shared assumptions used in every cost figure below.** Stated once so they can be corrected once.

| Input | low | base | high | Basis |
|---|---|---|---|---|
| FX, IDR/USD | 16,000 | 16,000 | 16,000 | [INFERENCE] Planning rate. §4a.5.2 used 15,500 for one conversion; the 3% difference does not move any conclusion here. **Not a forecast.** |
| Fully-loaded multiplier on base salary | ×1.50 | ×1.68 | ×1.92 | **Reused verbatim from §4a.5.3.** Do not re-derive. |
| BD / partnerships person, monthly base | Rp 10.0m | Rp 14.0m | Rp 18.0m | [INFERENCE] Set above §4a.5.2's concierge band (Rp 4.5/6.5/9.5m) because the role carries negotiation and commercial judgement. **Unverified — no Indonesian BD salary source was collected.** |
| → BD fully-loaded, per productive hour | **US$6** | **US$8** | **US$11** | Derived: (base × multiplier) ÷ 176 h/mo. |
| Integration-capable engineer, monthly base | Rp 18.0m | Rp 28.0m | Rp 40.0m | [INFERENCE] **Unverified — no Indonesian engineering salary source was collected. This is the weakest input in 4c and it drives the break-even.** Flag 4c-G1. |
| → Engineer fully-loaded, per eng-week | **US$390** | **US$679** | **US$1,109** | Derived: (base × multiplier) ÷ 4.33 wk/mo. |
| Support agent fully-loaded, per minute | US$0.052 | US$0.065 | US$0.109 | Derived from §4a.5.2/4a.5.3 concierge band ÷ (22 d × 8 h × 60 min). |

> ⚠️ **The engineer figure is payroll, and payroll is the wrong price.** The brief states there is **no dedicated engineering capacity** — the D2/D3 build already consumes the team. So the true cost of an eng-week spent on partner integration support is not US$679, it is **the roadmap slip it causes**. Every eng-week number below is therefore a *floor*, and the break-even derived from it is **optimistic by construction**. Read the break-even as "no better than this."

### 4c.0.1 🔴 The finding that reorders the bracket: Pool B cannot pass an industry-standard API certification

Viator's published certification requirements [VERIFIED, Tier A — partnerresources.viator.com/travel-commerce/certification] gate production access on, among ~60 discrete checks:

- "Real-time availability and pricing checks"
- "Booking hold" as a distinct endpoint operation
- `/availability/schedules/modified-since` availability ingestion
- A minimum **120-second timeout** tolerance
- "Calendar view with available dates"

**Pool B has no real-time availability.** That is not a gap against a nice-to-have; it is a gap against the *gating* requirement of the only certification regime in this category we could document. The consequence:

> **D3 built on Pool B cannot be a real-time booking API. It can only be a request-to-book (on-request) API.** And a request-to-book API, from an agent's seat, is a slower, more brittle version of the WhatsApp message they already send. **The API adds latency to a workflow whose only defect is latency.**

This single fact propagates through the whole bracket: it is the dominant cost driver in §4c.2, it is why the D3 break-even in §4c.3 is fragile, it is the reason the AI-quoting assistant in §4c.4 can manufacture unhonourable quotes, and it is the strongest argument in §4c.7 for riding an existing rail instead of building D3.

**Restated as a dependency, not an opinion:** closing Pool B's availability gap is worth more than building D3. Doing D3 first spends the scarce resource on the wrong layer.

### 4c.0.2 The rezio read — directionally right, but mis-specified for D2

The established finding is that a partner dashboard is a supply-acquisition subsidy, not a revenue line (rezio: ~US$50–217/mo, no commission, ceiling ~US$6M ARR on 5,000+ operators). **Tested against 4c's evidence, the framing holds for supplier backends and is the wrong shape for D2.**

Corroborating evidence found here:
- **Viator's API certification carries no stated fee** [VERIFIED, Tier A]. The distributor absorbs the entire onboarding cost of its own integrators. Consistent with "not a revenue line."
- **Holibob ships a white-label experience website and a connectivity-manager function as channel enablement**, not as a subscription product [VERIFIED, Tier A — Holibob's own IATA AirTechZone submission and partner.documentation.holibob.tech]. Again: distribution plumbing given away to move volume.

But the correction matters:

| | rezio (supply-side) | D2 (demand-side) |
|---|---|---|
| Who logs in | The operator who *owns* the inventory | The agent who *resells* the inventory |
| What the login is for | Manage own product, distribute outward | Search, quote, book someone else's product |
| Where money comes from | Subscription (thin) | **Margin on the bookings it moves** |
| Correct P&L treatment | Product with its own revenue line | **Cost of goods sold on a distribution channel** |

> **The corrected read: D2 has no revenue line of its own and should never be given one. It is a channel, and its only KPI is retained margin on Pool B volume, net of the cost of running the channel.** Any plan that scores D2 on partner count, seats, or subscription revenue is measuring the wrong thing — that is precisely the trap rezio's US$6M ceiling illustrates. Do not charge partners for D2 access; a subscription would suppress the volume that is the actual product.

### 4c.0.3 The margin that is actually available to D2/D3

Pool B is ~27% gross, ~24% after payments and FX (applying Pool A's ~3pp payments+FX drag). **But B2B means giving the partner a real net rate — that spread is the offer.** SatuSatu retains only the remainder.

| | low | base | high |
|---|---|---|---|
| Pool B gross margin | 27% | 27% | 27% | 
| less payments + FX | ~3pp | ~3pp | ~3pp |
| less spread conceded to partner (volume-tiered) | 16pp | 12pp | 8pp |
| **= SatuSatu retained net margin on B2B Pool B** | **8%** | **12%** | **16%** |

[INFERENCE] throughout. The conceded-spread band is the planning variable, not an observation — **no SatuSatu net-rate card was available.** Flag 4c-G2. Every contribution figure in 4c scales linearly off this row, so it is the highest-leverage number to replace with a real one.

---

## 4c.1 Partner acquisition and onboarding economics

### 4c.1.1 Cost and time per partner, step-decomposed

Bands are [INFERENCE] built from BD hours × the US$6/8/11 loaded hourly rate in §4c.0. The hour counts are not sourced — **no TAA-specific partner-onboarding time study was found** (flag 4c-G3). The one external anchor: Apiable states that generic API partner onboarding runs "two weeks of elapsed time and involves three people on your team" [Tier C — vendor, and it is their pre-improvement straw man, so treat as an upper anchor only].

| Step | low (h) | base (h) | high (h) | What drives the high end |
|---|---|---|---|---|
| Source + qualify | 2 | 5 | 12 | Cold outbound to an agency with no Bali volume |
| Commercial negotiation + net-rate agreement | 1 | 3 | 8 | Partner counter-proposes tiers, exclusivity, or marketing support |
| KYC / entity verification (NIB, NPWP, sector tourism licence, bank account) | 0.5 | 1.5 | 4 | Scanned documents, mismatched legal names, expired licence |
| Activation + training (portal walkthrough, first assisted booking) | 2 | 5 | 12 | Non-technical staff, multiple seats, WhatsApp hand-holding |
| **Subtotal — prepay partner** | **5.5** | **14.5** | **36** | |
| **Cost @ US$6 / 8 / 11 per hour** | **US$33** | **US$116** | **US$396** | |
| *Add if credit terms are extended* | *+3* | *+8* | *+20* | *Financial DD, references, limit setting, guarantor* |
| **Cost — credit partner** | **US$51** | **US$180** | **US$616** | |

**Calendar time per partner** [INFERENCE]: **1–2 weeks** (prepay, motivated, self-serve-capable) / **3–6 weeks** (base) / **3–4 months** (credit underwriting + sector-licence verification + the partner's own procurement).

> 🔴 **The structural finding: the expensive part of B2B onboarding is not account creation, it is credit underwriting.** It roughly doubles the labour cost and multiplies calendar time by 3–8×. Given no disclosed funding since Nov 2022, credit is also a working-capital exposure (§4c.6). **Launching D2 prepay-wallet-only removes the entire expensive half of onboarding and the entire AR function at once.** This is the cheapest single decision in 4c and it requires no engineering beyond a wallet balance.

### 4c.1.2 ⚠️ Testing the hypothesis: is D2 a high-volume, low-ACV motion?

**Verdict: yes for D2, and the same claim is false for D3.** The two must not share a motion.

First, partner annual value. Built from bookings × ticket price × retained margin (§4c.0.3):

| | low partner | base partner | high partner |
|---|---|---|---|
| Bali attraction bookings/yr through SatuSatu | 200 | 800 | 3,000 |
| Avg ticket, USD | 20 | 25 | 30 |
| Partner GMV/yr | US$4,000 | US$20,000 | US$90,000 |
| × retained net margin (8 / 12 / 16%) | US$320 | **US$2,400** | US$14,400 |

[INFERENCE] on every row. **Unverified — no SatuSatu partner-volume data and no Indonesian agent Bali-attraction spend distribution were found** (flag 4c-G4). Sanity anchor: GlobalTix reports 12,000+ agents and 25M tickets/yr [VERIFIED, Tier B, established elsewhere in this study] → ~2,080 tickets/agent/yr *on average*, but the Arival finding that **>70% of experience operators are small or micro** applies symmetrically to the agent population: this is a power law, the mean is not the mode, and the modal agent sits at or below the "low partner" column.

Now the test. Two things decide whether human-assisted sales is affordable:

1. **Payback on base-case acquisition cost (US$116) against base-case first-year contribution.** D2 contribution before any integration cost is US$2,400 − US$784 variable support (§4c.2) = **US$1,616**. Payback: **~26 days of the partner's first year.** Affordable by a wide margin.
2. **The same test on the modal *small* partner.** Contribution US$320 − US$174 = **US$146/yr**. Acquisition US$33–116. Payback **3 months to just under a year.** Still positive — thin, but positive.

> 🔴 **The counter-intuitive result, and it is the most actionable finding in §4c.1: Indonesian labour arbitrage makes human-assisted B2B onboarding affordable at an ACV where SaaS orthodoxy says it cannot be.** At US$6–11 per loaded BD hour, a 10-hour hand-held onboarding costs US$60–110 — under a year's contribution from even the *smallest* viable partner, and under 7% of the base partner's. The standard advice ("low ACV forces pure self-serve") is a conclusion drawn from US/EU salary structures and **does not transfer.** SatuSatu should staff human onboarding and should not build self-serve tooling to avoid it.

**Where this breaks:** the same arithmetic applied to D3 gives the opposite answer, because D3 adds eng-weeks, and engineering is both 85–170× more expensive per hour than BD *and* the constrained resource. See §4c.3.4.

### 4c.1.3 What AI actually compresses in onboarding — and why it is not worth building

Attacking each step with the most favourable realistic assumption:

| Step | AI mechanism | Realistic compression | Hours saved (base) |
|---|---|---|---|
| Source + qualify | Enrichment / ICP scoring | 10–25% | 0.5–1.3 |
| Negotiation + contracting | Template assembly, clause redlining | 30–50% | 0.9–1.5 |
| KYC / entity verification | Document extraction from NIB/NPWP/licence scans | 40–70% | 0.6–1.1 |
| Activation + training | In-portal AI assistant answering setup questions | 20–40% | 1.0–2.0 |
| **Total** | | **~20–35%** | **3.0–4.9 h** |

[INFERENCE] on every compression figure. The activation row is deliberately conservative: **§4a.2's definitional critique of deflection applies unchanged here** — an AI onboarding assistant deflects *questions*, it does not deflect the *relationship*, and a new B2B partner asking "will this actually work for my clients" is not asking a knowledge-base question.

**The arithmetic that settles it:** 3.0–4.9 hours × US$8 = **US$24–39 saved per partner.** At 100 partners that is **US$2,400–3,900 per year** — roughly 4–10 eng-weeks of payback at US$390–1,109/week, i.e. the tooling costs more than it saves for years, and it is charged against the constrained resource.

> **Verdict: do not build AI onboarding tooling.** Take whatever ships inside the CRM/e-sign stack already in use and stop there. **The labour arbitrage that makes human onboarding affordable is the same fact that makes automating it uneconomic.** The AI budget belongs in §4c.2's availability loop, which is 20–75× larger and grows with volume.

---

## 4c.2 Partner support economics

### 4c.2.1 Why the B2B queue is a different animal — and where the cost actually sits

The brief's framing is correct — a broken partner booking damages the *partner's* customer relationship, so the stakes and SLA differ. But the collected evidence says the dominant B2B support cost at SatuSatu is not the incident queue at all.

**Pool B has no real-time availability. Therefore 100% of Pool B bookings require a human confirmation round with the Balinese operator.** This is not a contact *rate* — it is a contact on every single transaction, and it scales linearly with volume forever. It never amortises.

| Cost line | low | base | high | Derivation |
|---|---|---|---|---|
| Agent minutes per on-request confirmation (message operator, await reply, notify partner) | 6 | 10 | 15 | [INFERENCE], unverified — flag 4c-G5 |
| × loaded per-minute rate (§4c.0) | US$0.052 | US$0.065 | US$0.109 | |
| **= cost per Pool B booking** | **US$0.39** | **US$0.65** | **US$1.64** | |
| × bookings/partner/yr (§4c.1.2) | 200 | 800 | 3,000 | |
| **= availability confirmation, per partner per year** | **US$78** | **US$520** | **US$4,920** | |

> 🔴 **At the high-volume end this consumes ~34% of the US$14,400 retained margin — and it is caused entirely by a missing availability feed, not by support quality.** This is the quantification of "Pool B's availability gap is the gating dependency." A partner who succeeds makes the problem worse, which is the signature of a broken unit economic.

### 4c.2.2 The exception/incident queue, and the benchmark that does not transfer

External cost-per-ticket benchmarks, all **Tier C (vendor content, no disclosed methodology)** and all US/EU-cost-structured:

| Source | Claim | Use |
|---|---|---|
| Lorikeet | "$2.70 for simple retail interactions to $60 or more for complex B2B support cases" | Directionally supports the B2B premium (~10–20×). **Do not use the absolute values.** |
| Tandem | "average ticket costs $25 to $35 to resolve at the growth stage" | US-cost anchor |
| Higher Logic | "a deflected ticket saves $15–$20" | US-cost anchor |
| Kustomer | AI gives "23% to 28% reduction in overall support" cost | Vendor claim, no methodology. **Reject as a planning input.** |

**Recomputed on SatuSatu's cost base** rather than imported: a B2B exception (booking amendment, pax-mix change, no-show dispute, operator failure) runs 3–5× a simple consumer contact because it involves three parties. At 20/40/75 agent-minutes: **US$1.04 / US$2.60 / US$8.18 per B2B exception** — i.e. **7–58× cheaper than the US benchmark.** Same conclusion as §4c.1.2: the labour arbitrage is the advantage, and it lowers the value of deflection automation rather than raising it.

| | low | base | high |
|---|---|---|---|
| Exception contacts per 100 bookings | 3 | 5.5 | 8 |
| Exceptions/partner/yr | 6 | 44 | 240 |
| Cost/exception | US$1.04 | US$2.60 | US$8.18 |
| **= exception queue, per partner per year** | **US$6** | **US$114** | **US$1,963** |

### 4c.2.3 Cost per partner served per year — and deflection separated from assist

| Component | low | base | high |
|---|---|---|---|
| Availability confirmation (§4c.2.1) | US$78 | US$520 | US$4,920 |
| Exception queue (§4c.2.2) | US$6 | US$114 | US$1,963 |
| Account management / QBR / re-training | US$18 | US$150 | US$660 |
| **Total support cost per partner per year (D2, no API)** | **US$102** | **US$784** | **US$7,543** |
| *Add ongoing API maintenance if integrated via D3* | *US$98* | *US$340* | *US$1,109* |
| **Total per partner per year (D3-integrated)** | **US$200** | **US$1,124** | **US$8,652** |

**Deflection vs assist, held strictly separate** (per §4a.2's definitions):

| Class | Mechanism | Realistic band | Note |
|---|---|---|---|
| **True deflection** — contact never reaches a human | Self-serve booking status, voucher re-issue, rate-card lookup, invoice download, cancellation-policy lookup | **25–45%** of *exception* contacts | [INFERENCE]. Agents are repeat, trained users — the single most deflectable audience there is, far more than consumers. |
| **True deflection on the availability loop** | **~0%** until Pool B has a feed | **0%** | 🔴 **No amount of AI deflects a question whose answer does not exist in any system.** This is the whole point. |
| **Assist** — human still handles it, faster | Drafted replies, auto-pulled booking context, operator-reply parsing, translation | **20–40%** handling-time reduction | [INFERENCE]. Materially larger prize than deflection here, because the availability loop is un-deflectable but *is* compressible. |

> **The one AI investment §4c.2 justifies: automate the operator-confirmation loop as an assist, not a deflection.** Send the operator a structured WhatsApp request, parse the free-text reply into a confirm/decline/alternative, and only escalate the ambiguous ones. A 20–40% handling-time cut on the availability line saves **US$16–2,000 per partner per year**, is the only cost line that grows with success, and — critically — it is *pre-confirmation*, so a mistake produces a re-ask rather than a wrong booking (contrast §4a.3's mutation blast radius).

---

---

## 4c.3 🔴 Integration support cost for D3, and the break-even

### 4c.3.1 ⚠️ What Viator's certification gates, and what it costs both sides

Viator runs a formal certification before production access. This is the best-documented precedent available and it is **[VERIFIED, Tier A]** — Viator's own partner resource centre and technical docs.

**Mechanics of the gate:**
- Two forms — a **back-end** certification form and a **front-end** certification form — submitted to `affiliateapi@tripadvisor.com`.
- "Your product ingestion/update strategy is one of our certification requirements and must be verified by us prior to your accessing the production server."
- "Allow a couple of days to receive your initial round of certification feedback, **depending on our current volume of demand**."
- A **test booking** and a **test cancellation** must be completed and verified.

**What it gates — nine categories, ~60 discrete checks:**

| Category | Representative gated items |
|---|---|
| Endpoint usage rules | `/products/modified-since`, `/availability/schedules/modified-since`, real-time search, real-time availability & pricing, **booking hold**, booking create/status/cancel, **minimum 120-second timeout** |
| User flow & navigation | search by destination/attraction/product code/freetext; filters on category, price, rating, duration, time of day, policies; sort order; breadcrumbs |
| Product details | title, images, description, inclusions, exclusions, additionalInfo, cancellationPolicy, languageGuides, itinerary, ticketInfo, logistics.start, logistics.end |
| **Age bands / passenger mix** | correct bucketing with `startAge`/`endAge` displayed; traveller-count limits enforced **per age band and per booking** |
| Availability & pricing | calendar view of available dates; per-person vs unit pricing; real-time verification |
| Product options | option title/description, start times, language-guide and pickup verification at option level |
| Booking process | booker identity, per-traveller booking questions, pickup, cancellation-policy disclosure, voucher delivery, **HTTPS**, **PCI DSS compliance** |
| Cancellation | refund amount disclosed *before* cancellation; test cancellation |
| Miscellaneous | correct Viator branding; **non-indexing of "Protected Viator Unique Content"** |

**Cost to the API provider (the side SatuSatu would be on):** **no fee is charged** [VERIFIED, Tier A — none stated anywhere in the certification documentation]. The provider absorbs it entirely. The phrase *"depending on our current volume of demand"* is itself the tell: **certification is a staffed, capacity-constrained human queue, not an automated gate.** [INFERENCE] 0.5–2 reviewer-days per round, 1–4 rounds per partner.

**Cost to the integrator:** building to ~60 checks *plus* front-end obligations (breadcrumbs, six filter dimensions, sort, branding rules, content non-indexing) *plus* **PCI DSS compliance**. For a small Indonesian agency, PCI DSS is plausibly the single largest hidden gate in the list — larger than the API work.

> 🔴 **The transferable lesson: Viator publishes extensive documentation and a sandbox, and still requires mandatory human certification of every integrator.** The best-resourced, best-documented API in this category **has not** engineered the human review out. Viator can carry a staffed cert queue because it amortises across thousands of partners. **SatuSatu cannot amortise it across twenty.** If SatuSatu ships D3, it chooses between (a) no certification — accepting partners who ship mispriced, stale, parity-breaking storefronts under SatuSatu's supply, or (b) a certification queue it cannot staff. Neither is acceptable.

### 4c.3.2 Cost per integrator, and time-to-first-booking

| | low | base | high |
|---|---|---|---|
| **SatuSatu eng-weeks per integrator** (kickoff, sandbox support, debugging partner-side errors, cert review, go-live watch) | 0.75 | 2.5 | 8.0 |
| **Cost @ US$390 / 679 / 1,109 per eng-week** | **US$293** | **US$1,698** | **US$8,872** |
| Calendar time, contract signed → certified | 2–4 wks | 8–12 wks | 6–9 months |
| **Realistic time-to-first-booking** | **3–6 wks** | **3–4 months** | **9+ months, or never** |
| Ongoing maintenance, per integrator per year | 0.25 ew | 0.5 ew | 1.0 ew |

[INFERENCE] on all rows. External anchors, both **Tier C and both rejected as planning inputs**: Apiable's "two weeks of elapsed time and three people" per partner (their own pre-improvement straw man); an unattributed LinkedIn claim of "six weeks to eight days." **Neither is usable.**

> ⚠️ **Never-launch rate is a real and dominant risk and we could not quantify it.** The industry pattern that a large share of signed API partners never transact is widely asserted and **we found no Tier A or Tier B measurement of it.** Marked unverified — flag 4c-G6. It matters enormously: every never-launched integrator is pure cost with zero contribution, and if the never-launch rate is 40–60% the effective cost per *productive* integrator is 1.7–2.5× the table above.

### 4c.3.3 ⚠️ What sandboxes, docs and AI integration copilots actually save

**Demanded evidence, not found.** Searching specifically for measured integration-time savings from self-serve sandboxes, documentation quality, or AI integration copilots returned **only vendor marketing (Tier C) with no disclosed methodology, no baseline definition, and no independent verification.** Every quantified claim located traced back to an API-management or B2B-SaaS vendor selling the thing being measured — exactly the bias the brief flagged.

**Therefore: no savings figure from this category is admitted into the cost model above.** The eng-week bands in §4c.3.2 are stated *inclusive* of good docs and a sandbox, because those are table stakes rather than a lever.

What *can* be asserted, and it is Tier A because it is observed live-product behaviour rather than a claim:
1. **Viator ships docs + sandbox + a published certification checklist and still mandates human review.** Docs and sandboxes shift *where* the human time is spent — from answering basic questions to adjudicating certification — they do not remove it.
2. The certification list is dominated by **semantic** correctness (age-band bucketing, per-person vs unit pricing, refund disclosure before cancellation, pickup at option level), not syntactic correctness. **Sandboxes catch syntax; certification exists because syntax was never the problem.** An AI copilot generating correct API calls does not make a partner display the right cancellation policy to the right traveller.

> **Planning position: assume sandbox and docs save 0–20% of integration eng-weeks, treat any vendor figure above that as unproven, and do not fund tooling on the strength of it.**

### 4c.3.4 🔴 The break-even

**Annualised fixed cost of running a D3 partner API** (docs site, sandbox with seeded data, key/secret management, versioning and changelog, breaking-change comms, partner-facing status/monitoring, security patching; build amortised over 3 years):

| | low | base | high |
|---|---|---|---|
| Build, amortised (8 / 16 / 30 eng-weeks ÷ 3 yr) | 2.7 ew | 5.3 ew | 10.0 ew |
| Run & maintain, per year | 6.0 ew | 16.0 ew | 40.0 ew |
| **Total, eng-weeks per year** | **8.7** | **21.3** | **50.0** |
| **Annualised fixed cost of D3** | **US$3,393** | **US$14,463** | **US$55,450** |

**Contribution per D3 partner per year** = retained margin (§4c.0.3) − support (§4c.2.3, D3 row) − integration support amortised over 3 yr:

| | low partner | base partner | high partner |
|---|---|---|---|
| Retained net margin | US$320 | US$2,400 | US$14,400 |
| less support (D3-integrated) | (US$200) | (US$1,124) | (US$8,652) |
| less integration support, amortised 3 yr | (US$98) | (US$566) | (US$2,957) |
| **= contribution per partner per year** | **+US$22** | **+US$710** | **+US$2,791** |

**Break-even partner count, N\* = annualised fixed ÷ contribution:**

| Contribution per partner ↓ / Fixed cost → | US$3,393 (low) | **US$14,463 (base)** | US$55,450 (high) |
|---|---|---|---|
| **+US$22** (modal small partner) | 154 partners | **657 partners** | 2,520 partners |
| **+US$710** (base partner) | 5 partners | **🔴 20 partners** | 78 partners |
| **+US$2,791** (large partner) | 2 partners | **5 partners** | 20 partners |

> # 🔴 The number: **~20 actively-transacting D3 partners** at base assumptions — equivalent to **~16,300 Pool B bookings per year through the API.**

**But the partner count is the less useful half of the answer.** Restating it as a *size* threshold is what BD can actually act on. At base assumptions, per Pool B booking: retained margin US$3.00, variable cost US$0.79 (availability confirmation US$0.65 + exception load US$0.14), **contribution margin US$2.21 per booking.** Per-partner D3-specific fixed cost is US$1,056/yr (ongoing API maintenance US$340 + amortised integration US$566 + account management US$150). So:

> # 🔴 **A D3 partner must transact ≥ ~480 Pool B bookings per year (~9 per week) merely to pay for its own integration and maintenance.** Below that threshold the partner is loss-making *before* any share of platform cost. Above ~2,000 bookings/yr, D3 break-even falls to ~4–5 partners.

**Three things this number says that the headline does not:**

1. **D3 is not a high-volume, low-ACV motion — it is the opposite.** The 480-booking floor excludes the modal agent, and the Arival evidence (**>70% of experience operators small or micro**, applied symmetrically to agents) says the modal agent is exactly who would sign up. **D3's viability rests entirely on acquiring the top decile.** Any D3 plan with a partner-count target rather than a partner-volume target is mis-specified.
2. **D2 breaks even ~4× more easily and tolerates small partners.** Same method, D2 annualised fixed = 5/12/24 eng-weeks = **US$1,950 / US$8,148 / US$26,616**; contribution = retained margin − D2 support = **+US$218 / +US$1,616 / +US$6,857**. → N\* = **~5 base-size partners, or ~37 small partners.** 🔴 **D2 pays for itself in a regime where D3 cannot.** Sequence follows directly: **D2 first, D3 only against named partners already exceeding 480 bookings/yr.**
3. **The break-even is optimistic by construction and the direction of error is known.** The engineer rate is payroll, not the shadow price of the constrained resource (§4c.0); never-launch rate is unquantified (4c-G6) and inflates cost per productive integrator by 1.7–2.5× if it is 40–60%; and no certification queue is costed in, because SatuSatu cannot afford one. **A defensible planning position is 25–40 partners, not 20.**

---

---

## 4c.4 B2B search and quoting behaviour

### 4c.4.1 Evidence that the B2B search shape is real, not assumed

The brief's characterisation (multi-pax, date-flexible, itinerary-shaped, margin-aware, quote-before-book) is largely **corroborated by Viator's certification requirements**, which are the strongest available evidence because they are what a major TAA distributor *forces every integrator to implement* [VERIFIED, Tier A].

| Claimed behaviour | Evidence | Verdict |
|---|---|---|
| **Multi-pax** | Viator certifies an entire category on **"Age Bands/Passenger Mix"** — correct bucketing with `startAge`/`endAge` displayed, traveller-count limits enforced **per age band and per booking**; and certifies **per-person vs unit pricing** as a distinct check | **Confirmed, Tier A.** Pax mix is a first-class, certifiable concern — not a UI nicety. |
| **Date-flexible** | Viator certifies a **"Calendar view with available dates"** and availability ingestion via `/availability/schedules/modified-since` | **Confirmed, Tier A.** Date-range browsing is mandatory, not optional. |
| **Quote before booking** | Viator certifies **"Booking hold"** as a distinct endpoint operation, separate from booking creation | **Confirmed, Tier A.** Hold-then-book is a standardised primitive in this category, which is exactly what quote-then-book requires. |
| **Itinerary-shaped** | Viator certifies filters on **duration** and **time of day**, plus `logistics.start` / `logistics.end` and product-option **start times** | **Partially confirmed, Tier A.** The *fields* required to assemble a day plan are all certified. No source shows agents searching in itinerary form. |
| **Margin-aware** | Net-rate contracts carry "strict confidentiality clauses. Disclosing wholesale prices to clients or competitors violates agreements" [Tier C — dmcquote]; TourCMS documents net rates as the alternative to agent commission [Tier C — vendor docs] | **Weakly supported.** The net-rate *mechanism* is documented; agent margin-aware *search behaviour* is not. |
| **Agent platform-selection behaviour generally** | ⚠️ **Nothing.** No Tier A or Tier B study of Indonesian (or SE Asian) travel-agent platform selection or search behaviour was found | **Unverified.** Flag 4c-G7 — a significant gap. |

### 4c.4.2 What this implies for ranking

Viator's certified filter set is the floor: **category, price, rating, duration, time of day, policies**, plus sort order. A B2B surface needs five dimensions Viator's consumer-derived list does not contain:

1. **Confirmation latency as a first-class signal.** Instant-confirm vs on-request, with an expected confirmation window. 🔴 **This is not a nice-to-have — it is the one ranking rule that prevents D2 from damaging partners.** Pool B is high-margin *and* on-request; Pool A is low-margin *and* instant. A margin-optimising ranker will systematically surface on-request inventory to an agent who is on the phone with a client, and the resulting "let me get back to you" is precisely the damage the brief warns about. **Margin-weighted ranking must be gated on latency, not merely annotated with it.**
2. **Net rate and retained margin**, agent-visible for their rate, operator-visible never (the confidentiality clause makes leakage a contract-termination event — §4c.5).
3. **Pax-mix feasibility as a filter, not a post-filter error.** Age-band and per-booking traveller limits are certified concerns; an agent searching for 2 adults + 3 children aged 4/7/11 should never see a product that will reject that mix at checkout.
4. **Cancellation latitude.** Agents' clients change plans; free-cancellation windows are a purchasing criterion in a way they are not for consumers. Viator certifies refund-amount disclosure *before* cancellation — treat cancellation terms as a ranked attribute.
5. **Geographic/temporal cluster compatibility** — so a result can be judged against the day already planned rather than in isolation.

### 4c.4.3 An AI quoting assistant — good fit, with one hard constraint

The genuine job is **itinerary assembly under constraints**: multi-day, multi-pax with age bands, geographically coherent, duration- and start-time-feasible, hitting a target agent margin, output as a branded quote the agent forwards to their client. This is a strong AI fit for reasons that are structural rather than hopeful:

- It is **generative plus constraint-satisfying** — the shape of problem where LLMs plus a solver beat both a human and a pure rules engine.
- It is **pre-booking**, so the blast radius of an error is a re-quote, not a wrong booking. **Contrast §4a.3: this sits on the safe side of the mutation boundary.** It is the best-positioned AI application in this bracket.
- It absorbs the *itinerary-shaped* query that no filter UI expresses well.

> 🔴 **The hard constraint: Pool B has no real-time availability, so every quote it produces is unconfirmable at the moment of quoting.** An assistant that emits confident quotes against on-request inventory manufactures exactly the failure the brief warns about — the agent forwards it, the client accepts, and the confirmation fails. Non-negotiable design rules: **(a)** every Pool B line carries an explicit *subject-to-confirmation* state visible in the client-facing output; **(b)** a short validity window — **24–72 h** [INFERENCE] — bounding both availability drift and FX drift (§4c.5.4); **(c)** the assistant must be able to compose an all-instant-confirm quote on request, which means Pool A parity inventory has a real job inside D2 even at 7% net.

---

## 4c.5 Rate and contract management — where the money leaks

### 4c.5.1 Documented failure modes

| # | Failure mode | Evidence | Tier | Money mechanism |
|---|---|---|---|---|
| 1 | **Net-rate confidentiality breach** — agent or downstream sub-agent exposes the wholesale rate | "Net rate contracts include strict confidentiality clauses. Disclosing wholesale prices to clients or competitors violates agreements and can result in losing..." — dmcquote | C | Loss of the contract, i.e. **loss of the exclusivity that is the entire D2/D3 thesis.** Worst-case severity in this table. |
| 2 | **Parity break via net-rate resale into public channels** — a partner publishes the net rate as a retail price | "OTAs can cause rate parity problems by reselling wholesale or net rates, offering mobile- or member-only discounts and displaying taxes or fees" — Mews; corroborated by altexsoft's rate-parity treatment | C / B | SatuSatu's own D2C price is undercut **by its own supply**, and the Balinese operator sees its product discounted below contract. |
| 3 | **Mispriced net rate below true cost** — manual rate sheets, seasonal tiers, 253 SKUs | Arival's guidance is that net rates are derived by working a percentage off retail while remaining viable — i.e. a **manual derivation per product** | B | Direct negative-margin sales, invisible until reconciliation. |
| 4 | **Stale contract / expired seasonal rate still live** | No direct source found — **unverified** | — | Selling at last season's rate while paying this season's cost. |
| 5 | **Unenforced volume tier** — partner retains a tier rate after volume falls | No direct source found — **unverified** | — | Permanent margin give-away; ratchets one way because nobody wants the downgrade conversation. |
| 6 | **Currency drift** — see §4c.5.4 | PBI 17/3/2015 + net-30 exposure | A | 19–63% of contribution on an adverse move. |

Modes 4 and 5 are asserted from the mechanism, not from a source. Flag 4c-G8.

### 4c.5.2 The 253-SKU observation that makes this urgent

At 253 SKUs with volume-tiered net rates across multiple tiers and seasons, the rate matrix is in the low thousands of cells. **Modes 3, 4 and 5 are spreadsheet-integrity failures, and spreadsheet-integrity failures are silent** — they surface at reconciliation, months after the money left. There is no AI problem here.

> **The fix is a nightly rule-check job, not a model.** Four deterministic assertions: (i) every live net rate ≥ cost × margin floor; (ii) no contract past its expiry serving live rates; (iii) every partner's assigned tier matches trailing-90-day volume; (iv) a parity crawl of partner storefronts against the contracted floor. This is scripted SQL and a scraper — **plausibly under 1 eng-week [INFERENCE]** — and it defends against the highest-severity item in the table (mode 1/2 leading to contract loss). **Highest return per eng-week of anything in 4c.**

### 4c.5.3 What contract terms must carry, given zero engineering capacity

Because engineering is the binding constraint, prefer **contractual** controls over **coded** ones:
- Explicit downstream-resale and sub-agent clauses — mode 1 and 2 are committed by the partner's *customers*, so the clause must bind downstream.
- A stated retail floor per SKU, not just a net rate.
- Tier eligibility with an automatic review date and an explicit downgrade mechanism (removes the human conversation that causes mode 5).
- Contract expiry defaulting to *suspension of rates*, not to rollover (removes mode 4 by construction).

### 4c.5.4 ⚠️ FX: USD display against IDR settlement

**A regulatory fact first, and it is decisive.** Bank Indonesia Regulation **No. 17/3/PBI/2015** obliges the use of Rupiah in the territory of Indonesia. Standard Chartered's customer summary of the regime states it plainly: **"The Rupiah must be used for all domestic transactions. The Rupiah must be used in all Invoice issuances for domestic payments."** Clifford Chance's briefing: "all corporations and individuals are required to use Rupiah as the currency for all transactions, cash and non-cash." [VERIFIED — Tier A for the regulation itself (bi.go.id publishes the PBI series); Tier B for the law-firm and bank summaries.]

**Consequence for the B2B contract, split by partner domicile:**

| Partner domicile | Invoicing | Where FX exposure sits |
|---|---|---|
| **Indonesian-domiciled agent** | **Must be IDR** for the domestic invoice | If SatuSatu quotes USD and invoices IDR, SatuSatu carries the drift between quote and invoice. **Displaying USD to an Indonesian agent is a compliance question, not just a UX choice.** |
| Foreign-domiciled (Singapore, Australia, India outbound) | USD permissible | SatuSatu pays Balinese operators in IDR while receiving USD → **SatuSatu carries the whole exposure.** |

**Why this is severe rather than annoying — the margin arithmetic:**

| | Pool B gross 27% | **B2B retained 8 / 12 / 16%** |
|---|---|---|
| A 3% adverse FX move consumes | 11% of gross margin | **19% / 25% / 38% of contribution** |
| A 5% adverse move consumes | 19% of gross margin | **31% / 42% / 63% of contribution** |

[INFERENCE] arithmetic on the §4c.0.3 bands. **The move sizes are illustrative — no IDR/USD volatility measurement was collected** (flag 4c-G9). Combine with net-30 credit (§4c.6) and every invoice is an **unhedged 30-day short-USD/long-IDR position**, entered by default, at a scale that can exceed the margin on the sale.

> **Mitigations, all zero-engineering, in preference order: (1) denominate B2B net rates in IDR** — eliminates the exposure and satisfies PBI 17/3/2015 for domestic partners in one move; **(2)** if USD is required for foreign partners, add a **stated FX reset threshold** (re-rate if spot moves beyond ~2%); **(3)** bound quote validity to 24–72 h (§4c.4.3), which caps drift per quote; **(4)** prepay wallet (§4c.6), which collapses the exposure window from 30 days to zero. These are contract clauses and a checkout setting. **None of them requires the engineering team.**

---

## 4c.6 Reconciliation, invoicing, credit and disputes

### 4c.6.1 Documented TAA / DMC B2B terms

All **Tier C** (operator T&Cs, agency blogs, DMC marketing), but **mutually corroborating across independent sources**, which is the best available:

| Term | Observed | Source |
|---|---|---|
| Deposit on confirmation | **20%** non-refundable | LinkedIn DMC/tour-operator sample T&Cs; experiencenorthdmc.com ("usually 20% of full cost") |
| Deposit, alternative | **30%** (vs 50% as the tighter norm) | dmcquote — cites "Lower deposit percentages: 30% instead of 50%" as a cash-flow concession |
| Balance due before travel | **21–35 days** prior | 21 days (B2B portal example); 35 days (LinkedIn sample T&Cs) |
| Credit line for repeat business | **net-30** | dmcquote: "For repeat business, some suppliers offer net-30 terms" |
| Settlement shape | **three moments** — deposit at confirmation, balance before travel, short post-travel reconciliation | paidaidmc.com |
| Generic B2B baseline | Net 30 is the most common US B2B term | Corpay |

**Planning band for SatuSatu:** deposit **20–30%**, balance **21–35 days** pre-travel, credit **net-30** for established partners only.

### 4c.6.2 The cost of running the function, and why credit should not launch

| | low | base | high |
|---|---|---|---|
| Finance FTE for AR, invoicing, reconciliation, disputes | 0.25 | 0.35 | 0.50 |
| Fully-loaded finance cost/yr (§4a.5 band × multiplier) | US$2,050 | US$3,440 | US$6,840 |
| Bad debt, % of B2B GMV | 0.5% | 1.0% | 3.0% |
| Bad debt at base-case D3 volume (20 partners × 800 × US$25 = US$400,000 GMV) | US$2,000 | US$4,000 | US$12,000 |
| **Total cost of the credit function** | **US$4,050** | **US$7,440** | **US$18,840** |

[INFERENCE] on the FTE and bad-debt bands — **no TAA-specific B2B bad-debt rate was found** (flag 4c-G10).

> # 🔴 Compare against the entire base-case D3 contribution: 20 partners × US$710 = **US$14,200 per year.** Running the credit function at base cost (**US$7,440**) consumes **52% of it**; at the high band it **exceeds total contribution by 33%.**

**Therefore: launch D2 prepay-wallet-only.** Top-up balance, book against balance, no invoice, no AR, no underwriting (removing the expensive half of §4c.1.1), no dispute-as-leverage, no 30-day FX window (§4c.5.4). The concession is real — credit terms are a genuine agent purchasing criterion (see §4c.7.2) — so treat credit as a **tier-3 privilege earned by trailing volume**, deliberately scarce, priced into the tier rather than given away as an onboarding sweetener. **Deposits and staged payments are the middle ground and are already the documented industry norm** (20–30% / 21–35 days), so offering them is not a disadvantage.

**What AI does here:** reconciliation matching (bank credit → invoice → booking → operator payable) is deterministic matching with a fuzzy tail; the fuzzy tail is a real AI fit but it is a *small* tail at 20 partners. **Do not build. Revisit above ~100 partners.**

---

## 4c.7 🔴 What does a partner integrate SatuSatu *for*?

### 4c.7.1 Testing the Pool B answer

Stated answer: exclusive direct-contracted Balinese long-tail supply. **The test that matters is not whether Pool B is attractive — it is whether Pool B competes for the thing SatuSatu is asking the agent to give up.**

Every documented incumbent competes on **breadth**:

| Incumbent | Documented position | Tier |
|---|---|---|
| **TBO Holidays** | 200,000+ global sightseeing products advertised to agents | B |
| **GlobalTix** | 12,000+ agents, 150,000+ experiences, 25M tickets/yr, cash-flow positive, holds the InJourney Borobudur mandate | B |
| **Golden Rama** | `rols.golden-rama.com` runs **both** an Agent Login and a Supplier Login, sells Attractions alongside its other lines | A (live product) |
| **Traveloka × Trip.com** | Pooling attractions inventory (June 2026) | B |
| **Holibob** | GraphQL partner API + white-label experience website + a "Connectivity Manager" function | A (own docs, IATA submission) |

**Nobody competes on Bali-only exclusivity.** Two readings — an unoccupied position, or an unoccupied position because it does not sell. The decisive consideration resolves it:

> 🔴 **An agent's platform choice is a workflow decision, not a product decision.** What they are buying is one login, one rate logic, one invoice, one AR relationship, one support contact. **A Bali-only 253-SKU supplier is, by construction, an *additional* login.** Golden Rama's architecture is the clearest evidence: an Agent Login *and* a Supplier Login on one rail, next to flights and hotels — the incumbent's moat is that it consolidates the agent's workflow, not that its Bali catalogue is better.
>
> **Therefore Pool B does not compete for the agent's platform slot. It competes for a line item inside whichever platform the agent already logs into.** This is the central finding of §4c.7, and it inverts the D2/D3 question: the problem is not building a better portal, it is **getting Pool B into portals that already exist.**

### 4c.7.2 What actually drives agent platform adoption

⚠️ **No Tier A or Tier B study of Indonesian or SE Asian agent platform-selection drivers was found** (flag 4c-G7). Ranked below by weight of *structural* evidence, with the basis stated. **This ranking is [INFERENCE] and is the single most valuable thing to replace with primary research — 10–15 agent interviews would settle it and cost almost nothing.**

| Rank | Driver | Basis | Can SatuSatu win it? |
|---|---|---|---|
| 1 | **Incumbent relationship + workflow consolidation** | Golden Rama's Agent+Supplier architecture; every incumbent leads with breadth | **No.** Structurally cannot — it is a 253-SKU single-destination supplier. |
| 2 | **Credit terms** | net-30 documented as the concession suppliers make for repeat business (dmcquote); deposit reduction 50%→30% framed as the lever | **Weakly, and it is the wrong fight** — §4c.6 shows credit consumes ~52% of contribution. |
| 3 | **Rate** | Net-rate contracting is the documented B2B mechanism (TourCMS, Arival) | **Yes on Pool B (27% gross gives real room). No on Pool A** — 10% gross, and the agent reaches it via GlobalTix anyway. |
| 4 | **Inventory uniqueness** | No source found showing agents switch platforms for unique inventory | **Yes — this is the only genuinely defensible claim.** But note it is ranked *below* rate and credit on available evidence, not above. |
| 5 | **Support quality** | Not evidenced as a selection driver | Yes, and cheap (§4c.2's labour arbitrage) — but it retains, it does not acquire. |

> **The uncomfortable read: the stated answer (Pool B exclusivity) is the driver for which we found the *least* supporting evidence, while the drivers with the most evidence behind them (workflow consolidation, credit) are ones SatuSatu either cannot win or cannot afford to win.** That does not make Pool B wrong — exclusive supply is genuinely the only asset here that an incumbent cannot copy. It means **Pool B is a reason to *stock* SatuSatu, not a reason to *integrate* it.**

### 4c.7.3 Minimum viable Pool B catalogue size

The right unit is **not SKU count — it is itinerary-slot coverage.** An agent's question is "can I build my client's Bali week from this?", and one unmissable gap sends them back to their incumbent for the whole itinerary.

A Bali itinerary draws on roughly 15–20 recognised demand clusters — Uluwatu/Kecak, Tanah Lot, Ubud core (rice terraces, Monkey Forest), waterfalls, Mount Batur sunrise, Nusa Penida day trip, Nusa Dua watersports, ATV, rafting, cooking class, spa, animal parks, waterpark, snorkel/dive, sunset cruise, Kintamani/coffee. A typical booked itinerary is **2–4 activities/day over 4–7 days = 8–28 booked slots.**

| | low | base | high |
|---|---|---|---|
| Clusters that must be covered | 10 | 15 | 20 |
| Credible exclusive options per cluster | 2 | 2–3 | 3 |
| **Minimum viable Pool B catalogue** | **~20 SKUs** | **~30–45 SKUs** | **~60 SKUs** |

[INFERENCE]. **The cluster list is general Bali destination knowledge, not a sourced taxonomy** — flag 4c-G11; it should be rebuilt from SatuSatu's own booking mix, which would take minutes internally.

**Two qualifications that matter more than the number:**
1. **Exclusivity must be *provable* to the agent, per SKU.** Against GlobalTix's 150,000 experiences, "we have unique Bali product" is unfalsifiable marketing. A per-SKU "not available via any aggregator" assertion the agent can spot-check is the actual asset. If Pool B products are also listed on GlobalTix or Klook by their operators, **Pool B is not exclusive and the entire D2/D3 thesis fails.** 🔴 **Verify this before anything else in this bracket — it is a one-day catalogue cross-check and it is a genuine kill criterion.**
2. **Below ~20 SKUs, do not launch a B2B offer at all** — it cannot fill an itinerary, and a partner who tries once and fails does not try twice.

### 4c.7.4 🔴 Lower-effort routes than an API, ranked

Given zero engineering capacity, effort is the currency. **Note that the two cheapest routes are also the two that work correctly with on-request inventory** — which is not a coincidence: an API's advantage is latency, and Pool B has no latency advantage to sell.

| # | Route | SatuSatu eng-weeks | Time to live | Works with on-request Pool B? | Reach |
|---|---|---|---|---|---|
| **1** | **Rate sheet / CSV feed + WhatsApp or email request-to-book** | **0** | **Days** | ✅ **Natively — it *is* the request-to-book workflow** | Only agents you contact |
| **2** | **List Pool B as a supplier on a rail agents already use** (GlobalTix, Holibob, Golden Rama's Supplier Login, an existing marketplace) | **0–2** (they integrate, or accept a feed) | Weeks–months (their queue) | ✅ Yes — aggregators handle on-request supply routinely | **GlobalTix alone: 12,000+ agents** |
| 3 | Hosted booking page + per-agent tracking code, net rate applied at checkout | 1–3 | 2–6 wks | ✅ Yes | Only agents you contact |
| 4 | Agent seats on the existing D2C site: net-rate flag + prepay wallet (**= minimum-viable D2**) | 3–8 | 1–3 months | ✅ Yes | Only agents you contact |
| 5 | White-label storefront per agent (the iVenture-style pattern; Holibob ships this as a white-label experience website) | 8–20 | 3–6 months | ⚠️ Degrades — a consumer-facing storefront that cannot confirm is a bad storefront | Agent's own customers |
| 6 | **D3 reseller API** | **24–46 + ongoing** (§4c.3.4) | **6–12 months** | ❌ **Cannot be certified to industry norms (§4c.0.1)** | ~20+ large agents required to break even |

[INFERENCE] on all effort figures. iVenture's white-label pass platform is taken from the brief and **was not independently verified in this bracket**; Holibob's white-label experience website and Connectivity Manager **were** verified from Holibob's own documentation and its IATA AirTechZone submission [Tier A].

### 4c.7.5 Should SatuSatu ride an existing rail? — the trade, stated as a trade

**The case for route 2 (list Pool B as a supplier on GlobalTix / Holibob / similar):**
- **Zero engineering** against the binding constraint.
- Reaches 12,000+ agents who are **already logged in** — it answers §4c.7.1's workflow finding instead of fighting it.
- The aggregator absorbs certification, integration support, agent onboarding, credit, AR, disputes and FX — i.e. **§4c.1, §4c.3, §4c.6 in their entirety.**
- SatuSatu **keeps the direct operator contract**, which is the actual asset.
- It is the only route that produces a real demand signal fast enough to inform whether D2/D3 is worth building at all.

**The case against, and it is not weak:**
- Hands SatuSatu's only exclusive, high-margin asset to the aggregator that also supplies Pool A — **deepening the dependency Pool A already represents.**
- Concedes 10–20 points of the 27%, plausibly landing Pool B near Pool A's economics — which dissolves the reason Pool B was tiered separately.
- **Disintermediation risk is concrete and directional:** GlobalTix sees the volume, sees which Balinese operators produce it, and contracts them directly. It has the scale, the agent base, and an Indonesian SOE mandate (Borobudur) demonstrating it does exactly this kind of direct sourcing.
- Once Pool B is on an aggregator, **it is no longer exclusive**, and per §4c.7.3 exclusivity is the whole thesis.

> **The trade is: distribution reach now, at the cost of the exclusivity that makes Pool B worth distributing.** The resolution is sequencing rather than choosing:
>
> 1. **Route 1 immediately (0 eng-weeks).** Rate sheet plus WhatsApp request-to-book to 20–40 named agents. This tests demand for Pool B at agent rates **inside the month**, with no build and no exclusivity given away. If Pool B does not sell by rate sheet, it will not sell by API — and that answer arrives before any engineering is spent.
> 2. **Route 4 next (3–8 eng-weeks) only if route 1 produces repeat buyers.** Prepay-wallet agent seats on the existing site. Breaks even at ~5 base-size or ~37 small partners (§4c.3.4) and tolerates the modal small agent.
> 3. **Route 2 selectively** — list the *replaceable* part of Pool B on an aggregator to buy reach; **withhold the genuinely irreplaceable SKUs.** Partial listing is the hedge, and it makes the disintermediation risk a bounded loss rather than an unbounded one.
> 4. **Route 6 (D3) only against named partners already transacting ≥480 Pool B bookings/yr through routes 1–4, and only after Pool B has a real-time availability feed.** Both conditions, not either.
>
> 🔴 **The two prerequisites for D3 are the availability feed and the exclusivity verification (§4c.7.3 qualification 1) — neither is an engineering project, and both are worth more than D3 itself.**

---

## 4c.8 Gaps and unverified items

| # | Gap | Why it matters | Status |
|---|---|---|---|
| **4c-G1** | Indonesian engineering salary — **drives the entire break-even** | A 2× error moves N\* from ~20 to ~10 or ~45 | Not found. Trivially resolvable internally. |
| **4c-G2** | SatuSatu's actual conceded B2B spread on Pool B | Every contribution figure scales linearly off it | Not found. Internal decision, not research. |
| **4c-G3** | Any TAA-specific partner-onboarding time study | All §4c.1 hour counts are inferred | Not found. May not exist publicly. |
| **4c-G4** | Indonesian agent Bali-attraction volume distribution | Decides whether the 480-booking floor excludes 70% or 99% of the market | Not found. **10–15 agent interviews would settle it.** |
| **4c-G5** | Minutes per on-request operator confirmation | Largest single cost line in §4c.2 | Not found. Measurable internally in a week. |
| **4c-G6** | **API partner never-launch rate** | If 40–60%, cost per productive integrator is 1.7–2.5× and N\* rises to 35–50 | Widely asserted, **no Tier A/B measurement found.** |
| **4c-G7** | Indonesian/SE Asian agent platform-selection drivers | §4c.7.2's ranking is inference; it decides the whole strategy | **Not found. Biggest research gap in 4c.** |
| **4c-G8** | Stale-contract and unenforced-tier failure modes | Asserted from mechanism, no source | Not found. |
| **4c-G9** | IDR/USD volatility measurement | The 3%/5% move sizes in §4c.5.4 are illustrative | Not collected. |
| **4c-G10** | TAA B2B bad-debt rate | Drives whether credit consumes 28% or 130% of contribution | Not found. |
| **4c-G11** | Sourced Bali demand-cluster taxonomy | Minimum-viable-catalogue figure rests on general knowledge | Not found externally; **rebuild from SatuSatu's own booking mix.** |
| **4c-G12** | Measured savings from sandboxes / docs / AI integration copilots | The brief asked for evidence; **only vendor marketing exists** | **Searched specifically. Tier C only. Rejected as planning input.** |
| **4c-G13** | iVenture Card's white-label platform mechanics and commercials | Cited in the brief as a comparator | **Not independently verified in this bracket.** |

---

*End of section 4c.*
