// Airline Agent Experience

Scans pass.
Agents fail.

AI agents are starting to browse, compare and book for real customers. Last Mile sends a real browser agent through real airline journeys and checks the only thing that matters: did the task actually complete?

// Benchmark — 7 airlines tested

Who completes the task

Last Mile benchmark: 7 airlines, ordered by measurable-run coverage, with tasks completed, access-gated runs and the top failing signal.
# Airline Completed / measurable Access-gated Top failing signal
1 Riyadh Air 0 / 15 1 Failure transparency
2 easyJet 0 / 2 Discoverability
3 Emirates 0 / 2 1 Input affordances
4 Lufthansa 0 / 2 1 Input affordances
5 airBaltic 0 / 1 Task completion
6 Delta 0 / 1 1 Task completion
7 Qatar Airways 0 / 1 Input affordances

Ordered by measurable-run coverage, not pass rate — every airline in this sample passed zero measurable runs. Riyadh Air accounts for 15 of the 24 measurable runs; the other 6 carriers contribute one or two tasks each. 2 additional runs failed on infrastructure grounds and are correctly excluded, not counted as fails.

Scope of this sweep

Methodological information
Measurable pass rate
0%

0 of 24 passed

Measurable runs
24

16 distinct tasks across 7 airlines

Blocked
4

4 of 28 runs

Airlines tested
7

airlines in this sweep

A spot-check, not an industry sweep — concentrated on one carrier. This is 7 airlines, 16 distinct tasks — Qatar Airways, Lufthansa, Riyadh Air, airBaltic, easyJet, Emirates, Delta. A mix of agent models drove — claude-sonnet-4-6 (15), kimi-k2.5:cloud (10), unattributed (6), run 2026-07 to 2026-08.

Top failure signals across the measurable sample: failureTransparency (11) · inputAffordances (5) · discoverability (4) · taskCompletion (4)

Access-gated is not failure

Access-gated is not failure. A booking-reference lookup that correctly asks for a booking reference is a site working as designed. 4 of 28 airline runs hit an honest access wall — a login, PNR, or payment gate the agent hit as expected. Gated runs are excluded from the Index — never counted as a pass, never counted against the site.

// The loop

Assess. Diagnose. Optimize.

Playground Try it in the playground — Friction Airways, a test airline we own. Open the playground (opens in a new tab)

Key takeaways

Takeaway_01

Zero of 7 airlines completed a measurable task.

Takeaway_02

Failure transparency is the top failure signal — 11 of 24 measurable runs went dark instead of erroring legibly.

Takeaway_03

The controlled site we own goes from 0% to 100% — hostile to ready, same tasks, same check.

// WebMCP case study

Judge WebMCP by the outcome it produces

Friction Airways is a controlled airline site we own, published in three WebMCP-readiness profiles. Same site, same 21-task measurable slice — only the profile changes.

Hostile
0% · 0/21
Median
71% · 15/21
Ready
100% · 21/21

Hostile's fails cluster on one signal: 18 of 21 measurable fails returned no machine-readable answer at all; ready has zero measurable fails. This is the evaluate → diagnose → optimize → retest loop in miniature, run on one controlled site, not an airline-industry claim.