A benchmark for flight-search intent parsing. Real-world queries in three languages — deduped, de-identified, hand-labeled — put to every model on the same terms; eleven independent graders decide whether each one caught the traveler's intent.

Run
2026-09-02T1859Z
Anchor date
01/01/2026
Cases
244
Languages
de · en · es
Prompt
v6
Repeats
1
Models
9
Calls
2,196
Fuzzy text similarity
cosine · threshold ≥ 0.72 · text-embedding-3-small

Leaderboard

ModelCases passedDimensions onewayroundtripmultiflexible-datefuzzy-locationinvalid-dateinvalid-others Errors

Dimension strip & tag columns — pass rate

  • 95%+
  • 80–95%
  • 50–80%
  • under 50%
  • never scored

Tag columns hover a header for its definition

flexible-date
a window rather than a day: “October or November”, “next 10 to 15 days”. Any date inside the window passes.
fuzzy-location
Places with no IATA code: “beach towns in Southeast Asia”, “West Europe”. evaluated against golden sets by meaning (cosine similarity).

Failed cases by tags

Language
ModelFailed casesErrored calls Where the failures sit