A benchmark for flight-search intent parsing. Real-world queries in three languages — deduped, de-identified, hand-labeled — put to every model on the same terms; eleven independent graders decide whether each one caught the traveler's intent.
| Model | Cases passed | Dimensions | oneway | roundtrip | multi | flexible-date | fuzzy-location | invalid-date | invalid-others | Errors | ||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.6-solmedium |
87%
212/244
|
86%
n=59
|
87%
n=153
|
67%
n=12
|
84%
n=116
|
83%
n=109
|
100%
n=7
|
100%
n=3
|
— | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 2 | gpt-5.6-terramedium |
86%
209/244
|
93%
n=59
|
86%
n=153
|
17%
n=12
|
79%
n=116
|
79%
n=109
|
100%
n=7
|
100%
n=3
|
— | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 3 | claude-opus-5medium |
85%
207/244
|
92%
n=59
|
86%
n=153
|
75%
n=12
|
86%
n=116
|
80%
n=109
|
0%
n=7
|
67%
n=3
|
— | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 4 | claude-sonnet-5medium |
79%
192/244
|
86%
n=59
|
80%
n=153
|
46%
n=13
|
77%
n=116
|
73%
n=110
|
0%
n=7
|
67%
n=3
|
2 | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 5 | gemini-3.5-flash-litemedium |
74%
181/244
|
85%
n=59
|
77%
n=153
|
8%
n=12
|
68%
n=116
|
64%
n=109
|
0%
n=7
|
67%
n=3
|
25 | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 6 | gemini-3.5-flashmedium |
74%
181/244
|
80%
n=59
|
78%
n=153
|
25%
n=12
|
72%
n=116
|
72%
n=109
|
0%
n=7
|
67%
n=3
|
32 | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 7 | gemini-3.8-flashmedium |
74%
181/244
|
83%
n=59
|
73%
n=153
|
8%
n=12
|
63%
n=116
|
63%
n=109
|
100%
n=7
|
100%
n=3
|
44 | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 8 | gpt-5.6-lunamedium |
71%
174/244
|
81%
n=59
|
71%
n=153
|
17%
n=12
|
66%
n=116
|
64%
n=109
|
57%
n=7
|
100%
n=3
|
— | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 9 | claude-haiku-4-5 |
53%
129/244
|
41%
n=59
|
56%
n=153
|
17%
n=12
|
47%
n=116
|
31%
n=109
|
71%
n=7
|
100%
n=3
|
— | ▸ | |||||||||||||||||||||||||||||||||||||||||||||||
Per-dimension pass rate over the repeats that scored it
Cost & latency
By language
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Model | Failed cases | Errored calls | Where the failures sit |
|---|