Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

τ²-Bench: Airline

τ²-Bench Airline tests whether a model can do an airline support agent's job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model doesn't need any domain knowledge; it's scored entirely on how well it executes tool calls across multi-step task trajectories. We run it continuously against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. We use this benchmark because it has a high floor, so we can assess provider variance and not model capability. Our routing algorithm for tool call requests uses these same signals to send traffic to the best performing endpoints.

Last benchmark run Sep 1, 2026, 9:01 PM UTC

PaperGitHub
Tool-call errors before → after Auto Exacto

3.9% → 4.1%

Tool-call error rate on models enrolled in Auto Exacto, OpenRouter's automatic provider optimization for tool-calling requests.
Model comparisonCost efficiencyTool-call reliabilityLeaderboardWhy we run itWhat scores tell youHow tasks are scoredMethodologyAPI access

Model comparison

Most Accurate

Favicon for anthropic
Anthropic: Claude Fable 5

81.5%

Best Value

Favicon for google
Google: Gemini 3.7 Flash

$0.077/task

Fastest

Favicon for z-ai
Z.ai: GLM 5.3

70s

Accuracy
Representative-run accuracy, best first.
Cost per task
Average cost per graded task, cheapest first.
Time per task
Average wall-clock time per task, fastest first; agents that loop or stall run long.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Tool-call reliability

Tool-call errors
Share of this benchmark's own requests where the model called a tool that doesn't exist, passed arguments that don't match the tool's schema, or emitted arguments that aren't valid JSON.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

#ModelStd dev
1
Anthropic: Claude Fable 5
Pareto
81.5%±1.1pp$0.982.3m5.84k
2
Google: Gemini 3.7 Flash
Pareto
80.6%--$0.0772.0m14.7k
3
Claude Opus 5
80.1%±2.0pp$0.502.3m7.34k
4
Z.ai: GLM 5.3
80.0%--$0.09170s4.66k
5
Amazon: Nova Micro 1.0
78.7%±2.0pp$1.281.7m5.69k
6
Qwen: Qwen3.8 27B
78.7%±2.0pp$0.0825.0m12.2k
7
Qwen: Qwen3.5 397B A17B
78.7%±3.6pp$0.0995.9m15.6k
8
DeepSeek: DeepSeek V4 Pro 0813
78.0%±1.1pp$0.116.0m18.8k
9
Anthropic: Claude Opus 4.5
77.7%±2.3pp$0.542.9m10.5k
10
Google: Gemini 3 Flash Preview
77.3%±2.0pp$0.122.5m20.5k
11
StepFun: Step 3.7 Flash
Pareto
77.3%--$0.0203.8m10.9k
12
Anthropic: Claude Opus 4.7
77.1%±1.6pp$0.411.7m5.72k
13
Anthropic: Claude Opus 4.8
76.9%±1.4pp$0.502.3m8.17k
14
NVIDIA: Nemotron 3 Ultra
76.9%±0.8pp$0.102.6m8.57k
15
Qwen: Qwen3.5-122B-A10B
76.7%±3.6pp$0.159.1m35.9k
16
OpenAI: GPT-5.6 Sol Pro
76.7%--$1.865.6m22.9k
17
Qwen: Qwen3.8 2.4T A95B
76.7%--$0.3110.1m20.4k
18
Anthropic: Claude Sonnet 5
76.7%±2.2pp$0.202.3m7.98k
19
OpenAI: GPT-5.6 Sol
76.7%±0.7pp$0.212.1m5.14k
20
Google: Gemma 4 31B
Pareto
76.1%±4.0pp$0.0165.5m8.28k
21
DeepSeek: DeepSeek V4 Pro 0423
76.0%±2.8pp$0.0422.7m7.59k
22
Anthropic: Claude Opus 4.6
76.0%±2.6pp$0.483.3m10k
23
SpaceXAI: Grok 4.6
76.0%--$0.354.1m11.6k
24
Anthropic: Claude Sonnet 4.6
75.9%±1.3pp$0.323.4m11.4k
25
Google: Gemini 3.5 Flash Lite
75.7%±1.7pp$0.111.7m25.1k
26
OpenAI: GPT-5.1
75.7%±4.3pp$0.438.7m37.6k
27
OpenAI: GPT-5.4
75.7%±0.3pp$0.302.5m10.8k
28
OpenAI: GPT-5.5
75.3%±2.7pp$0.502.6m8.28k
29
Google: Gemini 3.1 Pro Preview
75.3%±2.0pp$0.362.3m12.8k
30
Z.ai: GLM 5
75.3%±4.2pp$0.0362.0m5.54k
31
Z.ai: GLM 5.2
75.2%±2.7pp$0.0382.9m7.66k
32
DeepSeek: DeepSeek V4 Flash 0423
Pareto
75.1%±3.3pp$0.0092.4m8.29k
33
Google: Gemini 3.5 Flash
74.7%±0.7pp$0.422.8m28.3k
34
Xiaomi: MiMo-V2.5-Pro
74.4%±3.9pp$0.0276.7m12.9k
35
Qwen: Qwen3.6 27B
74.1%±4.8pp$0.109.8m23.6k
36
Z.ai: GLM 5.1
74.1%±3.6pp$0.0572.4m5.73k
37
DeepSeek: DeepSeek V4 Flash Vision Exp
74.0%--$0.0313.1m15.9k
38
MoonshotAI: Kimi K2.6
73.7%±2.7pp$0.0613.8m9.41k
39
Anthropic: Claude Sonnet 4.5
73.6%±3.3pp$0.333.4m10.3k
40
Qwen: Qwen3.5-35B-A3B
73.3%±7.9pp$0.0567.3m34.1k
41
OpenAI: GPT-5.6 Terra
73.3%±1.3pp$0.1001.6m4.17k
42
OpenAI: GPT-5.2
73.3%--$0.243.5m12.5k
43
Z.ai: GLM 5.3 Flash
Pareto
73.3%--$0.0052.6m4.09k
44
DeepSeek: DeepSeek V4 Flash 0731
73.2%±6.3pp$0.0096.0m16.7k
45
Z.ai: GLM 4.7
73.1%±4.5pp$0.0362.5m5.47k
46
Google: Gemini 3.6 Flash
73.0%±0.3pp$0.242.4m22.9k
47
Google: Gemini 3.1 Flash Lite
72.3%±1.0pp$0.0972.4m51.2k
48
DeepSeek: DeepSeek V3.2
72.1%±2.6pp$0.0234.6m8.07k
49
Meta: Muse Glimmer 30B
72.0%±3.0pp$0.0413.4m13.5k
50
MoonshotAI: Kimi K2.7 Code
71.8%±3.4pp$0.0664.4m7.45k
51
Z.ai: GLM 4.5 Air
71.4%±3.9pp$0.0152.3m4.45k
52
OpenAI: GPT-5.6 Luna Pro
71.3%±1.3pp$0.0645.8m34.2k
53
MoonshotAI: Kimi K2.5
71.2%±2.7pp$0.0293.2m6k
54
MiniMax: MiniMax M2.7
71.1%±4.3pp$0.0182.3m5.87k
55
Qwen: Qwen3.6 35B A3B
71.1%±7.9pp$0.0433.7m24.1k
56
MoonshotAI: Kimi K3
70.9%±2.0pp$0.348.1m9.11k
57
MiniMax: MiniMax M3
70.9%±6.4pp$0.0232.2m6k
58
Xiaomi: MiMo-V2.5
70.9%±2.4pp$0.0084.9m14.5k
59
Auto Router (Beta)
70.7%--$0.242.5m17.9k
60
OpenAI: GPT-5.6 Luna
70.7%±0.0pp$0.0091.7m5.4k
61
OpenAI: GPT-5 Mini
69.7%±3.0pp$0.1311.6m58.4k
62
MoonshotAI: Kimi K2 Thinking
69.2%±10.8pp$0.0441.8m3.91k
63
DeepSeek: DeepSeek V3.1 Terminus
68.6%±4.1pp$0.0396.8m11.6k
64
Xiaomi: MiMo-V2-Flash
68.6%±5.7pp$0.00544s3.62k
65
OpenAI: GPT-5
68.4%±3.8pp$0.6813.4m61.2k
66
DeepSeek: DeepSeek V3.2 Exp
68.3%±4.7pp$0.0557.6m8.77k
67
Google: Gemma 4 26B A4B
68.3%±3.9pp$0.0174.8m13k
68
Anthropic: Claude Sonnet 4
68.0%±2.0pp$0.332.9m8.95k
69
Thinking Machines: Inkling
67.7%±4.3pp$0.07969s3.59k
70
Qwen: Qwen3.5-9B
67.5%±5.0pp$0.0249.6m31.3k
71
Anthropic: Claude Haiku 4.5
67.3%±3.5pp$0.122.0m12.8k
72
OpenAI: GPT-5.2 Chat
67.3%--$0.1180s4.28k
73
MiniMax: MiniMax M2.1
67.0%±4.0pp$0.01580s5.43k
74
OpenAI: GPT-5.4 Nano
67.0%±0.3pp$0.0332.5m17.5k
75
Z.ai: GLM 4.6
66.6%±11.2pp$0.0363.4m6.05k
76
OpenAI: gpt-oss-120b
64.4%±3.6pp$0.0133.9m13.4k
77
MiniMax: MiniMax M2.5
64.1%±4.8pp$0.0162.8m6.61k
78
OpenAI: GPT-5.4 Mini
63.0%±5.0pp$0.0922.2m12.5k
79
Thinking Machines: Inkling Small
62.0%±3.3pp$0.0474.5m3.91k
80
Z.ai: GLM 4.7 Flash
61.9%±6.1pp$0.0082.6m7.83k
81
Google: Gemini 2.5 Pro
61.3%±0.7pp$0.223.2m15.1k
82
Qwen: Qwen3 235B A22B Thinking 2507
61.3%±4.1pp$0.0586.8m20.6k
83
DeepSeek: DeepSeek V3.1
60.0%±10.4pp$0.0596.2m9.11k
84
Qwen: Qwen3 Coder Next
58.5%±6.3pp$0.02459s2.73k
85
Google: Gemini 2.5 Flash
57.7%±0.3pp$0.02979s8.17k
86
Ling-3.0-flash
57.3%--$0.0122.0m16.8k
87
MoonshotAI: Kimi K2 0905
54.9%±5.2pp$0.112.6m3.13k
88
DeepSeek: R1 0528
54.9%±6.0pp$0.08313.1m17k
89
NVIDIA: Nemotron 3 Nano 30B A3B
52.3%±2.6pp$0.0247.5m50k
90
OpenAI: gpt-oss-20b
51.4%±2.5pp$0.02318.7m96.5k
91
OpenAI: GPT-4.1
48.7%±0.7pp$0.1251s2.66k
92
OpenAI: GPT-5 Nano
47.6%±4.5pp$0.04311.6m99.7k
93
Google: Gemini 2.5 Flash Lite
47.3%--$0.0242.9m45.1k
94
Qwen: Qwen3 Coder 480B A35B
47.0%±5.2pp$0.04564s1.87k
95
Qwen: Qwen3 Next 80B A3B Instruct
46.8%±3.5pp$0.01560s2.25k
96
OpenAI: GPT-4o
46.3%±1.7pp$0.2052s2.08k
97
Qwen: Qwen3 235B A22B Instruct 2507
46.3%±4.3pp$0.0181.7m2.24k
98
OpenAI: GPT-5.3 Chat
46.0%--$0.09965s2.75k
99
OpenAI: GPT-4o (2024-08-06)
44.7%±1.3pp$0.2147s2.22k
100
Meta: Llama 4 Maverick
44.6%±2.3pp$0.0411.8m1.94k
101
OpenAI: GPT-4.1 Mini
44.0%±1.3pp$0.02974s2.45k
102
Mistral: Mistral Small 4
43.7%±2.3pp$0.0141.6m10.2k
103
Qwen: Qwen3 30B A3B
43.7%±6.0pp$0.0153.1m10.1k
104
Qwen: Qwen3 Coder 30B A3B Instruct
43.1%±0.6pp$0.0243.6m3.15k
105
DeepSeek: DeepSeek V3 0324
42.6%±4.5pp$0.0405.5m8.13k
106
Qwen: Qwen3 14B
42.3%±1.0pp$0.0308.7m20.1k
107
Qwen: Qwen3 32B
42.0%±10.0pp$0.0157.6m12.9k
108
Qwen: Qwen3 VL 235B A22B Instruct
40.8%±4.3pp$0.0342.2m3.31k
109
OpenAI: GPT-4o (2024-05-13)
40.0%--$1.1372s3.25k
110
Meta: Llama 3.3 70B Instruct
38.9%±3.5pp$0.01243s529
111
DeepSeek: DeepSeek V3
38.3%±0.3pp$0.0714.7m7.51k
112
Qwen: Qwen3 30B A3B Instruct 2507
37.5%±3.6pp$0.0161.8m2.58k
113
Qwen: Qwen3 VL 30B A3B Instruct
34.3%±3.1pp$0.0733.5m6.68k
114
Qwen: Qwen3 VL 8B Instruct
33.1%±5.1pp$0.02064s2.6k
115
Meta: Llama 3.1 8B Instruct
30.9%±6.5pp$0.00666s1.6k
116
OpenAI: GPT-4o-mini
27.0%±3.0pp$0.02166s3.95k
117
Qwen2.5 72B Instruct
26.0%±1.3pp$0.0799.8m5.67k
118
Mistral: Mistral Nemo
18.7%±6.3pp$0.0201.8m2.7k
119
Qwen: Qwen2.5 7B Instruct
16.7%--$0.01758s2.49k
120
OpenAI: GPT-4.1 Nano
10.3%±0.3pp$0.00761s2.32k

Why we run this benchmark

It's a tool-calling benchmark that is hard to game. Grading depends on live tool-call trajectories rather than memorized answers, so it resists training-data leakage better than Q&A-style evals. It exercises every tool-calling failure mode (wrong arguments, skipped policy checks, giving up, hallucinated confirmations) at a relatively low cost per run. The relative scores also carry more signal than the absolute ones. The same model can score differently across providers, and those deltas are what Exacto routing uses to pick higher-accuracy endpoints.

Each task is a simulated airline support conversation with a scripted user, a toolbox (flight search, booking changes, refunds, loyalty policies), and a gold reference solution. A task passes only if the final database state and the messages to the user match the reference; partial credit is not awarded.

What the scores can and can't tell you

There is still headroom. Top models fail roughly one in five tasks, and the airline domain is the hardest τ²-Bench split. Accuracy differences here separate models that follow multi-step policies from ones that merely chat well.

The floor is high, though. Many tasks reward inaction. A refusal task with an empty gold action list passes for any agent that changes nothing. Even weak models score well above zero, so the meaningful spread sits at the top of the range.

The benchmark is public, so tasks may appear in training corpora. Contamination inflates scores less here than in Q&A-style evals, though, since a leaked task still has to be executed correctly, step by step, against a live database.

Scoring fidelity has limits. The checker verifies two things: the final database hash and exact substring matches in the agent's messages. Each task's natural-language assertions ("agent should refuse the cancellation") are metadata, and no judge model reads the transcript. So a savings calculation fails if the agent says "$23,552.50" when the checker greps for "23553".

The user simulator matters too. We pin it to gemini-2.5-flash so agent scores stay comparable, but the sim is itself an LLM with failure modes of its own. It can stop the conversation before the agent finishes, leak its hidden task instructions, or keep a stuck agent looping until the 200-step ceiling kills the run. Swapping the sim model shifts absolute scores, which is why cross-paper τ²-Bench numbers rarely line up exactly.

How a task is scored

Every task ships a gold solution: a list of tool calls, strings the agent must say, and natural-language assertions. After the conversation ends, the checker replays the gold tool calls against a fresh database and compares hashes with the agent's final database. It then greps the agent's messages for each required string. The reward is the product of those two checks:

reward = db_match × communicate_met   // each ∈ {0, 1}
db_match        = hash(agent DB) == hash(gold DB)
communicate_met = every required string appears in an agent message
any run that hits MAX_STEPS instead of a clean stop scores 0 outright

The rollouts below are from real runs, with gemini-2.5-flash as the user simulator throughout.

reward = 1

Pass: three changes in one request, all three land

Task 17 agent: openai/gpt-5.1

For reservation FQ8APE: add 3 checked bags, swap the passenger to Omar Rossi, and upgrade basic economy to economy, paying with a gift card.

  • Database must match the gold state: update_reservation_flights (economy upgrade), update_reservation_passengers, and update_reservation_baggages with exact arguments
  • communicate_info is empty, so no string check applies
db ✓communicate ✓USER_STOP

The agent looked up the user, found the right reservation among several, confirmed the changes and payment method, then made all three writes: passenger swap, cabin upgrade, and bags. The final database hashes match the gold state and the run ends on USER_STOP, so reward is 1. This is what the eval is designed to measure: multi-step tool use under policy constraints, done correctly.

Methodology

Scores aggregate all successful runs, weighted by task count, with a minimum of 45 graded tasks per model-provider pair. A model's headline score uses its default routing (not pinned to a provider) when one exists; otherwise it falls back to the median provider. The standard deviation is measured across runs for that representative result. Cost, time, and token figures are per-task averages from the same runs. Best value is the cheapest Pareto-optimal model within 5 points of the top score.

These are the same measurements that power Exacto routing. See the docs for how routing works, or browse all models to try one.

API access

These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.

GET https://openrouter.ai/api/v1/benchmarks?source=openrouter
Authorization: Bearer <API key>

Use task_type=agentic to filter to tau_bench_verified_airline. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.