Introducing ITSMBench

ITSMBench tests whether AI can actually close IT tickets, not just talk about them. Every task runs inside a real enterprise environment: dozens of connected systems, thousands of records, the same complexity a service desk deals with daily. To pass, an agent doesn’t just name the problem. It has to trace the actual root cause, follow the runbook, fix it, and leave every system it touched clean and consistent.

Leaderboard

ITSMBench score

claude-opus-5
glm-5.2
gpt-5.6-sol
gpt-5.6-terra
grok-4.5
0%20%40%60%80%100%$0.00$0.50$1.00$1.50$2.00Avg cost per task
ModelPass@1Avg costOut tok
1
claude-opus-5 [high]
50.56%±1.81$1.5219.7K
2
claude-opus-5 [medium]
50.56%±1.84$1.1113.8K
3
grok-4.5 [high]
45.39%±2.47$0.7117.4K
4
gpt-5.6-sol [xhigh]
43.82%±2.43$1.6912.8K
5
claude-opus-5 [low]
41.01%±2.07$0.779.6K
6
gpt-5.6-sol [high]
40.90%±2.31$1.269.7K
7
grok-4.5 [medium]
40.00%±2.74$0.4612.6K
8
gpt-5.6-sol [medium]
31.24%±2.27$0.796.3K
9
gpt-5.6-terra [xhigh]
30.11%±2.82$0.6512.1K
10
grok-4.5 [low]
29.66%±1.92$0.307.2K
11
gpt-5.6-terra [high]
28.99%±2.14$0.457.8K
12
gpt-5.6-sol [low]
24.04%±1.99$0.443.7K
13
gpt-5.6-terra [low]
20.22%±1.80$0.244.1K
14
gpt-5.6-terra [medium]
20.00%±2.14$0.305.1K
15
glm-5.2 [high]
19.78%±2.28$0.2621.9K
16
glm-5.2 [xhigh]
19.10%±1.80$0.3332K
17
gpt-5.6-sol [off]
17.30%±2.07$0.493.2K
18
gpt-5.6-terra [off]
11.24%±1.59$0.242.6K

For a fair comparison, the leaderboard runs every model through the same open-source Pi harness. Running models through their own harness (like Codex or Claude Code) changes both score and cost — the technical report contains more details about this.

Sample task

Prompt

Nadia Rahman on the finance team just posted in the #it-helpdesk Slack channel: CrowdStrike Falcon threw a malware detection on her workstation this morning, right after she opened a file she thought was a vendor invoice. She says the machine feels fine now and wants the alert looked at and closed out.

You're the IT responder on shift. Pick this up and handle it in the systems we operate. Work out what actually happened, contain anything that needs containing, and get the environment back to a safe state before you consider this resolved.

Built in collaboration withNew Measure