Microsoft researchers have released ThinkingBox, an open sandbox and benchmark that grades AI agents on what they actually changed in a backend database, not on what they said. The headline result is uncomfortable: across 121,680 trials, 79,853 attempts failed, and 67.24 percent of those failures ended without any final tool error. Even the best models, Claude Opus 5.5 and Opus 5, passed all 20 attempts on only 241 of 507 tasks (47.53 percent). The Hugging Face write-up puts the thesis in one line: "A trajectory is a claim. Database state is the evidence."
If you are building agents that touch orders, claims, bookings, or tickets, this is the most practical evaluation idea of the week. It is also a useful companion to our complete guide to AI benchmarks and to the Epoch Capabilities Index result for Claude Opus 5.5, because it measures something those aggregate scores do not: consistency on real workflows.

TL;DR: ThinkingBox in one table
| Question | Answer |
|---|---|
| Who built it? | Microsoft researchers, published with Hugging Face (blog) |
| What does it measure? | Whether the final database state and side effects match the required end state |
| How big is it? | 507 tasks, each run 20 times on isolated, freshly reset backends |
| Domains | Retail, travel and hospitality, auto insurance, neobank IT, consulting IT/HR |
| Headline numbers | 121,680 trials; 79,853 failed attempts; best models pass 20 of 20 on 241 of 507 tasks |
| Top pass@1 | Claude Opus 5.5 at 67.16 percent, per the Hugging Face post |
| Is it free? | Framework is MIT licensed on GitHub; dataset on Hugging Face |
| Paper | arXiv 2608.19741, "One Success Isn't Reliability" |
| Do I need Docker? | Yes, plus Python, a Typesense database, MCP servers, and model endpoints |
Why grading the transcript is not enough
Most agent benchmarks, and most home-grown evals, score one of two things: whether the final message sounds right, or whether the agent called the expected tools. Both are proxies. An agent can say "I have updated your reservation" after writing the wrong date, or call the correct refund tool with the wrong amount, or touch a second record it was never supposed to modify.
ThinkingBox removes the proxy. Each task defines a required end state, and deterministic judges check the terminal database and any side effects against it. The paper's authors argue that existing benchmarks miss the messy parts of real work: gathering information, following policy, coordinating tools, and keeping state correct.
The failure breakdown in the Hugging Face post is what makes this actionable. Among the 79,853 failed trials, roughly 77.6 percent had wrong field values, 43 percent created unintended side effects, and 25 percent missed required changes. These overlap, since one trial can fail in several ways. By root cause, 79.9 percent traced to tool handling, 10.3 percent to wrong updates, 7.0 percent to incomplete resolution, and 2.9 percent to no action at all.
Tool handling dominating the list matches what builders already see. If you want to reduce that class of error, the design of the tool surface matters as much as the model, which is the theme of our post on MCP tool descriptions and selection reliability.
How the benchmark works
The setup is deliberately simple to reason about.
- Each of the 507 tasks is a business workflow, such as modifying a retail order, filing an insurance claim, rebooking travel, or resolving an IT ticket.
- Tools are exposed as MCP servers, so any agent that speaks the Model Context Protocol can be plugged in.
- Every run gets an isolated, freshly reset backend, so one attempt cannot contaminate the next.
- After the agent stops, deterministic judges compare the database to the required end state and look for forbidden side effects.
- Each task is repeated 20 times, and results are reported three ways.
The three metrics are the real contribution:
| Metric | Meaning | What it tells you |
|---|---|---|
| pass@1 | Share of single attempts that succeed | Typical first-try quality |
| pass@20 | Tasks solved at least once in 20 tries | Capability ceiling, what the model can do on a good day |
| observed 20 of 20 | Tasks that pass every recorded attempt | Dependability, what you can promise a customer |
Reporting all three exposes the gap between "can do" and "reliably does".
What the numbers say about today's models
The Hugging Face post reports that Claude Opus 5.5 leads at 67.16 percent pass@1 overall. The same post notes that Kimi-K3 solves 93.89 percent of tasks at least once but only 13.41 percent consistently. That is a huge spread: a model that can eventually solve nearly everything is dependable on a small fraction.
The paper abstract frames the same pattern with slightly different figures for the models it names: Claude Opus 5 falls from 66.50 percent (pass@1) to 47.53 percent (20 of 20) as the bar moves from one attempt to 20 of 20, and Kimi-K3 falls from 57.37 percent to 17.60 percent. The decimals differ between the two write-ups, likely because of run configuration or version naming, so read the paper's tables before quoting a specific cell. The direction is the same everywhere: reliability collapses as you demand repeatability.
Domain matters too. Retail averages 59.52 percent pass@1, while auto insurance averages 33.83 percent. Insurance workflows carry more policy rules and more fields that must be right, which is exactly the sort of task a business will actually automate.
What it costs to get a dependable result
The post also estimates cost per dependable task: about $6.80 for GPT-5.4, $7.45 for GPT-6 Astra, and $7.80 for Claude Opus 5.5. These are the Hugging Face post's numbers, and they depend on how many retries a dependable result needs. The practical point is that sticker price per token and cost per reliable outcome are different things. A cheaper model that needs more attempts, or more human cleanup, can cost more per finished task. For background on how fast model prices move, see our GPT-5.6 launch breakdown.
Why "terminated cleanly" is the dangerous signal
The paper's third claim deserves attention: many failed trials end cleanly after valid state-changing actions. In a dashboard, that looks like success. The agent did not crash, did not loop, did not hit a rate limit, and produced a confident summary.
This matters because most production monitoring watches for errors and timeouts. A silent wrong write produces neither. You find out later, from a customer, an auditor, or a reconciliation job. ThinkingBox turns that hidden class of failure into a number you can track.
What people are asking
Is this just another leaderboard?
Not quite. Leaderboards rank models on a score. ThinkingBox is closer to a test harness: the environment, the tasks, and the judges are open, so you can add your own domain and run your own agent. That makes it a template for building evals, not only a place to compare names. For another example of benchmark design aimed at agents, see Terminal-Bench 2.0 and Agents' Last Exam.
Does 47.5 percent mean agents are useless?
No. It means unsupervised, fully autonomous completion of multi-step stateful workflows is not yet dependable for most tasks, at least in this benchmark's setup. Many real deployments add human review for high-risk writes, narrow the tool surface, or retry with verification. The benchmark gives you a way to measure how much those mitigations help.
Are the judges perfect?
Deterministic state checks are far less noisy than an LLM grader, but they are only as good as the required end states someone wrote. A task with an overly strict spec could mark a valid alternative solution as a failure, and a loose spec could miss a harmful side effect. The authors' choice to check side effects explicitly is a strength, but it is worth sampling failures by hand before trusting any aggregate.
How comparable are the cross-model numbers?
Agent scaffolding, prompts, and tool descriptions all shift results. Treat the model ranking as a snapshot under one harness, not a universal ordering.
How to apply the idea to your own agents
You do not need ThinkingBox to steal its method. A minimal version for any agent that writes to a database looks like this:
- Write the end state first. For each workflow, list the exact rows and fields that must change, and the rows that must not.
- Snapshot before and after. Reset to a known fixture, run the agent, then diff the data.
- Check side effects. Assert that unrelated tables and records are unchanged.
- Repeat. Run each task at least 10 to 20 times and report pass@1 and all-runs-pass separately.
- Bucket failures. Tag each failure as wrong value, extra side effect, missed change, or no action. Fix the biggest bucket first, which per this benchmark is usually tool handling.
A starter prompt for generating the checks with a coding agent:
Read our order-management tool schemas. For each workflow in /evals/workflows.md,
write a deterministic checker that loads the fixture database, runs the agent
transcript's tool calls, then asserts (1) required field values, (2) no changes
to unrelated records, (3) no duplicate writes. Output a pytest file.
If your agents have broad tool access, pair this with the MCP security guide: a state-based eval catches accidental damage, and a security review catches deliberate abuse.
Running ThinkingBox yourself
The Hugging Face post points to the OpenEnv environment under envs/thinkingbox_env, the microsoft/ThinkingBox-Bench dataset, and the microsoft/thinkingbox repository. The post says the setup needs Python 3.11 or newer, Docker, a Typesense database, MCP servers, and an OpenEnv server, then episodes are scored from the command line. The framework README describes a tb command line with subcommands for starting the session proxy, running inference against a dataset, a terminal UI, and aggregate metrics. At the time of writing the repository showed about 85 stars and 16 commits on main, so expect rough edges and read the issues before betting a pipeline on it.
Limitations and open questions
- Synthetic domains. The 507 workflows are well designed but still stand-ins for your company's systems.
- Single harness. Results reflect one agent loop and tool setup. A better scaffold may raise numbers.
- Cost estimates are indicative. The per-dependable-task figures come from the benchmark's own runs and pricing at that time.
- Version naming is inconsistent across sources. The blog and abstract refer to Opus models slightly differently, which is why we avoid quoting a single canonical decimal.
What this means for what you build
Treat reliability as its own metric, separate from capability. If a task must succeed every time, measure observed all-runs-pass, not pass@1 or pass@20. Grade the database, not the chat. And budget for verification steps, because the benchmark shows that clean-looking runs hide wrong writes.
For deeper background, start with the agent fundamentals guide, then the benchmarks guide.
Related reading
- AI benchmarks: the complete guide (2026)
- Claude Opus 5.5 tops the Epoch Capabilities Index
- Terminal-Bench 2.0 explained
- Agents' Last Exam benchmark
- MCP tool descriptions and selection reliability
- MCP security guide
- What are AI agents? The complete guide
Primary sources: Hugging Face blog, arXiv paper, GitHub repository.
Figures and repository details are accurate as of October 4, 2026 and may change as the authors update the benchmark.
