EMPVResearch NotesIT
EMPV / RESEARCH NOTESNOTE / 008

TECH NOTE / AI AGENTS

ThinkingBox evaluates agents by the final state of the system

The Microsoft and Hugging Face benchmark repeatedly runs 507 business workflows and checks databases and resulting side effects. An agent can finish without an error and report completion while still leaving the system in the wrong state.

Published2026-10-06 ยท 7 min
ThinkingBoxAI agentsMCPAgent evaluation
01 / The benchmark

ThinkingBox checks what remains in the system after the agent finishes.

On October 3, Microsoft and Hugging Face published a new overview of ThinkingBox, an environment and benchmark for evaluating agents that work on stateful business workflows. The accompanying paper, revised on October 1, describes 507 tasks across retail, auto insurance, travel, neobank and IT/HR support scenarios.

Each task starts from a defined initial state, exposes Model Context Protocol (MCP)-compatible tools and ends with executable checks over the backend state. The goal is not to decide whether the agent's answer sounds correct, but whether the database and resulting side effects actually match the requested outcome.

02 / A tool call is not an outcome

A plausible sequence of actions can still leave the wrong record behind.

The authors use a support ticket for a delayed delivery as an example. The agent checks the order, tracking, customer profile and policy, opens a ticket and then closes it as resolved. The sequence looks orderly, but the carrier still has an open exception: the correct ticket state should have been "hold", not "solved".

ThinkingBox therefore treats the agent trajectory as a claim and the terminal state as evidence. For tasks with structured effects, deterministic judges compare what changed with the required end state and reject wrong values, missing actions or unintended extra changes.

03 / Reliability

Succeeding once and succeeding every time are different metrics.

The 507 tasks are repeated 20 times from a clean state. The benchmark separates average single-run success, whether a task can be solved at least once, and how many tasks are completed correctly in all 20 recorded attempts.

That distinction changes how the results read. In the report, Claude Opus 5.5 reaches 67.16% pass@1 but completes 241 of 507 tasks in all 20 runs, or 47.53%. Kimi-K3 solves 476 of 507 tasks at least once but only 68 in all 20 attempts. The benchmark is not only asking what a model can do; it measures how often the same workflow stays correct when repeated.

04 / Silent failures

Many failures do not produce an error signal that the surrounding system can rely on.

In a common-set analysis across 12 models, the authors report 121,680 valid trials. Of those, 79,853 fail the executable checks. Yet 67.24% of the failures still terminate cleanly, include at least one state-changing action and report no final tool error.

Within those failures, executable checks find wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%; the categories overlap. A monitor that only watches exceptions, formally valid tool calls or the agent's final text can therefore mark work as successful even when the system of record shows an incomplete or incorrect outcome.

05 / Operational implications

For agents that write to business systems, the success criterion should live outside the agent.

The practical design lesson is to separate execution from verification. If an agent updates a CRM, ticketing system, order record or customer master, the outcome can be checked by reading the authoritative state after the action and comparing it with explicit conditions, rather than treating the model's final message as proof.

The authors also suggest classifying tool and system errors so retries target recoverable failures, reducing the tool surface to what the workflow needs, and keeping human approval for changes that are difficult to reverse. ThinkingBox makes these choices measurable because each attempt starts in isolation and exposes the effects it produced.

06 / What it shows and what it does not

It is a reproducible benchmark, but the public workflows are still synthetic reconstructions.

The 507 tasks model enterprise patterns, but the post states that every public workflow is a synthetic reconstruction and does not contain real customers. In addition, 477 tasks are graded on state alone, while 30 add a narrow binary response rubric where a requirement cannot be expressed as a simple database value.

The advantage is that the environment, data and checks are inspectable. The ThinkingBox framework is open source, the v1.0 benchmark is released separately, and its OpenEnv integration can run episodes and return a pass/fail result.

For a company, this does not replace evaluation on its own processes. It does provide a concrete criterion to bring into internal tests: define the correct final state first, then verify whether the agent reaches it repeatedly.

EMPV / TAKEAWAY

If an agent changes business systems, "done" is not evidence: the result should be read from the final system state and verified more than once.

EMPV / SHARE

Share this Research Note.

Research Notes