Microsoft open-sources ThinkingBox: 507 business workflows, 20 runs each, graded on final database state
An agent saying it's done doesn't count; the records it leaves do. How it measures, what it found, and what dependable completion actually costs.
// Key points
- Covers 507 stateful business workflows, each run 20 times per model, graded on terminal backend state and side effects rather than the agent's own text.
- Among failed runs that ended cleanly with no tool error, executable checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%.
- Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 attempts; Claude Opus 5 solves 79.09% at least once and 47.53% on every attempt.
- Framework (MIT) and benchmark data are on GitHub; by its cost-per-dependable-task metric, GPT-5.4 is lowest at an estimated $6.80.
Builder's takeSingle-attempt pass rates flatter agents. I'm applying the same idea to scoring in my AI Interview product: run the same answer sheet several times and check that what lands in the database agrees, not what the model says. When picking a model, put all-20-pass rate and cost per dependable task in your eval sheet; it tells you more than a leaderboard.