Jump to channels Log in
‹ AI Model Releases
0

The Agent Said It Was Done. The Database Disagreed. huggingface.co Wildcard

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on whether they actually change database records correctly, not just whether they make valid tool calls or sound right. Testing 12 language models on 507 stateful business tasks run 20 times each, they found that 67% of failed attempts still produced clean tool calls with no reported errors—revealing wrong field values, unintended side effects, or incomplete actions. Only three models (GPT-6 Astra, Claude Opus 5.5, Claude Opus 5) retained most of their single-attempt accuracy across repeated trials, while others like Kimi-K3 solved more tasks at least once but failed consistently, and roughly 80% of failures stemmed from tool-handling errors rather than reasoning problems.

Log in to comment.

0 comments

No comments yet.