Benchmarks like this measure whether the task got done, but a lot of senior behavior is the messy integration layer. Whether the agent notices a tool call silently failed, whether it recovers when a server returns something unexpected. That rarely shows up in a pass/fail score. Does the benchmark capture any of the tool-use reliability, or just final task success?