Home ﹥ Hot News > Artificial Intelligence > Validation > V24.9 Brain-X τ-bench report 2026-07-24
Links:https://github.com/u7490637/XRM-SSD/blob/main/Test%20results ...
Tau-bench AI Agent Evaluation Benchmark - V24.9
Brain-X Phase 123 Physical Edition: Test Report Summary
Test Overview
-
Test Name: tau-bench AI Agent Evaluation Benchmark - V24.9 Phase 123 Physical Edition
-
Timestamp: 2026-07-24T12:41:25.942211
-
Platform & Agent: Windows, PhysicalAgent (Physical Mode Enabled)
-
Total Tasks & Iterations: 45 tasks across 5 iterations
-
Total Duration: 8.39 seconds
Performance Results
-
Success Count: 44 / 45 tasks
-
Success Rate: 97.78%
-
Average Score: 0.89 / 1.0
-
Average Steps: 33.64 steps per task
-
Total Tool Calls: 57
Domain Statistics
-
Retail: 15 tasks total, 14 successful (93.33% success rate, avg score: 0.86, avg steps: 30.0)
-
Airline: 10 tasks total, 10 successful (100% success rate, avg score: 0.90, avg steps: 33.2)
-
Banking: 10 tasks total, 10 successful (100% success rate, avg score: 0.90, avg steps: 35.3)
-
Travel: 10 tasks total, 10 successful (100% success rate, avg score: 0.90, avg steps: 37.9)
⚙️ Physical Metrics & Tool Statistics
-
Tool Call Latency: Average 58.27 ms (ranging from 30 ms to 80 ms)
-
Tool Call Jitter: 13.90 ms
-
Error Rate: 0.0% across 56 tool executions (with only 1 isolated service unavailable error encountered in a retail task)
-
Signal Interrupts & Checkpoint Restores: 0