AI

会走路,不等于会绕路(中英切换)

会走路,不等于会绕路。

代理评测还在迷信一条干净的工具链:工具全绿、路径笔直,过关就算会用工具。真实线上是另一回事——超时、404、库存负数、语义脏数据。会走路与会绕路,是两条能力。

会走路,不等于会绕路

隐式故障,信任先塌

Zhu 等《ToolMaze》(arXiv:2606.05806) 用 270 个工具、400 条基任务 ×(干净 + 四种扰动)共 2000 条实例,把故障拆成显式/隐式 × 瞬时/永久(P1–P4),拓扑从线性到多分支 DAG(C1–C4)。核心发现不是「有故障就掉分」,而是 隐式语义故障 把代理打穿:在带提示的设置下,隐式相对显式的 Perturbation Recovery Rate(PRR)平均再掉约 37.15%(瞬时约 53.75%,永久约 20.54%)。结构合法、语义脏——代理照单全收,毒值一路传。

Claude-Sonnet-4-6 在非扰动(NP)上 TSR 77%,一进 P1–P4 仍大幅塌;P4(隐式永久)上,跨模型平均往往 PRR <20%、Recovery Cost >70%——不是偶尔失手,是系统性过信脏输出。

规模在抬走路,绕路几乎不动

开源权重上,非扰动任务成功率 TSR(NP) 每升一个数量级参数约涨 17.85 个百分点;PRR 只涨约 4.88 pp——故障容忍随规模改善大约慢 3.66×。会走路会随规模上来;会绕路不会自动跟上。失败感知提示能抬 +1.5% 到 +20.8%,但扰动下的 TSR 仍远低于 NP——提示是补丁,不是替代品。

今晚:在注入故障下看恢复

别再用干净路径通过率单独当 promote/kill。发版闸门并列挂 注入工具故障后的恢复分(尤其静默/语义类):显式瞬时、显式永久、隐式瞬时、隐式永久各抽一批;记 PRR 与恢复成本,不只记 pass。今晚先问一句:这条代理在工具撒谎时,还会不会绕路?

判断很简单:会走路,不等于会绕路——闸门看恢复,不只看 happy path。

Happy path is not recovery.

Agent eval still worships a clean tool chain: green tools, straight path, pass means “can use tools.” Production is the other story — timeouts, 404s, negative inventory, semantically dirty payloads. Walking the happy path and recovering when tools fail are decoupled skills.

Happy path is not recovery.

Implicit faults; trust collapses first

Zhu et al., ToolMaze (arXiv:2606.05806) run 270 tools and 400 base tasks × (clean + four perturbation modes) = 2000 instances, splitting failures into explicit/implicit × transient/permanent (P1–P4) on DAGs from linear to multi-branch (C1–C4). The punchline is not “noise hurts scores.” It is that implicit semantic failures break agents: with failure-aware prompts, the implicit-vs-explicit Perturbation Recovery Rate (PRR) gap averages about 37.15% (53.75% transient, 20.54% permanent). Structurally valid, semantically wrong — agents trust it and poison the rest of the plan.

Claude-Sonnet-4-6 hits 77% TSR on the non-perturbed (NP) path and still collapses under P1–P4; under P4 (implicit permanent), average PRR often sits under 20% with Recovery Cost over 70% — not bad luck, systemic over-trust in corrupted outputs.

Scale lifts walking; recovery barely moves

On open-weight models, non-perturbed task success TSR(NP) rises about 17.85 pp per order of magnitude in parameters; PRR rises only about 4.88 pp — fault-tolerance improves roughly 3.66× slower than clean execution. Walking scales; recovery does not hitch a free ride. Failure-aware prompts help (+1.5% to +20.8%), yet perturbed TSR stays far below NP — a hint is a patch, not a substitute.

Tonight: score recovery under injected faults

Stop promoting or killing on clean-path pass alone. Hang a recovery score under injected tool faults (especially silent/semantic) beside any happy-path TSR: sample explicit-transient, explicit-permanent, implicit-transient, implicit-permanent; log PRR and recovery cost, not only pass. Ask one question tonight: when the tool lies, does this agent still find a detour?

The judgment is simple: happy path is not recovery — the gate scores recovery, not only the clean walk.