终点过了,不等于过程对了。
代理跑完一个电脑任务,最容易被验收的是最后那个交付物:文件在不在、字段对不对、功能验过没。真正难的是路上怎么走——子目标有没有按顺序完成、有没有点偏、有没有在已经做完之后继续空转。终点绿了,只说明交卷对了;过程对了,才说明路没白走。

同一模型,终点与过程差一整截
NVIDIA《OSWorld-Pro: Process-based Evaluation for Computer Use Agents》(arXiv:2609.24890) 把电脑使用代理的验收从「最终交付物」拆成 305 个任务、2,814 个渐进子目标,并用 67,264 条人工标注压过程。任务平均 9.2 个串行子目标、跨 3.45 个应用(OSWorld 大约 1.34)。
硬数字在同一模型上就撕开了。Minimax M3 在 OSWorld 终点评测 75.2%,到 OSWorld-Pro 过程评测只剩 28.9%。Claude Opus 5 从 83.4% 掉到 75.7%。开源权重顶端在 OSWorld 常 >80%,在 OSWorld-Pro 最好也只到 55.1%(Qwen 3.8 Flash Next)。终点绿,不是过程绿。
终点看不见:空转、点偏、绕路

过程评测把终点验收看不见的病摊开。Claude Opus 5 有一次在子目标都完成后,仍做了 59 步与子目标无关的动作。点击推进率:Minimax M3 / Kimi K3 大约 39.0–41.9%,强模型大约 72.4–92.5%——想关窗却点不到 ×。
效率也差一整截。成功完成可行子目标的平均步数:GPT-5.6-Sol 4.9,Minimax M3 21.2。碰到不可行子目标还不放手:Opus 5 9.5 步就认了,Sonnet 5 拖到 39.8。两边都能「交卷」,路上的执念差 4×。
今晚:别只验交付物
发版前别只跑功能验收脚本。并列记:全程子目标完成率(不是终点分)、成功子目标平均步数、有没有子目标完成后的空转、点击 / 键盘类动作推进率。交付物绿、路上空转 59 步——算没过。今晚先问一句:这条代理交的是看得见的终点,还是走得对的过程?
判断很简单:终点过了,不等于过程对了——交卷过关,不是路上过关。
Crossing the finish is not getting the process right.
When an agent finishes a computer-use task, the easy acceptance check is the final deliverable: file present, fields correct, functional verifier green. The hard part is how it got there — whether sequential subgoals completed in order, whether clicks landed, whether it kept looping after the work was already done. Finish-line green only means the hand-in passed. Process right means the path was earned.

Same model; outcome and process tear apart
NVIDIA, OSWorld-Pro: Process-based Evaluation for Computer Use Agents (arXiv:2609.24890) splits computer-use acceptance from final deliverables into 305 tasks and 2,814 progressive subgoals, backed by 67,264 human annotations. Tasks average 9.2 sequential subgoals across 3.45 apps (OSWorld about 1.34).
The hard cut appears on matched models. Minimax M3 scores 75.2% on OSWorld outcome evaluation, then 28.9% on OSWorld-Pro process evaluation. Claude Opus 5 drops from 83.4% to 75.7%. Top open-weight models often clear >80% on OSWorld, yet peak at 55.1% on OSWorld-Pro (Qwen 3.8 Flash Next). Finish-line green is not process green.
What the finish line cannot see: loops, misses, detours

Process scoring surfaces failures outcome checks miss. Claude Opus 5 once ran 59 subgoal-irrelevant steps after all subgoals were already complete. Click progression likelihood: about 39.0–41.9% for Minimax M3 / Kimi K3, versus about 72.4–92.5% for stronger models — intending to close a window and missing the ×.
Efficiency gaps the same way. Steps per successfully completed feasible subgoal: GPT-5.6-Sol 4.9, Minimax M3 21.2. On infeasible subgoals: Opus 5 gives up around 9.5 steps; Sonnet 5 persists to 39.8. Both can hand in; stubbornness differs by 4×.
Tonight: do not only verify the deliverable
Before promote, do not only run the functional verifier. Log side by side: full subgoal completion rate (not the outcome score), mean steps per successful subgoal, whether post-completion irrelevant loops appear, click / keyboard progression rates. Green deliverable with a 59-step idle loop is a fail. Ask one question tonight: is this agent shipping a visible finish line — or a path that was walked correctly?
The judgment is simple: crossing the finish is not getting the process right — hand-in pass is not path pass.