长得像,不等于跑得通。
代理复刻一个应用时,最容易交出一张「看起来对」的壳:按钮在、布局齐、首屏不崩。真正难的是点下去之后——状态怎么变、Undo 回不回得去、算出来的数对不对。长得像,可能只是静态结构过关;跑得通,才是行为过关。

结构最容易过,交互落下一大截
Alibaba Token Hub《RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents》(arXiv:2609.22000) 用 RecreationBench 做留出评测:250 个任务,Ubuntu / macOS / Windows / Android / Web 各 50。隐藏用例分 Prog(程序化断言)与视觉断言;参考应用本身是 oracle。
结果里有一条很硬的分界:代理复现静态界面结构,远比复现交互与计算输出靠谱。在原生平台的 16 组模型×平台对比里,structure 通过率最高的占 15 / 16;最弱的一类永远是 button、computation 或 interaction,相对 structure 落后 9.1–34.4 个百分点。Web 用另一套分类,静态内容仍压过交互 28.2–36.7 个百分点。
具体到 Ubuntu 上的 PhotoCollage:照片转过再 Undo,复刻版能把排列摆回去,照片本身还停在旋转态——首屏可以长得对,历史状态已经错了。
总分 58%,全过只有 2.8%

GPT-6 Astra 总分领先约 58.1%(文中亦写 58.06%),但 100% Prog 全过的任务只有 2.8%;其余模型 ≤ 0.8%。把门槛降到 90% Prog:Astra 17.6%,Claude Opus 5 5.5%,GLM-5.3 与 Qwen3.8-Max-0902 各 2.0%。平均分不是「这个应用跑通了」的概率。
交付物也偏瘦、偏整块:89.4% 比参考更小,recreation-to-reference LOC 中位比 16.9%;92.3% 源文件更少,83.8% 更集中在最大那个文件。最后一步闭环——最终编辑 → 重启动 → 再检查——只在 23.6–47.5% 的轨迹里完成(Claude Opus 5 23.6% → Qwen3.8-Max-0902 47.5%)。没闭环,就更难发现「长得对、跑不通」。
今晚:别用首屏当验收
发版前别只截一张「看起来齐」的图。并列记:Prog 全过率(不是总分)、交互 / 计算类断言是否过、最终 edit 之后有没有 relaunch 再看。首屏绿、Undo 后状态错——算没过。今晚先问一句:这条代理交的是看得见的壳,还是点得动、回得去的状态机?
判断很简单:长得像,不等于跑得通——静态结构过关,不是行为过关。
Looking right is not running right.
When an agent recreates an app, the easy deliverable is a shell that looks right: buttons present, layout tidy, first screen does not crash. The hard part is what happens after you click — how state changes, whether Undo actually undoes, whether computed values are correct. Looking right can mean static structure passed. Running right means behavior passed.

Structure passes; interaction trails far behind
Alibaba Token Hub, RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents (arXiv:2609.22000) holds out RecreationBench: 250 tasks, 50 each on Ubuntu, macOS, Windows, Android, and Web. Hidden suites mix Prog (programmatic) and visual assertions; the running reference is the oracle.
One hard cut shows up in outcomes: agents reproduce static interface structure more reliably than interactions and computed outputs. Across 16 model–platform native comparisons, structure has the highest mean pass rate in 15 / 16; the weakest category is always button, computation, or interaction, trailing structure by 9.1–34.4 percentage points. On Web’s different taxonomy, static-content still beats interaction by 28.2–36.7 pp.
Concretely, Ubuntu PhotoCollage: after a rotate and Undo, the recreation restores the arrangement but leaves the photo rotated — the opening screen can look fine while history state is already wrong.
~58% overall; 2.8% full Prog pass

GPT-6 Astra leads overall at about 58.1% (paper also reports 58.06%), yet passes all programmatic tests on just 2.8% of tasks; every other model ≤ 0.8%. At 90% Prog coverage: Astra 17.6%, Claude Opus 5 5.5%, GLM-5.3 and Qwen3.8-Max-0902 2.0% each. An average is not the probability that an application is fully reconstructed.
Delivered apps stay smaller and more monolithic: 89.4% smaller than reference LOC (median recreation-to-reference ratio 16.9%); 92.3% fewer source files; 83.8% more concentrated in the largest file. The final edit → relaunch → inspect loop completes in only 23.6–47.5% of trajectories (Claude Opus 5 23.6% → Qwen3.8-Max-0902 47.5%). Without that loop, “looks right, runs wrong” stays easy to miss.
Tonight: do not accept the first screen
Before promote, do not only grab a tidy screenshot. Log side by side: full Prog pass rate (not the overall mean), whether interaction / computation assertions pass, whether the final edit was relaunched and inspected. Green first screen with wrong Undo state is a fail. Ask one question tonight: is this agent shipping a visible shell — or a state machine you can click and reverse?
The judgment is simple: looking right is not running right — static structure is not action-conditioned behavior.