AI

代理在付语言税(中英切换)

This post is currently available in Chinese only. The original text follows.Open the Chinese page →

代理在付语言税。

代理评测还把英语当成默认世界。任务绿了、榜单好看,就被当成「可以上线」。可真实用户不讲英语时,同一条任务会更贵、更爱在工具上空转,还更容易逃回英语前缀。语言不是皮肤,是运行时成本与失败形态。

代理的语言税:低资源语言输入可达 2×

同样任务,输入能翻一倍

Peng 等《BabelArena》(arXiv:2609.23490) 用 BabelFlow 把四套代理基准对齐到 23 种语言,得到 702 个任务身份、16,146 条实例。五款前沿模型在同一 VitaBench 任务上,低资源语言的总输入 token 是英语的 1.69–2.06×;固定前缀(系统提示、工具定义、首轮用户消息)甚至到 1.89–2.51×。轮次呢?只有 0.97–1.16×——贵在每轮上下文变厚,不在多聊几轮。

高资源组相对英语也要 1.10–1.23×。从高到低,五个模型的 Pass^1 与语言一致性都掉:LC 掉 3.8–14.8 个点。没有一家模型横扫四个基准族——Qwen-3.8-Max 常领跑准确率,语言一致性却经常落后;GPT-5.6-Terra 的 LC 更高。任务分与「说对方的语言」不是一回事。

换语言,先卡在工具上

换语言,先卡在工具上

低资源语言的失败形态也变了。抽样 1,600 条 VitaBench 失败轨迹:泰语、泰米尔语里,工具使用与控制流错误占比明显高于英语和中文——答案错了还在其次,先在调用与循环上翻车。购物规划(DP-Shop)更狠:低资源轨迹轮次 1.14–1.71×、工具调用 1.31–1.80×;产品搜索空结果从英语的 6.1–17.1% 飙到 53.7–68.7%,多出来的工具调用里 49.8–79.9% 就是反复搜空。

语言不一致时,91.2% 的切换目标是英语——常见形态是工具调用前的英语开场白:「Let me check…」。代理一紧张就逃回训练里最熟的语言,用户却还在用自己的母语等结果。

今晚:把语言税写进闸门

别再用英语评测单点绿当多语上线许可。发版前至少并排记三件事:相对英语的输入倍数、工具空转 / 空搜占比、语言一致性(尤其工具前缀)。按语言资源档抽检,不按「支持 N 种语言」的营销句;同一任务在低资源语言上多烧一倍 token 却更爱卡工具,就算没过。今晚先问一句:这条代理换一门语言,账单和失败方式换了没有?

判断很简单:代理在付语言税——英语绿了,不等于多语便宜,也不等于多语稳。

Agents pay a language tax.

Agent eval still treats English as the default world. A green English run and a pretty leaderboard get read as “ship it.” For users who do not speak English, the same task gets more expensive, stalls more often on tools, and slips back into English prefixes. Language is not a skin. It is runtime cost and a different failure shape.

The agent language tax: low-resource input up to 2×

Same task, up to 2× input

Peng et al., BabelArena (arXiv:2609.23490) adapt four agent benchmarks into 23 languages with BabelFlow — 702 task identities, 16,146 instances. On matched VitaBench tasks, five frontier models burn 1.69–2.06× English total input tokens in the low-resource group; constant prefix (system, tools, first user turn) hits 1.89–2.51×. Turns? Only 0.97–1.16×. The tax is denser context per turn, not longer chats.

Even the high-resource group sits at 1.10–1.23× English. From high to low resource, every model loses Pass^1 and language consistency — LC drops 3.8–14.8 points. No model sweeps all four families: Qwen-3.8-Max often leads accuracy yet lags LC; GPT-5.6-Terra leads LC. Task score and “speaking the user’s language” are decoupled.

Change the language; stall on tools first

Change the language; agents stall on tools first

Failure modes shift too. On 1,600 failed VitaBench trajectories, Thai and Tamil show larger shares of tool-use and control-flow errors than English and Chinese — wrong answers matter less than broken calls and unproductive loops. Shopping planning (DP-Shop) is sharper: low-resource runs use 1.14–1.71× turns and 1.31–1.80× tool calls; empty product searches jump from 6.1–17.1% in English to 53.7–68.7% in low-resource languages, and 49.8–79.9% of the extra calls are those repeated empty searches.

When language consistency breaks, 91.2% of annotated switches land in English — often an English preface before a tool call (“Let me check…”). Under stress the agent flees to its most familiar training language while the user is still waiting in theirs.

Tonight: put the language tax on the gate

Stop treating a green English eval as a multilingual ship permit. Before promote, log three numbers side by side: input multiple vs English, tool idle / empty-search share, and language consistency (especially pre-tool prefixes). Sample by resource tier, not by “supports N languages” marketing. Same task, double tokens, more tool stalls — that is a fail. Ask one question tonight: when this agent changes language, do the bill and the failure mode change with it?

The judgment is simple: agents pay a language tax — English green is not multilingual cheap, and not multilingual stable.