跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分44

研究揭示 LLM 全流程为何默认偏向美式英语:1,813 组 AmE–BrE 变体与 DiAlign 方法

AI 导读

研究以 1,813 组匹配的美式英语(AmE)与英式英语(BrE)变体为资源,提出免训练的 DiAlign 方法,从数据分布证据估计区域对齐度。AmE 在全部 6 个预训练语料和 21 个后训练数据集中均被偏好,tokenizer 表示更紧凑、预测成本更低,且在中性英语提示下仍是默认生成偏好;改用英式英语提示可向 BrE 偏移,但无法稳定消除 AmE 默认。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose "English (US)" as a primary English setting despite the global diversity of English. We ask: How does "English (US)" become the default? We study this question as structural bias, examining how geopolitical histories of data curation, digital dominance, and linguistic standardization intersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six audited pretraining corpora and 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study of structural bias across major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.
Comments: Preprint
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
Cite as: arXiv:2604.04204 [cs.CL]
  (or arXiv:2604.04204v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2604.04204

arXiv-issued DOI via DataCite

Submission history

From: Mir Tafseer Nayeem [view email]
[v1] Sun, 5 Apr 2026 17:59:34 UTC (2,505 KB)
[v2] Mon, 28 Sep 2026 00:34:52 UTC (3,097 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org