跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分44

合成预预训练在规模扩展下依然有效,但并非源于语法先验

AI 导读

一项覆盖 500M 至 7B 四种参数规模、PT 预算最高 100B token 的研究显示,合成非自然语言数据上的预预训练(PPT)带来的下游性能与 token 效率提升可随规模保持,在 3B 规模下至少节省 21B PT token。但与先前工作不同,研究未发现一致证据表明这些收益来自语法先验,收益实际来自改善长程检索的 PPT 任务,且仅在 PT 数据混合中缺少网页文本时才会减弱。

正文

View PDF HTML (experimental)

Abstract:Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
Comments: Preprint. Under review
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.39827 [cs.CL]
  (or arXiv:2609.39827v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.39827

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Atsuki Yamaguchi [view email]
[v1] Wed, 30 Sep 2026 14:27:37 UTC (168 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org