HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分59
研究量化野生 AI 生成网页文本对预训练的价值,提出新缩放律并开源 WildAI 语料
AI 导读
论文研究野生 AI 生成网页文本对语言模型预训练的影响,发现 FineWeb 过滤后 2026 年 6 月网页数据中 27.5% 的 token 被 Pangram 标为 AI 生成,8 月升至 31.1%。
正文
Abstract:Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at this https URL.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.40295 [cs.CL] |
| (or arXiv:2609.40295v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.40295 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jenna Russell [view email]
[v1]
Wed, 30 Sep 2026 17:50:22 UTC (1,545 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org