HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分47
让 LLM 学会从上下文学习:基于扰动公开文档的合成训练方法
AI 导读
研究者构建了一套无需人工标注的合成流水线,通过扰动改写公开文档、生成需依赖文档推理的问题与评分标准,从 3.5k 篇文档生成约 10k 样本,用于训练学生模型。
正文
Abstract:Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.
| Comments: | 18 pages, 3 figures, 10 tables |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7; I.2.6 |
| Cite as: | arXiv:2609.33642 [cs.CL] |
| (or arXiv:2609.33642v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33642 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chengyue Jiang [view email]
[v1]
Sun, 27 Sep 2026 15:05:22 UTC (271 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org