跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分38

AdaTutoRank:面向 RAG 与深度研究的自适应辅导优化文档集重排序

AI 导读

AdaTutoRank 是一种集合级重排序器,采用自适应辅导优化(ATO)训练,在九项评分维度、三级层次结构下为冷启动提供银标签、为强化学习提供奖励、为蒸馏提供提示。

正文

View PDF HTML (experimental)

Abstract:Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.
Comments: Project Page: this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.32472 [cs.CL]
  (or arXiv:2609.32472v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.32472

arXiv-issued DOI via DataCite

Submission history

From: Kailin Jiang [view email]
[v1] Sat, 26 Sep 2026 11:05:34 UTC (11,983 KB)
[v2] Tue, 29 Sep 2026 03:10:15 UTC (11,983 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org