HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分41
CorpusMap:面向智能体检索的实体语料地图
AI 导读
研究者提出 CorpusMap,一种围绕语料中反复出现的实体组织文档的导航层,将每个实体表示为 Entity Page,聚合并链接所有提及该实体的文档,形成可供智能体遍历的实体—文档图。
正文
Abstract:Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37226 [cs.CL] |
| (or arXiv:2609.37226v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37226 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Soyeong Jeong [view email]
[v1]
Tue, 29 Sep 2026 10:39:57 UTC (205 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org