HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-08-03精选AI 评分74
UEmbed:统一稀疏与稠密的多模态嵌入模型
AI 导读
UEmbed 是一种仅解码器的多模态嵌入模型,可在单次因果前向传播中同时生成稀疏词级和稠密表示,通过可学习特殊 token 与词汇表分区突破单 token 信息瓶颈。
推荐理由
UEmbed 让一个解码器模型同时输出稀疏与稠密多模态嵌入,稀疏模式精度接近稠密且原生适配倒排索引,对兼顾效率和精度的检索系统,省去了维护多套模型的开销。
正文
Abstract:Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass. UEmbed appends N learnable special tokens to the input and partitions the vocabulary into N disjoint subsets. Each token's causal hidden state predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. Trained on public data, we release UEmbed at 2B, 4B, and 9B scales. UEmbed-9B reaches 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming multimodal embedding models trained on publicly available data (e.g., RzenEmbed). On BEIR, UEmbed also remains competitive with strong dense and sparse baselines. Furthermore, we demonstrate the practical utility of UEmbed across three dimensions: effectiveness, efficiency, and agentic applications. Overall, UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2608.02583 [cs.CV] |
| (or arXiv:2608.02583v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.02583 arXiv-issued DOI via DataCite |
Submission history
From: Tingyu Song [view email]
[v1]
Mon, 3 Aug 2026 17:54:11 UTC (5,799 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org