跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

IDSpect:用 IDS 分解与细粒度奖励提升中文文本渲染准确率

AI 导读

针对 OCR 奖励把汉字当作原子字符、忽略部件与空间结构的问题,研究者提出 IDSpect:用表意文字描述序列(IDS)将目标文本确定性分解为 IDS token,并把裁剪区域的视觉 IDS 预测与目标序列对齐,配合整字语义奖励提供细粒度反馈。

正文

View PDF HTML (experimental)

Abstract:Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.37569 [cs.CV]
  (or arXiv:2609.37569v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.37569

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xingsong Ye [view email]
[v1] Tue, 29 Sep 2026 13:37:32 UTC (3,880 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org