HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分37
用提示词分歧做偏好监督:无需像素级标注的开放词汇语义分割适配方法
AI 导读
研究者提出偏好引导适配框架,用二元偏好替代密集掩码监督,把不同提示词模板对同一图像产生系统性差异的"提示词分歧"现象转化为内置偏好监督来源。该方法从跨模板高不确定性区域挖掘局部偏好查询,并用 Region-Localized Preference Optimization(RLPO)配合一致性正则化适配 OVSS 模型。
正文
Abstract:Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences. Our code is available at this https URL.
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.34528 [cs.CV] |
| (or arXiv:2609.34528v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34528 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hyun-Kurl Jang [view email]
[v1]
Mon, 28 Sep 2026 08:00:20 UTC (4,862 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org