HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分47
AutoRef:面向智能体多参考图像生成的 Harness 自动优化
AI 导读
AutoRef 通过编码智能体迭代重写 harness 代码,在模型冻结的前提下自动优化多参考图像生成流程,并发现 AutoRef-Harness。该 harness 将开源 FLUX.2 [klein] 4B 在 MultiBanana 基准的四参考任务上从 5.72 提升至 7.37,追平或超过 Nano Banana Pro 和 GPT-Image-1.5 等闭源模型。
正文
Abstract:Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.
| Comments: | Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.35530 [cs.CV] |
| (or arXiv:2609.35530v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35530 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuta Oshima [view email]
[v1]
Mon, 28 Sep 2026 16:13:37 UTC (24,957 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org