跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-14精选AI 评分74

Boogu-Image-0.1 发布:开源统一多模态理解与生成模型,训练成本仅约 40 万美元

AI 导读

Boogu-Image-0.1 系列开源统一多模态理解与生成模型发布,包含 Base、Turbo、Edit 和 Edit-Turbo 四个变体,支持高质量文生图、快速推理、指令编辑及中英双语文本渲染。

推荐理由

开源多模态模型终于有一个能打闭源系统的了,40万美元训练成本刷到接近GPT-Image-2的水平,想自己部署图像生成和编辑产品的开发者可以认真看看。

正文

Authors:Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S.Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei

View PDF HTML (experimental)

Abstract:We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.13125 [cs.CV]
  (or arXiv:2607.13125v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2607.13125

arXiv-issued DOI via DataCite

Submission history

From: Guoxuan Chen [view email]
[v1] Tue, 14 Jul 2026 17:52:05 UTC (47,729 KB)
[v2] Sat, 18 Jul 2026 17:28:17 UTC (47,613 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org