HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分47
MaLiang-Harness:可编程图像与视频生成框架,GPT-6-Astra 双榜生成成功率 100%
AI 导读
MaLiang-Harness 是一个将 MLLM 驱动的视觉生成组织为持续构建、检查与修订流程的统一框架,用于弥合程序可执行但违反构图、外观或运动要求的 Program-to-Visual(P2V)差距。
正文
Abstract:Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at this https URL.
| Comments: | 22 pages, 11 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.34309 [cs.CV] |
| (or arXiv:2609.34309v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34309 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haoyu Zhao [view email]
[v1]
Mon, 28 Sep 2026 04:48:06 UTC (19,901 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org