HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分34
Adaptive Reward Routing:面向联合音视频扩散模型的前向过程 RL 动态多奖励优化
AI 导读
针对联合音视频扩散模型的多奖励强化学习,研究者提出 Adaptive Reward Routing,在 DiffusionNFT 前向过程 RL 中同时自适应调整更新位置与奖励协调方式。
正文
Abstract:Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.37200 [cs.CV] |
| (or arXiv:2609.37200v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37200 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Songlin Yang [view email]
[v1]
Tue, 29 Sep 2026 10:18:01 UTC (2,002 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org