跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分35

EvolvingAvatar:随对话演进自适应调整的交互式 3D 头部生成

AI 导读

EvolvingAvatar 是一种因果生成器,通过测试时训练在交互过程中适配用户面部视频与双人音频,其双人上下文预测目标无需测试时目标动作标注即可提供自监督信号。

正文

View PDF HTML (experimental)

Abstract:Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
Comments: Project Page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.35616 [cs.CV]
  (or arXiv:2609.35616v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.35616

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Junjie Chen [view email]
[v1] Mon, 28 Sep 2026 16:56:07 UTC (32,314 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org