ai.hackcv
论文精选 60arXiv

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.

AI 解读论文

提出了一种分层框架,可在分钟级别上生成与音乐同步的高质量舞蹈视频。

核心方法
将舞蹈生成过程分为全局关键帧规划和局部时间精炼两部分,使用时间映射RoPE嵌入实现动态帧率适应,通过基于光流的损失函数提高动作连续性,并通过运动速度控制保持快速动作中的高保真细节。
适合谁读
研究者、工程师、产品团队
要解决的问题
现有的方法无法生成长时间(超过20秒)、高质量、与音乐同步的舞蹈视频,存在时间漂移、身份不一致和重复运动模式的问题。
关键实验
通过广泛的实验,展示了该框架在生成长时间视频中的优越性能,并且在不同舞蹈类型中表现出强大的适用性和鲁棒性。
主要贡献
提出了能够生成超过一分钟的720p/30fps高质量舞蹈视频的新框架,解决了长时间生成视频中的多个技术难题。
意义与局限
此研究突破了当前视频生成模型的时间限制,为娱乐和内容创作领域提供了新的工具,但可能在更复杂场景中存在一定局限性。
领域:cs.CV作者:Mingyang Huang、Peng Zhang、Li Hu
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考