Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.
提出了一种分层框架,可在分钟级别上生成与音乐同步的高质量舞蹈视频。
- 核心方法
- 将舞蹈生成过程分为全局关键帧规划和局部时间精炼两部分,使用时间映射RoPE嵌入实现动态帧率适应,通过基于光流的损失函数提高动作连续性,并通过运动速度控制保持快速动作中的高保真细节。
- 适合谁读
- 研究者、工程师、产品团队
- 要解决的问题
- 现有的方法无法生成长时间(超过20秒)、高质量、与音乐同步的舞蹈视频,存在时间漂移、身份不一致和重复运动模式的问题。
- 关键实验
- 通过广泛的实验,展示了该框架在生成长时间视频中的优越性能,并且在不同舞蹈类型中表现出强大的适用性和鲁棒性。
- 主要贡献
- 提出了能够生成超过一分钟的720p/30fps高质量舞蹈视频的新框架,解决了长时间生成视频中的多个技术难题。
- 意义与局限
- 此研究突破了当前视频生成模型的时间限制,为娱乐和内容创作领域提供了新的工具,但可能在更复杂场景中存在一定局限性。