ai.hackcv
论文精选 65arXiv

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators· 前向过程 RL 在 MeanFlow 生成器的应用

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow samples with average velocities. To bridge this gap, we introduce MeanFlowNFT. Inspired by the MeanFlow identity, which bridges average and instantaneous velocities, we construct an induced instantaneous-velocity predictor. We apply the DiffusionNFT objective to this predictor, making reward optimization well-defined for MeanFlow. Sampling remains based on the average velocity, preserving MeanFlow's fast few-step generation. We further prove that MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee. Experiments on image and video generation show that MeanFlowNFT consistently improves baselines. Moreover, it outperforms prior state-of-the-art RL-tuned few-step generators on most metrics ($6$ of $8$ on SD3.5-M), and can even surpass multi-step RL-tuned diffusion while using only a few sampling steps. For instance, on Wan 2.1, $4$-step MeanFlowNFT reaches a VBench score of $84.33$, surpassing $50$-step LongCat-Video RL ($82.57$).

AI 解读论文

前向过程 RL 优化 MeanFlow 生成器,提升多步采样模型效率。

核心方法
通过构建诱导的瞬时速度预测器,并应用 DiffusionNFT 的优化目标,使 MeanFlow 生成器能够基于平均速度进行快速多步采样。
适合谁读
研究者 / 工程师
要解决的问题
如何将前向过程的强化学习方法应用到 MeanFlow 生成器中,以优化其性能和生成质量。
关键实验
在图像和视频生成任务上进行了实验,展示了 MeanFlowNFT 的性能提升。例如,在 Wan 2.1 数据集上,4 步 MeanFlowNFT 达到了 84.33 的 VBench 分数,超过了 50 步的 LongCat-Video RL (82.57)。
主要贡献
1. 引入 MeanFlowNFT,使得 MeanFlow 生成器可以利用前向过程的 RL 进行优化。 2. 证明了 MeanFlowNFT 继承了 DiffusionNFT 的严格策略改进保证。 3. 实验表明 MeanFlowNFT 在多数指标上优于已有的前向 RL 调优的多步生成器。
意义与局限
该研究弥合了扩散模型和流模型在前向过程 RL 上的差距,提高了生成模型的效率和质量。局限在于需要验证更多任务和数据集上的表现。
领域:cs.CV作者:Yushi Huang、Xiangxin Zhou、Jun Zhang
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考