ai.hackcv
论文精选 60arXiv

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States· Puffin-World:用原生3D世界状态扩展统一多模态模型

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

AI 解读论文

Puffin-World 提出了一种集成了物理理解、空间模拟和3D世界生成的多模态模型。

核心方法
通过集成物理状态(重力场和经纬度)、几何状态(深度)和外观状态(图像)的原生3D世界状态模型,以及一个支持多样任务和灵活动作的统一全相机表示(Omni-Camera),Puffin-World 模型能够生成物理上一致且视觉上稳定的世界,并通过真实世界的相机属性实现这一目标。
适合谁读
研究者 / 工程师
要解决的问题
如何在不依赖外部离线模块的情况下,构建能够理解和交互3D世界的统一多模态模型。
关键实验
实验展示了Puffin-World在生成物理上一致且视觉上稳定的3D世界方面的性能,以及其在模仿和自校准世界探索等闭合回路应用中的能力。具体实验数据和设置详见论文。
主要贡献
实现了一个能够跨未来帧传播物理动态的策略,同步合成了未来视图及其底层几何结构,构造了Puffin-16M数据集,包含1500万视觉-语言-相机三元组和100万具有多样和挑战性动作的轨迹。
意义与局限
Puffin-World 的提出为多模态AI模型的发展提供了新的方向,特别是在3D环境理解与生成领域。它能够促进更复杂真实场景中的应用,但也存在计算量大和依赖高质量数据集的局限。
领域:cs.CV作者:Kang Liao、Yihang Luo、Xiao-Ming Wu
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考