ai.hackcv
论文精选 60arXiv

Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations· 使用3D基础模型进行零样本新视角深度合成

3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.

AI 解读论文

利用3D基础模型的内部表示进行新视角的零样本深度预测。

核心方法
论文提出Z3D方法,通过在3D基础模型的内部表示上执行潜在扩散,从而估计新视角下的点云图,实现零样本深度合成。
适合谁读
研究者、工程师
要解决的问题
当前3D视觉模型难以高效地从新视角中推断出场景的3D结构,特别是在没有训练样本的情况下。
关键实验
在多个数据集上验证了Z3D方法预测新视角深度图的有效性与准确性。
主要贡献
首次证明3D基础模型的内部表示能够解码隐藏表面,并且提出了Z3D,能够在多个数据集上预测新视角下的真实深度图。
意义与局限
为3D视觉任务提供了一种新的零样本学习方法,但可能受限于3DFM内部表示的泛化能力和特定场景的特点。
领域:cs.CV作者:Denis M. Akola、David F. Fouhey
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考