ai.hackcv
论文精选 65arXiv

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence· 从被动视频到可编辑体验:具身智能的物理基础体验合成

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling ge

AI 解读论文

提出Pegasus框架,将人类操作视频转换为机器人可学习的数据。

核心方法
通过结构化知识转移,将人类视频中的任务图通过可操作性和约束图转化为机器人规划图,利用层次化的可操作性隐空间建模对象状态、可操作性与任务之间的关系。
适合谁读
研究者
要解决的问题
解决具身智能中人类形态与机器人硬件之间的体现差距问题,使得机器人可以利用大量在线人类操作视频进行学习。
关键实验
展示了Pegasus框架在不同的机器人任务上的应用,验证了其可行性和有效性;具体实验细节和数据未提供。
主要贡献
构建了一个低资源框架Pegasus,能够有效缩小人类和机器人之间的体现差距,提升机器人学习效率。
意义与局限
为具身智能提供了一种新的数据处理方法,有助于机器人更好地利用人类数据提升自身能力,但转化过程的准确性和泛化能力需要进一步研究。
领域:cs.AI作者:Jia Luo
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考