ai.hackcv
论文精选 82arXiv

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination· RxBrain:结合语言-视觉推理与想象的具身认知基础模型

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task de

领域:cs.AI
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考