FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy· FabriMAE:我信任自己?
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and state
研究视觉-语言-动作模型的自评估机制,通过熵评估任务成功与否。
- 核心方法
- 提出了一种基于Markov注意力熵的自评估方法,利用模型内部的视觉模态熵来区分任务成功与失败。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 现有的视觉-语言-动作模型(VLA)在没有外部监督的情况下难以自评估其动作生成的可靠性。
- 关键实验
- 通过在多个VLA模型上的实验验证了该方法的有效性,展示了内部熵与任务成功率之间的相关性。
- 主要贡献
- 首次展示了不同VLA架构在任务成功与失败时的内部视觉模态熵具有一致性,为自评估提供了新思路。
- 意义与局限
- 该研究有助于提高VLA模型的自主性和可靠性,但目前仅在特定任务上进行了验证。