Adaptive Policy Portfolios for Robust Markov Decision Processes
Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memoryless randomized policies synthesized offline and paired with a lightweight online selector. Robust regret is a natural measure of portfolio quality: for each plausible environment, it measures the loss of the best portfolio member relative to the policy that would have been optimal had that environment been known. Related regret objectives were studied by Ghavamzadeh et al. (2016) with an emphasis on approximations and relaxations for safe policy improvement. We give a complexity-theoretic account of portfolio certificati
研究适配型策略组合在部分可识别环境中的健壮性优化问题。
- 核心方法
- 论文提出了适应性策略组合的概念,即在离线阶段合成有限的记忆无策略集,并在线上阶段通过轻量级的选择器动态选择最合适的策略。
- 适合谁读
- 适合研究者和工程师阅读,特别是对强化学习和决策理论感兴趣的读者。
- 要解决的问题
- 这篇论文旨在解决在部分可识别的未知环境中,单个策略过于保守的问题,以及如何有效地利用有限的在线信息进行策略优化。
- 关键实验
- 未提供
- 主要贡献
- 主要贡献在于提供了一个理论框架来评估策略组合的质量,并探讨了在部分可识别环境下的健壮性策略优化。
- 意义与局限
- 该研究为动态环境下的决策提供了新的理论基础,但可能因假设的限制而在实际应用中存在局限。