离散时间随机控制中通用可测策略的战略测度与最优性性质

On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies

Mathematics of Operations Research · 2023
被引 0
ABS 3

中文导读

研究了Borel状态和行动空间下通用可测策略的战略测度优化问题,并将其应用于风险中性和风险敏感的马尔可夫决策过程,证明了最优值函数的可测性及ε-最优策略的存在性。

Abstract

This paper concerns discrete-time infinite-horizon stochastic control systems with Borel state and action spaces and universally measurable policies. We study optimization problems on strategic measures induced by the policies in these systems. The results are then applied to risk-neutral and risk-sensitive Markov decision processes to establish the measurability of the optimal value functions and the existence of universally measurable, randomized or nonrandomized, ϵ-optimal policies, for a variety of average cost criteria and risk criteria. We also extend our analysis to a class of minimax control problems and establish similar optimality results under the axiom of analytic determinacy. Funding: This work was supported by grants from DeepMind, the Alberta Machine Intelligence Institute (AMII), and Alberta Innovates-Technology Futures (AITF).

随机控制马尔可夫决策过程最优控制数学经济学