抽象动态规划中的正则策略

Regular Policies in Abstract Dynamic Programming

SIAM Journal on Optimization · 2017
被引 12
ABS 3

中文导读

研究了抽象动态规划中贝尔曼方程及值迭代、策略迭代的复杂行为,提出正则策略概念,统一解决多类无折扣模型中的分析与算法难题。

Abstract

We consider challenging dynamic programming models where the associated Bellman equation, and the value and policy iteration algorithms commonly exhibit complex and even pathological behavior. Our analysis is based on the new notion of regular policies. These are policies that are well-behaved with respect to value and policy iteration, and are patterned after proper policies, which are central in the theory of stochastic shortest path problems. We show that the optimal cost function over regular policies may have favorable value and policy iteration properties, which the optimal cost function over all policies need not have. We accordingly develop a unifying methodology to address long standing analytical and algorithmic issues in broad classes of undiscounted models, including stochastic and minimax shortest path problems, as well as positive cost, negative cost, risk-sensitive, and multiplicative cost problems.

动态规划贝尔曼方程随机最短路径问题数学优化