On the Behavior of the Optimal Value Operator of Dynamic Programming
研究了状态空间和行动空间为有限维欧氏空间时,离散时间动态规划的最优性方程,给出了保证最优值算子行为良好的充分条件,这些条件比现有文献更弱,并应用于N阶段库存控制模型。
The optimality equation of discrete time dynamic programming is considered when state space and action space are finite dimensional Euclidean spaces. Based on a measurable selection theorem we give an elementary derivation of sufficient conditions to assure that the optimal value operator behaves well. For our model these conditions are weaker than those described in the existing literature. Under related conditions it is easy to prove that an optimal Markovian strategy exists for a finite stage Markovian stochastic optimization problem, and that the optimal strategies are completely characterized by the minimum sets of the optimality equations. This is illustrated with a general N-stage inventory control model.