Easy Affine Markov Decision Processes
研究了一类可分解的仿射马尔可夫决策过程,其值函数和最优策略均为线性,可通过求解一个小型辅助MDP精确计算,避免了维度灾难和离散化误差,适用于渔业管理、产能组合管理和商品采购。
Individuals, firms, and governments often face the challenge of making optimal decisions in a dynamic setting amidst a changing and uncertain environment. Although Markov decision processes (MDPs) provide a powerful modeling framework for such problems, solving an MDP is generally difficult. In “Easy Affine Markov Decision Processes,” Ning and Sobel characterize decomposable affine MDPs, which can be solved easily. An MDP in this class has continuous multidimensional controlled states and actions, and Markov-modulated uncontrolled states. With affine immediate reward and transition dynamics and affine constraints on actions, these MDPs have a linear value function and a linear optimal policy. The linear coefficients can be computed easily and exactly by solving a small auxiliary MDP, which frees the solution of the original MDP from the curse of dimensionality and errors caused by discretizing the continuous state and action spaces. The applicability of decomposable affine MDPs is illustrated in fishery management, capacity portfolio management, and commodity procurement.