大提前期马尔可夫决策过程中半开环策略的渐近最优性

Asymptotic Optimality of Semi-Open-Loop Policies in Markov Decision Processes with Large Lead Times

Operations Research · 2023

被引 1

人大 AFT50UTD24ABS 4*

Menglong Li · 香港城市大学
Xingyu Bai · 伊利诺伊大学厄巴纳-香槟分校
Xin Chen · 佐治亚理工学院
Alexander Stolyar · 伊利诺伊大学厄巴纳-香槟分校

中文导读

研究了大提前期下马尔可夫决策过程的半开环策略，证明其渐近最优性，为库存管理等延迟控制问题提供理论依据。

Abstract

A generic way to verify asymptotic optimality of semi-open-loop policies for a wide class of MDPs with large lead times. In many real-life inventory models, order lead times can result in uncertain effects of inventory decisions. However, as the lead time grows large, one would naturally postulate that the effect of the delayed order depends weakly on the current inventory level and, thus, intuit that decoupling the delayed order with the current inventory level may provide good heuristics. Motivated by these examples, in “Asymptotic Optimality of Semi-open-Loop Policies in Markov Decision Processes with Large Lead Times,” Bai et al. consider a generic Markov decision process (MDP) with one delayed control and one immediate control. For MDPs defined on general spaces with uniformly bounded cost functions and a fast mixing property, they construct a periodic semi-open-loop policy for each lead time value and show that these policies are asymptotically optimal as the lead time goes to infinity. For MDPs defined on Euclidean spaces with linear dynamics and convex structures, they impose another set of conditions under which constant delayed-control policies are asymptotically optimal.

马尔可夫决策过程库存管理渐近最优性半开环策略提前期

阅读原文 ↗