基于距离模糊集的分布鲁棒部分可观测马尔可夫决策过程

Distributionally Robust Partially Observable Markov Decision Process with Distance-Based Ambiguity Sets

IISE Transactions · 2025
被引 0
ABS 3

中文导读

研究了在随机奖励和转移观测概率分布未知时,使用距离模糊集建模不确定性,最大化最坏情况下的期望总折扣奖励,并证明了值函数的凸性,开发了可求解的算法。

Abstract

We consider a distributionally robust partially observable Markov decision process (DR-POMDP), where the joint distributions of random rewards and transition-observation probabilities are unknown to the decision maker. We model this distribution uncertainty using distance-based ambiguity sets, which contain all possible distributions within a statistical distance relative to a reference distribution. Our objective is to seek optimal policies that maximize the expected total discounted reward under the worst-case distribution. We first prove that under mild conditions, the value function of DR-POMDP is convex with respect to the belief state. Building upon this convexity, we develop tractable reformulations of the value function for several widely used distance metrics. We further investigate the asymptotic convergence of the optimal value and optimal policy for the proposed DR-POMDP. We show that the optimal value of the DR-POMDP provides a lower bound on the out-of-sample value of the distributionally robust optimal policy, with a certain probability. We adapt the heuristic search value iteration2 algorithm (HSVI2) to solve the proposed DR-POMDP with distance-based ambiguity sets. Computational studies are conducted to demonstrate the utility of the proposed DR-POMDP with distance-based ambiguity sets. Our results show that the DR-POMDP model yields better out-of-sample performance compared with the nominal model, especially for small-size training samples.

运筹学决策理论随机优化机器学习