带约束的平均连续时间马尔可夫决策过程中混合策略的最优性

Optimality of Mixed Policies for Average Continuous-Time Markov Decision Processes with Constraints

Mathematics of Operations Research · 2016
被引 2
ABS 3

中文导读

研究了带N个约束的平均连续时间马尔可夫决策过程,证明每个极端点由确定性平稳策略生成,且存在混合最优策略,混合策略数不超过N+1个。

Abstract

This article concerns the average criteria for continuous-time Markov decision processes with N constraints. We show the following; (a) every extreme point of the space of performance vectors corresponding to the set of stable measures is generated by a deterministic stationary policy; and (b) there exists a mixed optimal policy, where the mixture is over no more than N + 1 deterministic stationary policies.

马尔可夫决策过程数学优化随机过程运筹学