Optimality of Mixed Policies for Average Continuous-Time Markov Decision Processes with Constraints
研究了带N个约束的平均连续时间马尔可夫决策过程,证明每个极端点由确定性平稳策略生成,且存在混合最优策略,混合策略数不超过N+1个。
This article concerns the average criteria for continuous-time Markov decision processes with N constraints. We show the following; (a) every extreme point of the space of performance vectors corresponding to the set of stable measures is generated by a deterministic stationary policy; and (b) there exists a mixed optimal policy, where the mixture is over no more than N + 1 deterministic stationary policies.