基于连续时间马尔可夫决策过程的方法及其在追逃示例中的应用

A Continuous-Time Markov Decision Process-Based Method With Application in a Pursuit-Evasion Example

IEEE Transactions on Systems, Man, and Cybernetics: Systems · 2015
被引 17
ABS 3

中文导读

提出一种连续时间马尔可夫决策过程方法,用于处理追逃问题中的不确定性,通过考虑状态间转移时间的影响,在经典追逃基准问题上得到接近微分博弈解析解的离散化解,且比传统方法对转移概率变化更具鲁棒性。

Abstract

This paper presents a novel method-continuous-time Markov decision process (CTMDP)-to address the uncertainties in pursuit-evasion problem. The primary difference between the CTMDP and the Markov decision process (MDP) is that the former takes into account the influence of the transition time between the states. The policy iteration method-based potential performance for solving the CTMDP and its convergence are also presented. The results obtained by MDP-based method demonstrate that it is a special case of CTMDP-based method involving the identity transition rate matrix. To compare the methods, a well-known pursuit-evasion problem, involving two identical cars, is solved as a benchmark. The CTMDP-based method can provide a discretization solution that is close to the analytical solution obtained by the differential game method. Besides, it shows strong robustness against changes in the transition probability, as compared with the traditional MDP-based method. To the best of our knowledge, this is the first attempt to validate the influence of the transition time between the states in such a pursuit-evasion scenario, or in a similar application, solved by an MDP-related model. The CTMDP-based method offers a new approach to solving the pursuit-evasion problem and can be extended to similar optimization applications.

连续时间马尔可夫决策过程追逃问题策略迭代鲁棒性优化方法