通过平稳策略启发式方法求解整体风险最小化问题

Solving overall risk minimization by stationary policy heuristics

Annals of Operations Research · 2026
被引 0 · 同刊同年前 10%
ABS 3

中文导读

针对无限时域动态规划中的整体风险最小化问题,提出一种启发式方法,通过限制为马尔可夫策略并局部近似风险度量,使问题可用标准算法求解,数值实验证明其高效性。

Abstract

Abstract We propose a heuristic method for solving the overall risk minimization problem, specifically the infinite-horizon dynamic programming problem, which uses a coherent risk measure of the overall discounted reward as its criterion. The original problem is very complex, and its optimal policies may depend on the entire history of previous decisions. The heuristic consists in restricting attention to Markov policies—those depending only on the current state—and in consecutive local approximations of the criterion, a static coherent risk measure, by its nested dynamic counterpart: a limit nested risk measure. The approximating problems with the nested measures can be reformulated using the Bellman equation and, as such, can be solved by standard algorithms. We introduce several variants of this heuristic. Consequently, we conduct three numerical experiments, all of which demonstrate the efficiency of our heuristic in finding effective solutions. Due to its speed, the heuristic is suitable for finding “good enough” solutions, which may be further refined.

动态规划风险度量启发式算法马尔可夫过程