面向风险厌恶随机优化的并行非平稳直接策略搜索

Parallel Nonstationary Direct Policy Search for Risk-Averse Stochastic Optimization

INFORMS journal on computing · 2017
被引 8
UTD 24ABS 3

中文导读

提出一种针对有限时域、离散时间马尔可夫决策问题的非平稳策略搜索算法,通过非参数响应面模型和并行无导数优化处理大状态空间和风险敏感准则,并在最优储能充电问题中验证了有效性。

Abstract

This paper presents an algorithmic strategy to nonstationary policy search for finite-horizon, discrete-time Markovian decision problems with large state spaces, constrained action sets, and a risk-sensitive optimality criterion. The methodology relies on modeling time-variant policy parameters by a nonparametric response surface model for an indirect parametrized policy motivated by Bellman’s equation. The policy structure is heuristic when the optimization of the risk-sensitive criterion does not admit a dynamic programming reformulation. Through the interpolating approximation, the level of nonstationarity of the policy, and consequently, the size of the resulting search problem can be adjusted. The computational tractability and the generality of the approach follow from a nested parallel implementation of derivative-free optimization in conjunction with Monte Carlo simulation. We demonstrate the efficiency of the approach on an optimal energy storage charging problem, and illustrate the effect of the risk functional on the improvement achieved by allowing a higher complexity in time variation for the policy. The online supplement is available at https://doi.org/10.1287/ijoc.2016.0733 .

随机优化风险敏感决策动态规划蒙特卡洛模拟能源存储