离策略强化学习:双时间尺度工业过程的最优运行控制

Off-Policy Reinforcement Learning: Optimal Operational Control for Two-Time-Scale Industrial Processes

IEEE Transactions on Cybernetics · 2017
被引 71
ABS 3

中文导读

针对快慢两种时间尺度的工业流程,提出一种基于离策略强化学习的无模型最优控制方法,通过实时数据求解最优设定点,并在浮选过程仿真中验证有效性。

Abstract

Industrial flow lines are composed of unit processes operating on a fast time scale and performance measurements known as operational indices measured at a slower time scale. This paper presents a model-free optimal solution to a class of two time-scale industrial processes using off-policy reinforcement learning (RL). First, the lower-layer unit process control loop with a fast sampling period and the upper-layer operational index dynamics at a slow time scale are modeled. Second, a general optimal operational control problem is formulated to optimally prescribe the set-points for the unit industrial process. Then, a zero-sum game off-policy RL algorithm is developed to find the optimal set-points by using data measured in real-time. Finally, a simulation experiment is employed for an industrial flotation process to show the effectiveness of the proposed method.

强化学习工业过程控制最优控制双时间尺度系统