Off-Policy Reinforcement Learning: Optimal Operational Control for Two-Time-Scale Industrial Processes
针对快慢两种时间尺度的工业流程,提出一种基于离策略强化学习的无模型最优控制方法,通过实时数据求解最优设定点,并在浮选过程仿真中验证有效性。
Industrial flow lines are composed of unit processes operating on a fast time scale and performance measurements known as operational indices measured at a slower time scale. This paper presents a model-free optimal solution to a class of two time-scale industrial processes using off-policy reinforcement learning (RL). First, the lower-layer unit process control loop with a fast sampling period and the upper-layer operational index dynamics at a slow time scale are modeled. Second, a general optimal operational control problem is formulated to optimally prescribe the set-points for the unit industrial process. Then, a zero-sum game off-policy RL algorithm is developed to find the optimal set-points by using data measured in real-time. Finally, a simulation experiment is employed for an industrial flotation process to show the effectiveness of the proposed method.