策略镜像下降天生探索动作空间

Policy Mirror Descent Inherently Explores Action Space

SIAM Journal on Optimization · 2025
被引 0
ABS 3
强化学习优化理论机器学习运筹学