批量Q*学习中的数据驱动知识迁移

Data-Driven Knowledge Transfer in Batch Q * Learning

Journal of the American Statistical Association · 2026
被引 0 · 同刊同年前 8%
ABS 4

中文导读

研究了在批量静态环境中利用源数据迁移知识到目标任务的方法,提出了Transfer Fitted Q-Iteration算法,并证明了其相比单任务学习能显著降低学习误差,适用于数据稀缺场景。

Abstract

In data-driven decision-making across marketing, healthcare, and education, leveraging large datasets from existing ventures is crucial for navigating high-dimensional feature spaces and addressing data scarcity in new ventures. We investigate knowledge transfer in dynamic decision-making by focusing on batch stationary environments and formally defining task discrepancies through the framework of Markov decision processes (MDPs). We propose the Transfer Fitted Q-Iteration algorithm with general function approximation, which enables direct estimation of the optimal action-state function Q* using both target and source data. Under sieve approximation, we establish the relationship between statistical performance and the MDP task discrepancy, highlighting the influence of source and target sample sizes and task discrepancy on the effectiveness of knowledge transfer. Our theoretical and empirical results demonstrate that the final learning error of the function is significantly reduced compared to the single-task learning rate.

市场营销医疗健康教育动态决策机器学习