Joint Resource Allocation for Time-Varying Underwater Acoustic Communication System: A Self-Reflection Adversarial Bandit Approach
针对时变水声通信系统,提出一种结合对抗性多臂赌博机和过时导频反馈的分层学习算法,无需先验信道信息即可在线优化信道选择与功率分配,平衡探索与利用以应对快速时变环境。
This study deals with a joint channel selection and power allocation problem for time-varying underwater acoustic communication system. Without any prior channel information, designing a highly adaptable resource allocation algorithm to cope with the fast time-varying environment is a very challenging issue. To address this issue, a hierarchical learning approach, which is combined with adversarial multiarmed bandit theory and outdated pilot-based feedback information, is proposed. The proposed learning approach can online optimize joint resource allocate strategy without any prior channel state information. Specifically, a hierarchical self-reflection learning structure is proposed to offer different learning manners and spaces for the actual played information and outdated feedback information, thereby balancing the exploitation and exploration to cope with the time-varying environment effectively. Further, an integration learning structure is proposed to alleviate the solving difficulty and policy explosion of joint multiple substrategies problem. The user can rapidly achieve a few superior strategies in low-dimension space, then efficiently search the expected optimal strategy in high-dimension space, as a result, the learning efficiency is significantly improved. The proposed algorithms show strong tolerance for delay and noncomplete information due to the elaborate learning structures. The superiority of the proposed algorithms is demonstrated through numerical results.