Learning anticipation is a reasoning paradigm in multi-agent reinforcement learning, where agents, during learning, consider the anticipated learning of other agents. There has been substantial research into the role of learning anticipation in improving cooperation among self-interested agents in general-sum games. Two primary examples are Learning with Opponent-Learning Awareness (LOLA), which anticipates and shapes the opponent's learning process to ensure cooperation among self-interested agents in various games such as iterated prisoner's dilemma, and Look-Ahead (LA), which uses learning anticipation to guarantee convergence in games with cyclic behaviors. So far, the effectiveness of applying learning anticipation to fully-cooperative games has not been explored. In this study, we aim to research the influence of learning anticipation on coordination among common-interested agents. We first illustrate that both LOLA and LA, when applied to fully-cooperative games, degrade coordination among agents, causing worst-case outcomes. Subsequently, to overcome this miscoordination behavior, we propose Hierarchical Learning Anticipation (HLA), where agents anticipate the learning of other agents in a hierarchical fashion. Specifically, HLA assigns agents to several hierarchy levels to properly regulate their reasonings. Our theoretical and empirical findings confirm that HLA can significantly improve coordination among common-interested agents in fully-cooperative normal-form games. With HLA, to the best of our knowledge, we are the first to unlock the benefits of learning anticipation for fully-cooperative games.
翻译:在多试剂强化学习中,预期学习是多试剂强化学习的一种推理范式,代理商在学习过程中考虑其他代理商的预期学习。对于学习预期对于改善普通游戏中自利剂间合作的作用进行了大量研究。有两个主要的例子是 " 学习与亲利者学习意识 " (LOLA),这预示和塑造了对手学习过程,以确保自利者在各种游戏中的合作,如循环囚犯的困境,以及 " Look-Ahead(LA) " (LA),它利用学习预期来保证游戏与周期行为趋同。到目前为止,尚未探讨将学习预期应用到全面合作游戏中来提高自我利益者之间合作的预期作用。在本研究中,我们的目标是研究学习对共同利益者之间协调的预期所产生的影响。我们首先说明,如果应用LOLA和L(L)两个对手的学习过程,就会降低彼此之间的协调,从而造成最坏的结果。随后,我们提议高端的游戏学习预测(HLA), 代理商首先预测其他代理人学习等级的学习效果。具体地说,HLA 将他们的理论分析结果分配给我们的共同代理人。</s>