情景援助深强化学习 (Scenario-Assisted Deep Reinforcement Learning)

Deep reinforcement learning has proven remarkably useful in training agents from unstructured data. However, the opacity of the produced agents makes it difficult to ensure that they adhere to various requirements posed by human engineers. In this work-in-progress report, we propose a technique for enhancing the reinforcement learning training process (specifically, its reward calculation), in a way that allows human engineers to directly contribute their expert knowledge, making the agent under training more likely to comply with various relevant constraints. Moreover, our proposed approach allows formulating these constraints using advanced model engineering techniques, such as scenario-based modeling. This mix of black-box learning-based tools with classical modeling approaches could produce systems that are effective and efficient, but are also more transparent and maintainable. We evaluated our technique using a case-study from the domain of internet congestion control, obtaining promising results.

翻译：深层强化学习被证明对利用非结构化数据培训代理机构非常有益,然而,由于生产代理机构不透明,难以确保它们遵守人类工程师提出的各种要求。在这份进行中的报告中,我们提出了一种加强强化学习培训过程(特别是其奖励计算方法)的方法,使人类工程师能够直接贡献其专业知识,使正在接受培训的代理机构更有可能遵守各种相关限制。此外,我们提出的方法允许利用基于情景的模型模型模型模型模型模型等先进工程技术来制定这些制约因素。这种黑箱学习工具与经典模型模型方法相结合,可以产生有效和高效的系统,但也更透明、更便于维护。我们用互联网交通拥挤控制领域的案例研究评估了我们的技术,取得了有希望的结果。

相关内容

深度强化学习

关注 154

深度强化学习 (DRL) 是一种使用深度学习技术扩展传统强化学习方法的一种机器学习方法。传统强化学习方法的主要任务是使得主体根据从环境中获得的奖赏能够学习到最大化奖赏的行为。然而，传统无模型强化学习方法需要使用函数逼近技术使得主体能够学习出值函数或者策略。在这种情况下，深度学习强大的函数逼近能力自然成为了替代人工指定特征的最好手段并为性能更好的端到端学习的实现提供了可能。

计算机科学课程与视频课件合集，Computer Science courses with video lectures

专知会员服务

37+阅读 · 2022年1月24日

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

专知会员服务

74+阅读 · 2020年7月6日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日