Ask-AC: 一种在演员评论家框架中的征询式顾问 (Ask-AC: An Initiative Advisor-in-the-Loop Actor-Critic Framework) - 专知论文

会员服务 ·

0

Continuity · Learning · 回合 · Agent · 可交换的 ·

2023 年 3 月 21 日

Ask-AC: An Initiative Advisor-in-the-Loop Actor-Critic Framework

翻译：Ask-AC: 一种在演员评论家框架中的征询式顾问

Shunyu Liu,Na Yu,Jie Song,Kaixuan Chen,Zunlei Feng,Mingli Song

Despite the promising results achieved, state-of-the-art interactive reinforcement learning schemes rely on passively receiving supervision signals from advisor experts, in the form of either continuous monitoring or pre-defined rules, which inevitably result in a cumbersome and expensive learning process. In this paper, we introduce a novel initiative advisor-in-the-loop actor-critic framework, termed as Ask-AC, that replaces the unilateral advisor-guidance mechanism with a bidirectional learner-initiative one, and thereby enables a customized and efficacious message exchange between learner and advisor. At the heart of Ask-AC are two complementary components, namely action requester and adaptive state selector, that can be readily incorporated into various discrete actor-critic architectures. The former component allows the agent to initiatively seek advisor intervention in the presence of uncertain states, while the latter identifies the unstable states potentially missed by the former especially when environment changes, and then learns to promote the ask action on such states. Experimental results on both stationary and non-stationary environments and across different actor-critic backbones demonstrate that the proposed framework significantly improves the learning efficiency of the agent, and achieves the performances on par with those obtained by continuous advisor monitoring.

翻译：尽管取得了令人满意的结果，但是目前最先进的交互式强化学习方案仍然依赖于从顾问专家处被动地接收监管信号，其形式为连续监测或预定义规则，这无疑会导致一种繁琐而昂贵的学习过程。本文介绍了一种新的主动征询式顾问框架，名为Ask-AC，它将单向的顾问指导机制替换为双向的学习者主动机制，从而能够在学习者和顾问之间实现定制化和有效的消息交换。Ask-AC的核心是两个互补的组件，即动作请求器和自适应状态选择器，可以轻松地嵌入各种离散演员评论家体系结构中。前者允许代理在存在不确定状态时主动寻求顾问干预，而后者则在环境发生变化时识别可能被前者忽略的不稳定状态，然后学习促进在这些状态上的询问行为。对于静态和非静态环境以及不同演员评论家骨架，实验结果表明，所提出的框架显著提高了代理的学习效率，并达到了与连续顾问监测获得的性能相当的水平。

0

相关内容

Continuity

让 iOS 8 和 OS X Yosemite 无缝切换的一个新特性。 > Apple products have always been designed to work together beautifully. But now they may really surprise you. With iOS 8 and OS X Yosemite, you’ll be able to do more wonderful things than ever before.

Source: Apple - iOS 8

斯坦福最新《强化学习》2023课程，Emma Brunskill主讲，附PPT下载

斯坦福最新《强化学习》2023课程，Emma Brunskill主讲，附PPT下载

专知会员服务

45+阅读 · 2023年1月17日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【硬核书】规划算法 (Planning Algorithm)，1023页pdf，Steven M. Illinois大学

【硬核书】规划算法 (Planning Algorithm)，1023页pdf，Steven M. Illinois大学

专知会员服务

167+阅读 · 2022年4月10日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【CVPR2020】自监督的深度视觉测程与在线适应，Self-Supervised Deep Visual Odometry

【CVPR2020】自监督的深度视觉测程与在线适应，Self-Supervised Deep Visual Odometry

专知会员服务

32+阅读 · 2020年5月14日

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

泡泡机器人SLAM

23+阅读 · 2019年1月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

【泡泡一分钟】神经SLAM：使用外部存储器让智能体学习探索环境

【泡泡一分钟】神经SLAM：使用外部存储器让智能体学习探索环境

泡泡机器人SLAM

12+阅读 · 2018年4月17日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

发射可调铂(II)配合物的设计和新型静电喷雾沉积电致发光器件的制备

国家自然科学基金

0+阅读 · 2015年12月31日

基于Click反应的离子型多孔有机框架同步富集和催化转化CO2研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于Irf8的突变体建立斑马鱼白血病模型

国家自然科学基金

0+阅读 · 2013年12月31日

基于生物催化/VEGFR/PDGFR双靶标介导的中药脂质体给药系统的构建及评价

国家自然科学基金

0+阅读 · 2013年12月31日

miR34c重启衰老清除急性髓系白血病干细胞与机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

复合阻挡涂层对SiC/Ti基复合材料界面反应的调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

Levy扩散过程与非局部偏微分方程

国家自然科学基金

1+阅读 · 2012年12月31日

基于动态混合故障模型和进化博弈论的可生存性分析方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

一种新的调控途径（Pax6-Sox2-EGFR）影响小鼠胚胎发育晚期神经干细胞的增值与迁移

国家自然科学基金

0+阅读 · 2008年12月31日

On the convergence of the MLE as an estimator of the learning rate in the Exp3 algorithm

Arxiv

0+阅读 · 2023年5月11日

Composition, Attention, or Both?

Arxiv

0+阅读 · 2023年5月11日

An Option-Dependent Analysis of Regret Minimization Algorithms in Finite-Horizon Semi-Markov Decision Processes

Arxiv

0+阅读 · 2023年5月10日

Best-Effort Adaptation

Arxiv

0+阅读 · 2023年5月10日

Towards Domain Generalization for ECG and EEG Classification: Algorithms and Benchmarks

Arxiv

0+阅读 · 2023年5月9日

E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation

Arxiv

0+阅读 · 2023年5月9日

A Comprehensive Survey on Source-free Domain Adaptation

Arxiv

10+阅读 · 2023年2月23日

MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration

Arxiv

12+阅读 · 2021年2月7日

Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

Arxiv

17+阅读 · 2020年9月8日

A Multi-Objective Deep Reinforcement Learning Framework

A Multi-Objective Deep Reinforcement Learning Framework

Arxiv

16+阅读 · 2018年6月27日

VIP会员

文章信息

相关主题

相关VIP内容

斯坦福最新《强化学习》2023课程，Emma Brunskill主讲，附PPT下载

斯坦福最新《强化学习》2023课程，Emma Brunskill主讲，附PPT下载

专知会员服务

45+阅读 · 2023年1月17日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【硬核书】规划算法 (Planning Algorithm)，1023页pdf，Steven M. Illinois大学

【硬核书】规划算法 (Planning Algorithm)，1023页pdf，Steven M. Illinois大学

专知会员服务

167+阅读 · 2022年4月10日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【CVPR2020】自监督的深度视觉测程与在线适应，Self-Supervised Deep Visual Odometry

【CVPR2020】自监督的深度视觉测程与在线适应，Self-Supervised Deep Visual Odometry

专知会员服务

32+阅读 · 2020年5月14日

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【博士论文】在低维与高维空间中对潜在表征的分析、建模与变换

《美军使用大语言模型技术生成领域特定文档》2025最新379页

【NeurIPS 2025】以语言为中心的全模态表征学习的可扩展性研究

智能体化多模态大语言模型综述

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

泡泡机器人SLAM

23+阅读 · 2019年1月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

【论文推荐】最新六篇强化学习相关论文—Sublinear、机器阅读理解、加速强化学习、对抗性奖励学习、人机交互

专知

17+阅读 · 2018年4月28日

【泡泡一分钟】神经SLAM：使用外部存储器让智能体学习探索环境

【泡泡一分钟】神经SLAM：使用外部存储器让智能体学习探索环境

泡泡机器人SLAM

12+阅读 · 2018年4月17日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

On the convergence of the MLE as an estimator of the learning rate in the Exp3 algorithm

Arxiv

0+阅读 · 2023年5月11日

Composition, Attention, or Both?

Arxiv

0+阅读 · 2023年5月11日

An Option-Dependent Analysis of Regret Minimization Algorithms in Finite-Horizon Semi-Markov Decision Processes

Arxiv

0+阅读 · 2023年5月10日

Best-Effort Adaptation

Arxiv

0+阅读 · 2023年5月10日

Towards Domain Generalization for ECG and EEG Classification: Algorithms and Benchmarks

Arxiv

0+阅读 · 2023年5月9日

E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation

Arxiv

0+阅读 · 2023年5月9日

A Comprehensive Survey on Source-free Domain Adaptation

Arxiv

10+阅读 · 2023年2月23日

MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration

Arxiv

12+阅读 · 2021年2月7日

Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

Arxiv

17+阅读 · 2020年9月8日

A Multi-Objective Deep Reinforcement Learning Framework

A Multi-Objective Deep Reinforcement Learning Framework

Arxiv

16+阅读 · 2018年6月27日

相关基金

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

发射可调铂(II)配合物的设计和新型静电喷雾沉积电致发光器件的制备

国家自然科学基金

0+阅读 · 2015年12月31日

基于Click反应的离子型多孔有机框架同步富集和催化转化CO2研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于Irf8的突变体建立斑马鱼白血病模型

国家自然科学基金

0+阅读 · 2013年12月31日

基于生物催化/VEGFR/PDGFR双靶标介导的中药脂质体给药系统的构建及评价

国家自然科学基金

0+阅读 · 2013年12月31日

miR34c重启衰老清除急性髓系白血病干细胞与机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

复合阻挡涂层对SiC/Ti基复合材料界面反应的调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

Levy扩散过程与非局部偏微分方程

国家自然科学基金

1+阅读 · 2012年12月31日

基于动态混合故障模型和进化博弈论的可生存性分析方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

一种新的调控途径（Pax6-Sox2-EGFR）影响小鼠胚胎发育晚期神经干细胞的增值与迁移

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员