AACHER: 以事后见经验重现为背景的各类行为者-批评者深强化学习 (AACHER: Assorted Actor-Critic Deep Reinforcement Learning with Hindsight Experience Replay) - 专知论文

会员服务 ·

0

评论员 · Learning · 经验回放 · Performer · 强化学习 ·

2022 年 10 月 24 日

AACHER: Assorted Actor-Critic Deep Reinforcement Learning with Hindsight Experience Replay

翻译：AACHER: 以事后见经验重现为背景的各类行为者-批评者深强化学习

Adarsh Sehgal,Muskan Sehgal,Hung Manh La

Actor learning and critic learning are two components of the outstanding and mostly used Deep Deterministic Policy Gradient (DDPG) reinforcement learning method. Since actor and critic learning plays a significant role in the overall robot's learning, the performance of the DDPG approach is relatively sensitive and unstable as a result. We propose a multi-actor-critic DDPG for reliable actor-critic learning to further enhance the performance and stability of DDPG. This multi-actor-critic DDPG is then integrated with Hindsight Experience Replay (HER) to form our new deep learning framework called AACHER. AACHER uses the average value of multiple actors or critics to substitute the single actor or critic in DDPG to increase resistance in the case when one actor or critic performs poorly. Numerous independent actors and critics can also gain knowledge from the environment more broadly. We implemented our proposed AACHER on goal-based environments: AuboReach, FetchReach-v1, FetchPush-v1, FetchSlide-v1, and FetchPickAndPlace-v1. For our experiments, we used various instances of actor/critic combinations, among which A10C10 and A20C20 were the best-performing combinations. Overall results show that AACHER outperforms the traditional algorithm (DDPG+HER) in all of the actor/critic number combinations that are used for evaluation. When used on FetchPickAndPlace-v1, the performance boost for A20C20 is as high as roughly 3.8 times the success rate in DDPG+HER.

翻译：动作学习和批评者学习是杰出且大多使用的深确定性政策强化学习方法的两个组成部分。由于演员和批评者学习在整个机器人学习中起着重要作用, DDPG 方法的性能相对敏感且不稳定。我们提议多动作- 批评性 DDPG 方法, 用于可靠的演员- 批评性学习, 以进一步提高 DDPG 的性能和稳定性。这个多动作- 批评性 DDPG 与 Hindsight 经验再游戏(HER) 整合, 以形成我们称为 AACHER 的新的深层次学习框架。 AACHER 使用多个演员或批评者的平均性能来取代DDPG 的单一演员或批评者。许多独立演员和批评者也可以更广泛地从环境中获取知识。我们在基于目标的环境中实施了我们提议的 AACCHER : AuboReach, Fetchrereach-v1, FreetPush-V1, FreetSled-v1, 和 Flick-Place-Place-v1。在我们的实验中,我们使用了各种AHR- CLE+C 10的AHR- dal- dal- dal- disal- dust- dust- 的组合中,我们使用了各种动作/C 和Axx- dust的AV1, 在ARC 的AVA- dis- dust 的A- disal- dow 的A- 和A- disal- disal- disal- dow 的演算中,我们用来用来用来展示式的A- dow 和A- dow 和A- dow 的演技法化的A- dow。

0

相关内容

评论员

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

51+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

高通量高灵敏度等离激元共振增强OI-RD光学生物传感方法及应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

聚精氨酸诱导肿瘤微环境的免疫活性及逆转cetuximab耐药性的调控机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

新型小分子化合物Me6促进辐射损伤肠组织再生修复的研究

国家自然科学基金

0+阅读 · 2014年12月31日

Hybrid加速结构的理论及预制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Nrf2-Keap1-ARE信号通路在脊髓损伤后胶质瘢痕形成中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

旋转弯曲微动疲劳裂纹萌生及扩展机理和寿命预测方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

Osteopontin-CCL5-Wnt参与肺癌相关间充质干细胞促进肺癌干细胞转移的研究

国家自然科学基金

0+阅读 · 2013年12月31日

microRNA调节肿瘤抑制因子Caliban应答DNA损伤的机制

国家自然科学基金

1+阅读 · 2012年12月31日

基于iRGD靶向载药脂质体-微泡复合体的超声成像引导给药治疗肿瘤的研究

国家自然科学基金

0+阅读 · 2012年12月31日

Analysis of Anomalous Behavior in Network Systems Using Deep Reinforcement Learning with CNN Architecture

Arxiv

0+阅读 · 2022年12月8日

Reinforcement Learning with Non-Exponential Discounting

Arxiv

0+阅读 · 2022年12月7日

What is the Solution for State-Adversarial Multi-Agent Reinforcement Learning?

Arxiv

0+阅读 · 2022年12月7日

Efficient Optimization with Higher-Order Ising Machines

Arxiv

0+阅读 · 2022年12月7日

Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2022年12月7日

Reinforcement Learning for UAV control with Policy and Reward Shaping

Arxiv

0+阅读 · 2022年12月6日

What is the Solution for State Adversarial Multi-Agent Reinforcement Learning?

Arxiv

0+阅读 · 2022年12月6日

Pretraining in Deep Reinforcement Learning: A Survey

Arxiv

21+阅读 · 2022年11月8日

Reinforcement Learning based Air Combat Maneuver Generation

Reinforcement Learning based Air Combat Maneuver Generation

Arxiv

91+阅读 · 2022年1月14日

Coding for Distributed Multi-Agent Reinforcement Learning

Arxiv

32+阅读 · 2021年1月7日

VIP会员

文章信息

相关主题

相关VIP内容

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

51+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

大语言模型智能体强化学习：全景综述

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

【伯克利博士论文】从推理服务到训练：面向大规模 LLM 智能体的高效系统

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

相关资讯

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Analysis of Anomalous Behavior in Network Systems Using Deep Reinforcement Learning with CNN Architecture

Arxiv

0+阅读 · 2022年12月8日

Reinforcement Learning with Non-Exponential Discounting

Arxiv

0+阅读 · 2022年12月7日

What is the Solution for State-Adversarial Multi-Agent Reinforcement Learning?

Arxiv

0+阅读 · 2022年12月7日

Efficient Optimization with Higher-Order Ising Machines

Arxiv

0+阅读 · 2022年12月7日

Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2022年12月7日

Reinforcement Learning for UAV control with Policy and Reward Shaping

Arxiv

0+阅读 · 2022年12月6日

What is the Solution for State Adversarial Multi-Agent Reinforcement Learning?

Arxiv

0+阅读 · 2022年12月6日

Pretraining in Deep Reinforcement Learning: A Survey

Arxiv

21+阅读 · 2022年11月8日

Reinforcement Learning based Air Combat Maneuver Generation

Reinforcement Learning based Air Combat Maneuver Generation

Arxiv

91+阅读 · 2022年1月14日

Coding for Distributed Multi-Agent Reinforcement Learning

Arxiv

32+阅读 · 2021年1月7日

相关基金

高通量高灵敏度等离激元共振增强OI-RD光学生物传感方法及应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

聚精氨酸诱导肿瘤微环境的免疫活性及逆转cetuximab耐药性的调控机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

新型小分子化合物Me6促进辐射损伤肠组织再生修复的研究

国家自然科学基金

0+阅读 · 2014年12月31日

Hybrid加速结构的理论及预制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Nrf2-Keap1-ARE信号通路在脊髓损伤后胶质瘢痕形成中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

旋转弯曲微动疲劳裂纹萌生及扩展机理和寿命预测方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

Osteopontin-CCL5-Wnt参与肺癌相关间充质干细胞促进肺癌干细胞转移的研究

国家自然科学基金

0+阅读 · 2013年12月31日

microRNA调节肿瘤抑制因子Caliban应答DNA损伤的机制

国家自然科学基金

1+阅读 · 2012年12月31日

基于iRGD靶向载药脂质体-微泡复合体的超声成像引导给药治疗肿瘤的研究

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员