提高使用非政策排级的逐步演变RL的抽样效率 (Improving Sample Efficiency in Evolutionary RL Using Off-Policy Ranking) - 专知论文

会员服务 ·

0

秩 · INTERACT · Learning · 粤港澳大湾区数字经济研究院 · 样本 ·

2022 年 8 月 22 日

Improving Sample Efficiency in Evolutionary RL Using Off-Policy Ranking

翻译：提高使用非政策排级的逐步演变RL的抽样效率

Eshwar S R,Shishir Kolathaya,Gugan Thoppe

Evolution Strategy (ES) is a powerful black-box optimization technique based on the idea of natural evolution. In each of its iterations, a key step entails ranking candidate solutions based on some fitness score. For an ES method in Reinforcement Learning (RL), this ranking step requires evaluating multiple policies. This is presently done via on-policy approaches: each policy's score is estimated by interacting several times with the environment using that policy. This leads to a lot of wasteful interactions since, once the ranking is done, only the data associated with the top-ranked policies is used for subsequent learning. To improve sample efficiency, we propose a novel off-policy alternative for ranking, based on a local approximation for the fitness function. We demonstrate our idea in the context of a state-of-the-art ES method called the Augmented Random Search (ARS). Simulations in MuJoCo tasks show that, compared to the original ARS, our off-policy variant has similar running times for reaching reward thresholds but needs only around 70% as much data. It also outperforms the recent Trust Region ES. We believe our ideas should be extendable to other ES methods as well.

翻译：进化策略( ES) 是一种基于自然进化理念的强大黑盒优化技术。在每一个迭代中, 关键步骤都包含基于某些健身分的排名候选解决方案。对于强化学习的ES 方法( RL), 这一排名步骤需要评估多种政策。目前, 是通过政策性方法进行的: 每项政策的得分都是通过使用该政策与环境进行多次互动来估算的。这导致大量浪费性互动, 因为排名完成后, 只有与最高等级政策相关的数据才能用于随后的学习。为了提高抽样效率, 我们提出一个新的非政策性排名替代方案, 其依据是健身功能的本地近似值。我们用最先进的ES 方法( ARS) 来展示我们的想法。 Mujoco 任务模拟显示, 与原始的ARS 相比, 我们的离政策变式在达到奖励阈值方面有着相似的运行时间, 但只需要大约70 %的数据。它也比最近的 Trust区域 ES 。我们相信, 我们的想法应该可以推广到其他ES 。

0

相关内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

地下水中二恶烷的纳米四氧化三铁/生物炭活化过硫酸盐高级氧化修复机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

柔性电子卷到卷制造中异质结构可控转移与层合机理

国家自然科学基金

0+阅读 · 2014年12月31日

MicroRNA调控BACE1在AD发病中的作用与机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

肝脏再生的细胞和分子调控机制

国家自然科学基金

0+阅读 · 2013年12月31日

脑内NG2细胞吞噬β淀粉样蛋白的能力及其代谢途径

国家自然科学基金

0+阅读 · 2012年12月31日

微纳结构Ag3PO4空心球/石墨烯异质结的构筑及光催化性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

去除水中微量As(Ⅲ)/ As(Ⅴ)的吸附剂制备及其构效关系研究

国家自然科学基金

0+阅读 · 2011年12月31日

等离子体助离子液体中可磁分离TiO2形成机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

信号转导通路和表观遗传模式在双酚A神经发育毒性中的作用

国家自然科学基金

0+阅读 · 2009年12月31日

以离子液体为溶剂的丙烯腈ARGET ATRP研究

国家自然科学基金

0+阅读 · 2009年12月31日

Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias

Arxiv

0+阅读 · 2022年10月6日

Learning convergence prediction of astrobots in multi-object spectrographs

Arxiv

0+阅读 · 2022年10月5日

A uniform kernel trick for high-dimensional two-sample problems

Arxiv

0+阅读 · 2022年10月5日

Learning Dynamic Abstract Representations for Sample-Efficient Reinforcement Learning

Arxiv

0+阅读 · 2022年10月4日

Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RL

Arxiv

0+阅读 · 2022年10月4日

Improving Robustness of Deep Reinforcement Learning Agents: Environment Attack based on the Critic Network

Arxiv

0+阅读 · 2022年10月3日

Learning GFlowNets from partial episodes for improved convergence and stability

Arxiv

0+阅读 · 2022年9月30日

Safe Exploration Method for Reinforcement Learning under Existence of Disturbance

Arxiv

0+阅读 · 2022年9月30日

S2P: State-conditioned Image Synthesis for Data Augmentation in Offline Reinforcement Learning

Arxiv

0+阅读 · 2022年9月30日

Feature Denoising for Improving Adversarial Robustness

Feature Denoising for Improving Adversarial Robustness

Arxiv

15+阅读 · 2018年12月9日

VIP会员

文章信息

相关主题

粤港澳大湾区数字经济研究院

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

大语言模型智能体强化学习：全景综述

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

【伯克利博士论文】从推理服务到训练：面向大规模 LLM 智能体的高效系统

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias

Arxiv

0+阅读 · 2022年10月6日

Learning convergence prediction of astrobots in multi-object spectrographs

Arxiv

0+阅读 · 2022年10月5日

A uniform kernel trick for high-dimensional two-sample problems

Arxiv

0+阅读 · 2022年10月5日

Learning Dynamic Abstract Representations for Sample-Efficient Reinforcement Learning

Arxiv

0+阅读 · 2022年10月4日

Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RL

Arxiv

0+阅读 · 2022年10月4日

Improving Robustness of Deep Reinforcement Learning Agents: Environment Attack based on the Critic Network

Arxiv

0+阅读 · 2022年10月3日

Learning GFlowNets from partial episodes for improved convergence and stability

Arxiv

0+阅读 · 2022年9月30日

Safe Exploration Method for Reinforcement Learning under Existence of Disturbance

Arxiv

0+阅读 · 2022年9月30日

S2P: State-conditioned Image Synthesis for Data Augmentation in Offline Reinforcement Learning

Arxiv

0+阅读 · 2022年9月30日

Feature Denoising for Improving Adversarial Robustness

Feature Denoising for Improving Adversarial Robustness

Arxiv

15+阅读 · 2018年12月9日

相关基金

地下水中二恶烷的纳米四氧化三铁/生物炭活化过硫酸盐高级氧化修复机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

柔性电子卷到卷制造中异质结构可控转移与层合机理

国家自然科学基金

0+阅读 · 2014年12月31日

MicroRNA调控BACE1在AD发病中的作用与机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

肝脏再生的细胞和分子调控机制

国家自然科学基金

0+阅读 · 2013年12月31日

脑内NG2细胞吞噬β淀粉样蛋白的能力及其代谢途径

国家自然科学基金

0+阅读 · 2012年12月31日

微纳结构Ag3PO4空心球/石墨烯异质结的构筑及光催化性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

去除水中微量As(Ⅲ)/ As(Ⅴ)的吸附剂制备及其构效关系研究

国家自然科学基金

0+阅读 · 2011年12月31日

等离子体助离子液体中可磁分离TiO2形成机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

信号转导通路和表观遗传模式在双酚A神经发育毒性中的作用

国家自然科学基金

0+阅读 · 2009年12月31日

以离子液体为溶剂的丙烯腈ARGET ATRP研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员