回首惊讶时的回顾:稳定神经近似的反向经验回放 (Look Back When Surprised: Stabilizing Reverse Experience Replay for Neural Approximation) - 专知论文

会员服务 ·

0

经验回放 · 近似 · UniFormer · 有偏 · Learning ·

2022 年 9 月 30 日

Look Back When Surprised: Stabilizing Reverse Experience Replay for Neural Approximation

翻译：回首惊讶时的回顾:稳定神经近似的反向经验回放

Ramnath Kumar,Dheeraj Nagaraj

Experience replay-based sampling techniques are essential to several reinforcement learning (RL) algorithms since they aid in convergence by breaking spurious correlations. The most popular techniques, such as uniform experience replay (UER) and prioritized experience replay (PER), seem to suffer from sub-optimal convergence and significant bias error, respectively. To alleviate this, we introduce a new experience replay method for reinforcement learning, called Introspective Experience Replay (IER). IER picks batches corresponding to data points consecutively before the 'surprising' points. Our proposed approach is based on the theoretically rigorous reverse experience replay (RER), which can be shown to remove bias in the linear approximation setting but can be sub-optimal with neural approximation. We show empirically that IER is stable with neural function approximation and has a superior performance compared to the state-of-the-art techniques like uniform experience replay (UER), prioritized experience replay (PER), and hindsight experience replay (HER) on the majority of tasks.

翻译：基于经验重现的抽样技术对若干强化学习算法至关重要,因为它们通过打破假的关联而有助于趋同。最受欢迎的技术,例如统一经验重现(UER)和优先经验重现(PER),似乎分别受到亚最佳趋同和重大偏差错误的影响。为此,我们引入了一种新的强化学习经验重现方法,称为内向体验重现(IER)。IER选择了与“突变”点之前连续数据点相对应的分批。我们提议的方法基于理论上严格的反向重现(RER),这可以显示消除线性近距离设置的偏差,但可以是神经近距离的次优劣。我们从经验上表明,IER与神经功能近似稳定,并且与统一的经验重现(UER)、优先经验重现(PER)和在大多数任务上的短视重现(HER)相比,其性能更高。

0

相关内容

经验回放

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

MARVELD1基因调控肝细胞癌介入治疗的机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

N2O差分吸收激光雷达单频激光源频率控制技术研究

国家自然科学基金

0+阅读 · 2015年12月31日

载脂蛋白Eε4与ε2等位基因相反调控晚发性阿尔茨海默病发病风险的神经网络机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于业务和用户认知的蜂窝无线网络能效优化与资源配置技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

OsDCL3b基因调控稻穗生长发育的遗传机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

母体慢性应激对子代注意缺陷多动障碍发生的影响及机制

国家自然科学基金

0+阅读 · 2011年12月31日

500MHz 5-cell超导高频腔高次模抑制方案研究

国家自然科学基金

0+阅读 · 2011年12月31日

新型中红外激光晶体Er3＋:CaReAlO4(Re=Y,Gd)的研究

国家自然科学基金

0+阅读 · 2009年12月31日

Ter94在Hedgehog信号转导途径中的作用机理

国家自然科学基金

0+阅读 · 2009年12月31日

HOXD13与GLI3基因在马蹄内翻足发病机制中的意义研究

国家自然科学基金

0+阅读 · 2009年12月31日

GriT-DBSCAN: A Spatial Clustering Algorithm for Very Large Databases

Arxiv

0+阅读 · 2022年11月6日

FedER: Federated Learning through Experience Replay and Privacy-Preserving Data Synthesis

Arxiv

0+阅读 · 2022年11月4日

An approach for benchmarking the numerical solutions of stochastic compartmental models

Arxiv

0+阅读 · 2022年11月4日

The Benefits of Model-Based Generalization in Reinforcement Learning

Arxiv

0+阅读 · 2022年11月4日

Approximate exploitability: Learning a best response in large games

Arxiv

0+阅读 · 2022年11月3日

Phase Transitions in Learning and Earning under Price Protection Guarantee

Arxiv

0+阅读 · 2022年11月3日

Learning Hypergraphs From Signals With Dual Smoothness Prior

Arxiv

0+阅读 · 2022年11月3日

An Improved Time Feedforward Connections Recurrent Neural Networks

Arxiv

0+阅读 · 2022年11月3日

Delivery by Drones with Arbitrary Energy Consumption Models: A New Formulation Approach

Arxiv

0+阅读 · 2022年11月2日

Faster Meta Update Strategy for Noise-Robust Deep Learning

Arxiv

11+阅读 · 2021年4月30日

VIP会员

文章信息

相关主题

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

大语言模型智能体强化学习：全景综述

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

【伯克利博士论文】从推理服务到训练：面向大规模 LLM 智能体的高效系统

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

GriT-DBSCAN: A Spatial Clustering Algorithm for Very Large Databases

Arxiv

0+阅读 · 2022年11月6日

FedER: Federated Learning through Experience Replay and Privacy-Preserving Data Synthesis

Arxiv

0+阅读 · 2022年11月4日

An approach for benchmarking the numerical solutions of stochastic compartmental models

Arxiv

0+阅读 · 2022年11月4日

The Benefits of Model-Based Generalization in Reinforcement Learning

Arxiv

0+阅读 · 2022年11月4日

Approximate exploitability: Learning a best response in large games

Arxiv

0+阅读 · 2022年11月3日

Phase Transitions in Learning and Earning under Price Protection Guarantee

Arxiv

0+阅读 · 2022年11月3日

Learning Hypergraphs From Signals With Dual Smoothness Prior

Arxiv

0+阅读 · 2022年11月3日

An Improved Time Feedforward Connections Recurrent Neural Networks

Arxiv

0+阅读 · 2022年11月3日

Delivery by Drones with Arbitrary Energy Consumption Models: A New Formulation Approach

Arxiv

0+阅读 · 2022年11月2日

Faster Meta Update Strategy for Noise-Robust Deep Learning

Arxiv

11+阅读 · 2021年4月30日

相关基金

MARVELD1基因调控肝细胞癌介入治疗的机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

N2O差分吸收激光雷达单频激光源频率控制技术研究

国家自然科学基金

0+阅读 · 2015年12月31日

载脂蛋白Eε4与ε2等位基因相反调控晚发性阿尔茨海默病发病风险的神经网络机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于业务和用户认知的蜂窝无线网络能效优化与资源配置技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

OsDCL3b基因调控稻穗生长发育的遗传机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

母体慢性应激对子代注意缺陷多动障碍发生的影响及机制

国家自然科学基金

0+阅读 · 2011年12月31日

500MHz 5-cell超导高频腔高次模抑制方案研究

国家自然科学基金

0+阅读 · 2011年12月31日

新型中红外激光晶体Er3＋:CaReAlO4(Re=Y,Gd)的研究

国家自然科学基金

0+阅读 · 2009年12月31日

Ter94在Hedgehog信号转导途径中的作用机理

国家自然科学基金

0+阅读 · 2009年12月31日

HOXD13与GLI3基因在马蹄内翻足发病机制中的意义研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员