通过数据重新平衡促进离线强化学习 (Boosting Offline Reinforcement Learning via Data Rebalancing) - 专知论文

会员服务 ·

0

Boosting（一种模型训练加速方式） · Learning · Performer · 约束 · 强化学习 ·

2022 年 10 月 17 日

Boosting Offline Reinforcement Learning via Data Rebalancing

翻译：通过数据重新平衡促进离线强化学习

Yang Yue,Bingyi Kang,Xiao Ma,Zhongwen Xu,Gao Huang,Shuicheng Yan

from arxiv, 8 pages, 2 figures

Offline reinforcement learning (RL) is challenged by the distributional shift between learning policies and datasets. To address this problem, existing works mainly focus on designing sophisticated algorithms to explicitly or implicitly constrain the learned policy to be close to the behavior policy. The constraint applies not only to well-performing actions but also to inferior ones, which limits the performance upper bound of the learned policy. Instead of aligning the densities of two distributions, aligning the supports gives a relaxed constraint while still being able to avoid out-of-distribution actions. Therefore, we propose a simple yet effective method to boost offline RL algorithms based on the observation that resampling a dataset keeps the distribution support unchanged. More specifically, we construct a better behavior policy by resampling each transition in an old dataset according to its episodic return. We dub our method ReD (Return-based Data Rebalance), which can be implemented with less than 10 lines of code change and adds negligible running time. Extensive experiments demonstrate that ReD is effective at boosting offline RL performance and orthogonal to decoupling strategies in long-tailed classification. New state-of-the-arts are achieved on the D4RL benchmark.

翻译：离线强化学习( RL) 受到学习政策和数据集之间分布变化的挑战。解决这个问题, 现有工作主要侧重于设计精密的算法, 以明示或隐含的方式限制所学政策接近行为政策。限制不仅适用于表现良好的行动,也适用于劣等行动,这限制了所学政策上层的性能。不调整两个分布的密度, 调整支持会放松限制, 同时仍然能够避免分配之外的行动。因此, 我们提出一个简单而有效的方法, 以重现数据集使分布支持保持不变的观察为基础, 来提升离线的 RL 算法。更具体地说, 我们制定更好的行为政策, 根据旧数据集的回归, 重新标注每个转换。我们用的方法( 以翻转为基础的数据重新平衡) 重新定位, 可以用不到10行的代码变化来实施, 并增加微小的运行时间。因此, 广泛的实验证明 ReD 能够有效地提升离线 RL 性运行, 以及或Thotognal 来在长期链接的分类中解析战略。

0

相关内容

Boosting（一种模型训练加速方式）

Boosting（一种模型训练加速方式）

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

RIP1/RIP3通路调控脑出血后细胞程序性坏死的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

苹果氮响应蛋白MdBT4与MdJAZ2互作调控花氰苷合成和果实着色的研究

国家自然科学基金

0+阅读 · 2014年12月31日

黑果枸杞果实成熟介导花青素合成的分子机制解析

国家自然科学基金

0+阅读 · 2014年12月31日

光电化学分解水制氢用Cu2O基光电极的晶体取向和界面结构的协同调控

国家自然科学基金

0+阅读 · 2013年12月31日

一类Monge-Ampère方程解的边界行为

国家自然科学基金

0+阅读 · 2013年12月31日

肺腺癌中SPLUNC1基因启动子区Snail调控元件功能的研究

国家自然科学基金

0+阅读 · 2011年12月31日

油菜BnICE1基因与MAP激酶信号途径在调控植物耐寒性中的相互作用研究

国家自然科学基金

0+阅读 · 2011年12月31日

以EGFR为识别靶位多靶点联合克服NSCLC EGFR TKIs耐药的基因干预研究

国家自然科学基金

0+阅读 · 2011年12月31日

ATP和ROS在BCL-2基因抑癌活性中的作用机制

国家自然科学基金

0+阅读 · 2011年12月31日

PGRMC1蛋白在肾癌中的功能及作用机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

Learning Efficient Policies for Picking Entangled Wire Harnesses: A Solution to Industrial Bin Picking

Arxiv

0+阅读 · 2022年11月21日

Prediction-aware and Reinforcement Learning based Altruistic Cooperative Driving

Arxiv

0+阅读 · 2022年11月19日

Fast Lifelong Adaptive Inverse Reinforcement Learning from Demonstrations

Arxiv

0+阅读 · 2022年11月19日

Exploring with Sticky Mittens: Reinforcement Learning with Expert Interventions via Option Templates

Arxiv

0+阅读 · 2022年11月17日

Controllable Data Generation by Deep Learning: A Review

Arxiv

15+阅读 · 2022年7月19日

Reinforcement Learning on Graph: A Survey

Arxiv

67+阅读 · 2022年4月13日

Recent Advances in Reinforcement Learning in Finance

Arxiv

11+阅读 · 2021年12月8日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

Learning Latent Representations to Influence Multi-Agent Interaction

Arxiv

11+阅读 · 2020年11月12日

A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions

A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions

Arxiv

79+阅读 · 2020年1月19日

VIP会员

文章信息

相关主题

Boosting（一种模型训练加速方式）

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】数据驱动决策中的激励、信息与不确定性

DGP双粒度提示框架：图增强大模型助力欺诈检测

【ICCV2025】ESSENTIAL：用于视频类增量学习的情景记忆与语义记忆整合

唯快不破：大型语言模型高效架构综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

相关论文

Learning Efficient Policies for Picking Entangled Wire Harnesses: A Solution to Industrial Bin Picking

Arxiv

0+阅读 · 2022年11月21日

Prediction-aware and Reinforcement Learning based Altruistic Cooperative Driving

Arxiv

0+阅读 · 2022年11月19日

Fast Lifelong Adaptive Inverse Reinforcement Learning from Demonstrations

Arxiv

0+阅读 · 2022年11月19日

Exploring with Sticky Mittens: Reinforcement Learning with Expert Interventions via Option Templates

Arxiv

0+阅读 · 2022年11月17日

Controllable Data Generation by Deep Learning: A Review

Arxiv

15+阅读 · 2022年7月19日

Reinforcement Learning on Graph: A Survey

Arxiv

67+阅读 · 2022年4月13日

Recent Advances in Reinforcement Learning in Finance

Arxiv

11+阅读 · 2021年12月8日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

Learning Latent Representations to Influence Multi-Agent Interaction

Arxiv

11+阅读 · 2020年11月12日

A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions

A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions

Arxiv

79+阅读 · 2020年1月19日

相关基金

RIP1/RIP3通路调控脑出血后细胞程序性坏死的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

苹果氮响应蛋白MdBT4与MdJAZ2互作调控花氰苷合成和果实着色的研究

国家自然科学基金

0+阅读 · 2014年12月31日

黑果枸杞果实成熟介导花青素合成的分子机制解析

国家自然科学基金

0+阅读 · 2014年12月31日

光电化学分解水制氢用Cu2O基光电极的晶体取向和界面结构的协同调控

国家自然科学基金

0+阅读 · 2013年12月31日

一类Monge-Ampère方程解的边界行为

国家自然科学基金

0+阅读 · 2013年12月31日

肺腺癌中SPLUNC1基因启动子区Snail调控元件功能的研究

国家自然科学基金

0+阅读 · 2011年12月31日

油菜BnICE1基因与MAP激酶信号途径在调控植物耐寒性中的相互作用研究

国家自然科学基金

0+阅读 · 2011年12月31日

以EGFR为识别靶位多靶点联合克服NSCLC EGFR TKIs耐药的基因干预研究

国家自然科学基金

0+阅读 · 2011年12月31日

ATP和ROS在BCL-2基因抑癌活性中的作用机制

国家自然科学基金

0+阅读 · 2011年12月31日

PGRMC1蛋白在肾癌中的功能及作用机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员