从离线强化学习中微调：挑战、权衡和实际解决方案 (Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions) - 专知论文

会员服务 ·

0

离线强化学习 · 微调 · 在线 · 强化学习 · 性能提升 ·

2023 年 3 月 30 日

Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

翻译：从离线强化学习中微调：挑战、权衡和实际解决方案

Yicheng Luo,Jackie Kay,Edward Grefenstette,Marc Peter Deisenroth

from arxiv, An abstract of this paper was accepted at RLDM 2022

Offline reinforcement learning (RL) allows for the training of competent agents from offline datasets without any interaction with the environment. Online finetuning of such offline models can further improve performance. But how should we ideally finetune agents obtained from offline RL training? While offline RL algorithms can in principle be used for finetuning, in practice, their online performance improves slowly. In contrast, we show that it is possible to use standard online off-policy algorithms for faster improvement. However, we find this approach may suffer from policy collapse, where the policy undergoes severe performance deterioration during initial online learning. We investigate the issue of policy collapse and how it relates to data diversity, algorithm choices and online replay distribution. Based on these insights, we propose a conservative policy optimization procedure that can achieve stable and sample-efficient online learning from offline pretraining.

翻译：离线强化学习允许在不与环境交互的情况下训练具有竞争力的智能体。对这些离线训练模型在线微调可以进一步提高性能。但是，我们应该如何理想地微调离线强化学习训练得到的智能体呢？离线强化学习算法原则上可以用于微调，但在实践中，它们的在线性能提升速度较慢。相反，我们发现可以使用标准的在线离策略算法实现更快的性能提升。然而，我们还发现这种方法可能受到策略崩溃的影响，即策略在初始的在线学习过程中会遭受严重的性能恶化。我们研究了策略崩溃问题及其与数据多样性、算法选择和在线回放分布的关系。基于这些见解，我们提出了一种谨慎的策略优化过程，可以实现离线预训练后稳定而且高效的在线学习。

0

相关内容

离线强化学习

离线强化学习

【AAAI2022】一种基于状态扰动的鲁棒强化学习算法

【AAAI2022】一种基于状态扰动的鲁棒强化学习算法

专知会员服务

35+阅读 · 2022年1月31日

【AAAI2021】自校正Q学习，Self-correcting Q-Learning

专知会员服务

17+阅读 · 2020年12月4日

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

专知会员服务

17+阅读 · 2020年7月14日

【CMU博士论文】用动态超参数优化改进深度学习训练和推理，Improving Deep Learning Training and Inference with Dynamic Hyperparameter Optimization

【CMU博士论文】用动态超参数优化改进深度学习训练和推理，Improving Deep Learning Training and Inference with Dynamic Hyperparameter Optimization

专知会员服务

55+阅读 · 2020年5月26日

【综述】生成式对抗网络(GANs)最新2020综述:挑战、解决方案和未来方向，Generative Adversarial Networks (GANs): Challenges, Solutions, and Future Directions

【综述】生成式对抗网络(GANs)最新2020综述:挑战、解决方案和未来方向，Generative Adversarial Networks (GANs): Challenges, Solutions, and Future Directions

专知会员服务

63+阅读 · 2020年5月12日

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

专知会员服务

32+阅读 · 2020年2月1日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

富信息环境下复杂可修系统动态维修决策研究

国家自然科学基金

3+阅读 · 2015年12月31日

虚拟人维修动作混合驱动及人机工效自动量化评估方法研究

国家自然科学基金

3+阅读 · 2014年12月31日

多跳认知无线电网络动态信道接入问题研究

国家自然科学基金

0+阅读 · 2014年12月31日

miR-628促进烧伤后胰岛素抵抗和骨骼肌消耗的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Nogo-A调控阿尔茨海默病Aβ代谢的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

碱性环境下高稳定性γ-谷氨酰转肽酶的分子设计与构建

国家自然科学基金

0+阅读 · 2012年12月31日

纠缠在实际应用中的动力学稳健性及相关问题的研究

国家自然科学基金

0+阅读 · 2011年12月31日

miR-124和miR-27对阿尔茨海默病BACE1基因影响的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

LTE-Advanced中继网络关键技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

心肌细胞新陈代谢系统中的若干动力学问题

国家自然科学基金

0+阅读 · 2009年12月31日

On Text-based Personality Computing: Challenges and Future Directions

Arxiv

0+阅读 · 2023年5月18日

Contrastive State Augmentations for Reinforcement Learning-Based Recommender Systems

Arxiv

0+阅读 · 2023年5月18日

Semantically Aligned Task Decomposition in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年5月18日

Automatic Design Method of Building Pipeline Layout Based on Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年5月18日

Discovering Individual Rewards in Collective Behavior through Inverse Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年5月17日

Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning

Arxiv

0+阅读 · 2023年5月17日

Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges

Arxiv

19+阅读 · 2023年5月17日

A Survey on Transformers in Reinforcement Learning

Arxiv

31+阅读 · 2023年1月8日

A Comprehensive Survey on Deep Clustering: Taxonomy, Challenges, and Future Directions

Arxiv

42+阅读 · 2022年6月15日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

VIP会员

文章信息

相关主题

离线强化学习

相关VIP内容

【AAAI2022】一种基于状态扰动的鲁棒强化学习算法

【AAAI2022】一种基于状态扰动的鲁棒强化学习算法

专知会员服务

35+阅读 · 2022年1月31日

【AAAI2021】自校正Q学习，Self-correcting Q-Learning

专知会员服务

17+阅读 · 2020年12月4日

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

专知会员服务

17+阅读 · 2020年7月14日

【CMU博士论文】用动态超参数优化改进深度学习训练和推理，Improving Deep Learning Training and Inference with Dynamic Hyperparameter Optimization

【CMU博士论文】用动态超参数优化改进深度学习训练和推理，Improving Deep Learning Training and Inference with Dynamic Hyperparameter Optimization

专知会员服务

55+阅读 · 2020年5月26日

【综述】生成式对抗网络(GANs)最新2020综述:挑战、解决方案和未来方向，Generative Adversarial Networks (GANs): Challenges, Solutions, and Future Directions

【综述】生成式对抗网络(GANs)最新2020综述:挑战、解决方案和未来方向，Generative Adversarial Networks (GANs): Challenges, Solutions, and Future Directions

专知会员服务

63+阅读 · 2020年5月12日

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

专知会员服务

32+阅读 · 2020年2月1日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

新质生成式AI赋能产业变革的实践与路径

用于多模态大模型的离散标记化：全面综述

Nature综述：金融网络中的物理学

【CMU博士论文】通信高效且差分隐私的优化方法

相关资讯

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

On Text-based Personality Computing: Challenges and Future Directions

Arxiv

0+阅读 · 2023年5月18日

Contrastive State Augmentations for Reinforcement Learning-Based Recommender Systems

Arxiv

0+阅读 · 2023年5月18日

Semantically Aligned Task Decomposition in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年5月18日

Automatic Design Method of Building Pipeline Layout Based on Deep Reinforcement Learning

Arxiv

0+阅读 · 2023年5月18日

Discovering Individual Rewards in Collective Behavior through Inverse Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年5月17日

Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning

Arxiv

0+阅读 · 2023年5月17日

Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges

Arxiv

19+阅读 · 2023年5月17日

A Survey on Transformers in Reinforcement Learning

Arxiv

31+阅读 · 2023年1月8日

A Comprehensive Survey on Deep Clustering: Taxonomy, Challenges, and Future Directions

Arxiv

42+阅读 · 2022年6月15日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

相关基金

富信息环境下复杂可修系统动态维修决策研究

国家自然科学基金

3+阅读 · 2015年12月31日

虚拟人维修动作混合驱动及人机工效自动量化评估方法研究

国家自然科学基金

3+阅读 · 2014年12月31日

多跳认知无线电网络动态信道接入问题研究

国家自然科学基金

0+阅读 · 2014年12月31日

miR-628促进烧伤后胰岛素抵抗和骨骼肌消耗的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Nogo-A调控阿尔茨海默病Aβ代谢的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

碱性环境下高稳定性γ-谷氨酰转肽酶的分子设计与构建

国家自然科学基金

0+阅读 · 2012年12月31日

纠缠在实际应用中的动力学稳健性及相关问题的研究

国家自然科学基金

0+阅读 · 2011年12月31日

miR-124和miR-27对阿尔茨海默病BACE1基因影响的分子机制

国家自然科学基金

0+阅读 · 2011年12月31日

LTE-Advanced中继网络关键技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

心肌细胞新陈代谢系统中的若干动力学问题

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员