任务不可知离线强化学习计划 (Latent Plans for Task-Agnostic Offline Reinforcement Learning) - 专知论文

会员服务 ·

0

Learning · 潜在 · 强化学习 · 控制器 · Performer ·

2022 年 9 月 19 日

Latent Plans for Task-Agnostic Offline Reinforcement Learning

翻译：任务不可知离线强化学习计划

Erick Rosete-Beas,Oier Mees,Gabriel Kalweit,Joschka Boedecker,Wolfram Burgard

from arxiv, CoRL 2022. Project website: http://tacorl.cs.uni-freiburg.de/

Everyday tasks of long-horizon and comprising a sequence of multiple implicit subtasks still impose a major challenge in offline robot control. While a number of prior methods aimed to address this setting with variants of imitation and offline reinforcement learning, the learned behavior is typically narrow and often struggles to reach configurable long-horizon goals. As both paradigms have complementary strengths and weaknesses, we propose a novel hierarchical approach that combines the strengths of both methods to learn task-agnostic long-horizon policies from high-dimensional camera observations. Concretely, we combine a low-level policy that learns latent skills via imitation learning and a high-level policy learned from offline reinforcement learning for skill-chaining the latent behavior priors. Experiments in various simulated and real robot control tasks show that our formulation enables producing previously unseen combinations of skills to reach temporally extended goals by "stitching" together latent skills through goal chaining with an order-of-magnitude improvement in performance upon state-of-the-art baselines. We even learn one multi-task visuomotor policy for 25 distinct manipulation tasks in the real world which outperforms both imitation learning and offline reinforcement learning techniques.

翻译：长旋线每天的任务由多个隐含子任务序列组成,这在离线机器人控制方面仍构成一项重大挑战。虽然一些先前旨在处理这一背景的方法有模仿和脱线强化学习的变体,但学到的行为一般是狭窄的,而且往往难以达到可配置长旋线的目标。由于这两种模式都具有互补的强项和弱项,我们建议一种新型的等级化方法,将两种方法的优势结合起来,从高维相机观测中学习任务、不可知的长旋线政策。具体地说,我们结合了一种低层次的政策,通过模仿学习和从离线强化学习潜在行为前期技能而学到潜伏技能的高级政策。在各种模拟和真正的机器人控制任务中进行的实验表明,我们的配方能够产生先前看不见的技能组合,通过“缝合”和潜伏技能,通过目标链,在状态基线上提高性能的顺序。我们甚至学习了一种多段对流自动操作政策,用于在现实世界上学习25种截然不同的强化操纵技术。

0

相关内容

Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

专知会员服务

41+阅读 · 2020年4月11日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

专知会员服务

84+阅读 · 2020年2月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

酮体代谢对急性脊髓损伤神经保护作用机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

补肾方调节MAVS介导的信号通路发挥抗炎、抗病毒作用机制及其主体疗效中药组

国家自然科学基金

0+阅读 · 2014年12月31日

复杂压电驱动系统动力学建模、分析与控制

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

无线接入网QoS保证的节能控制建模与协同策略优化

国家自然科学基金

0+阅读 · 2012年12月31日

轻金属硼基氢化物复合材料的制备及储氢性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

硒对血管内皮细胞蛋白质巯基亚硝基化的作用及分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Spiking神经网络在移动机器人感知及控制中的应用研究

国家自然科学基金

1+阅读 · 2011年12月31日

知识地图的拓扑与演化特性研究及在e-Learning中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

认知无线电系统中的非凸优化资源分配算法

国家自然科学基金

1+阅读 · 2009年12月31日

SAM-RL: Sensing-Aware Model-Based Reinforcement Learning via Differentiable Physics-Based Simulation and Rendering

Arxiv

0+阅读 · 2022年10月27日

Rewarding Episodic Visitation Discrepancy for Exploration in Reinforcement Learning

Arxiv

0+阅读 · 2022年10月26日

In-context Reinforcement Learning with Algorithm Distillation

Arxiv

0+阅读 · 2022年10月25日

Entity Divider with Language Grounding in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月25日

MARLlib: Extending RLlib for Multi-agent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月11日

Transformers are Meta-Reinforcement Learners

Arxiv

15+阅读 · 2022年6月14日

MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration

Arxiv

12+阅读 · 2021年2月7日

Learning Latent Representations to Influence Multi-Agent Interaction

Arxiv

11+阅读 · 2020年11月12日

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Arxiv

20+阅读 · 2020年3月10日

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Arxiv

26+阅读 · 2020年2月10日

VIP会员

文章信息

相关主题

相关VIP内容

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

专知会员服务

41+阅读 · 2020年4月11日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

【牛津大学】深度残差强化学习，Deep Residual Reinforcement Learning

专知会员服务

84+阅读 · 2020年2月18日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【ICCV2025教程】基础模型遇见具身智能体

军事机器学习设计：关于开发自动化任务摘要系统的梯次化设计科学研究 | 2025最新93页

扩散模型中的缓存方法综述：迈向高效的多模态生成

【ICCV2025教程】《迈向视觉语言模型的全面推理》

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

相关论文

SAM-RL: Sensing-Aware Model-Based Reinforcement Learning via Differentiable Physics-Based Simulation and Rendering

Arxiv

0+阅读 · 2022年10月27日

Rewarding Episodic Visitation Discrepancy for Exploration in Reinforcement Learning

Arxiv

0+阅读 · 2022年10月26日

In-context Reinforcement Learning with Algorithm Distillation

Arxiv

0+阅读 · 2022年10月25日

Entity Divider with Language Grounding in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月25日

MARLlib: Extending RLlib for Multi-agent Reinforcement Learning

Arxiv

0+阅读 · 2022年10月11日

Transformers are Meta-Reinforcement Learners

Arxiv

15+阅读 · 2022年6月14日

MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration

Arxiv

12+阅读 · 2021年2月7日

Learning Latent Representations to Influence Multi-Agent Interaction

Arxiv

11+阅读 · 2020年11月12日

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Arxiv

20+阅读 · 2020年3月10日

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Q-value Path Decomposition for Deep Multiagent Reinforcement Learning

Arxiv

26+阅读 · 2020年2月10日

相关基金

酮体代谢对急性脊髓损伤神经保护作用机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

补肾方调节MAVS介导的信号通路发挥抗炎、抗病毒作用机制及其主体疗效中药组

国家自然科学基金

0+阅读 · 2014年12月31日

复杂压电驱动系统动力学建模、分析与控制

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

无线接入网QoS保证的节能控制建模与协同策略优化

国家自然科学基金

0+阅读 · 2012年12月31日

轻金属硼基氢化物复合材料的制备及储氢性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

硒对血管内皮细胞蛋白质巯基亚硝基化的作用及分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Spiking神经网络在移动机器人感知及控制中的应用研究

国家自然科学基金

1+阅读 · 2011年12月31日

知识地图的拓扑与演化特性研究及在e-Learning中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

认知无线电系统中的非凸优化资源分配算法

国家自然科学基金

1+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员