追回被盗资产举措前身:变革者,在视觉强化学习方面有国家行动奖励代表 (StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning) - 专知论文

会员服务 ·

0

Learning · 变换 · 表示 · 强化学习 · 归纳偏好 ·

2023 年 1 月 4 日

StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning

翻译：追回被盗资产举措前身:变革者,在视觉强化学习方面有国家行动奖励代表

Jinghuan Shang,Kumara Kahatapitiya,Xiang Li,Michael S. Ryoo

from arxiv, Accepted to ECCV 2022. Our code is available at https://github.com/elicassion/StARformer

Reinforcement Learning (RL) can be considered as a sequence modeling task: given a sequence of past state-action-reward experiences, an agent predicts a sequence of next actions. In this work, we propose State-Action-Reward Transformer (StARformer) for visual RL, which explicitly models short-term state-action-reward representations (StAR-representations), essentially introducing a Markovian-like inductive bias to improve long-term modeling. Our approach first extracts StAR-representations by self-attending image state patches, action, and reward tokens within a short temporal window. These are then combined with pure image state representations -- extracted as convolutional features, to perform self-attention over the whole sequence. Our experiments show that StARformer outperforms the state-of-the-art Transformer-based method on image-based Atari and DeepMind Control Suite benchmarks, in both offline-RL and imitation learning settings. StARformer is also more compliant with longer sequences of inputs. Our code is available at https://github.com/elicassion/StARformer.

翻译：强化学习(RL)可被视为一个序列建模任务:根据过去国家行动回报经验的顺序,一个代理预测下一个行动的顺序。在这项工作中,我们为视觉RL提议国家行动回报变换器(StARexer),该变换器明确模拟短期国家行动回报表(StAR-respresentations),基本上引入了类似于Markovian的诱导偏向,以改进长期建模。我们的方法首先通过在短时间窗口内自动显示图像状态的补丁、动作和奖赏符号来提取追回被盗资产。然后,这些都与纯粹的图像表现(作为革命性特征提取)相结合,对整个序列进行自我注意。我们的实验显示,追回被盗资产变换器在离线和模拟学习环境中都比基于图像的以阿塔里和深海控制套装基准更优异。追回资产前还更符合较长的输入序列。我们的代码可在 https://githhub.com/elicasion/Strave中查阅。

0

相关内容

Learning

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

专知会员服务

41+阅读 · 2020年4月11日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

新书分享：强化学习最新书稿《强化学习导论》（Reinforcement Learning An Introduction）第二版出炉

新书分享：强化学习最新书稿《强化学习导论》（Reinforcement Learning An Introduction）第二版出炉

专知会员服务

118+阅读 · 2019年10月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

Caspases切割ARF-BP1调控p53信号通路的机制

国家自然科学基金

0+阅读 · 2014年12月31日

CD8+CD28- T 细胞在激素非依赖性前列腺癌形成中的调控作用

国家自然科学基金

0+阅读 · 2012年12月31日

PTBP1介导的survivinΔEx3过表达调控胶质母细胞瘤微血管增生的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

基质中成纤维细胞在乳腺癌内分泌耐药中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

Skp2泛素化调控Aurora B的作用与机制

国家自然科学基金

0+阅读 · 2012年12月31日

肿瘤细胞中凋亡抑制蛋白CFLAR乙酰化调控的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

γ#27688;基丁酸通过肿瘤抗原TRAK1(MGb2-Ag)调控胃癌细胞生长的机制

国家自然科学基金

0+阅读 · 2009年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

去乙酰化转移酶（HDAC)抑制剂MS-275对胃癌细胞的选择性杀伤作用及机制

国家自然科学基金

0+阅读 · 2009年12月31日

食管癌细胞中PI3K/AKT-HIF1α36890;路对糖酵解的影响

国家自然科学基金

0+阅读 · 2008年12月31日

Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning

Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning

Arxiv

0+阅读 · 2023年3月2日

Reinforced Labels: Multi-Agent Deep Reinforcement Learning for Point-feature Label Placement

Arxiv

0+阅读 · 2023年3月2日

Masked Distillation with Receptive Tokens

Arxiv

0+阅读 · 2023年3月2日

Expert-Free Online Transfer Learning in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年3月2日

Preference Transformer: Modeling Human Preferences using Transformers for RL

Arxiv

0+阅读 · 2023年3月2日

Does Zero-Shot Reinforcement Learning Exist?

Arxiv

0+阅读 · 2023年3月1日

CRC-RL: A Novel Visual Feature Representation Architecture for Unsupervised Reinforcement Learning

Arxiv

0+阅读 · 2023年3月1日

STIR$^2$: Reward Relabelling for combined Reinforcement and Imitation Learning on sparse-reward tasks

Arxiv

0+阅读 · 2023年2月28日

Behavior Prior Representation learning for Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月28日

CURL: Contrastive Unsupervised Representations for Reinforcement Learning

Arxiv

17+阅读 · 2020年4月28日

VIP会员

文章信息

相关主题

相关VIP内容

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

专知会员服务

41+阅读 · 2020年4月11日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

新书分享：强化学习最新书稿《强化学习导论》（Reinforcement Learning An Introduction）第二版出炉

新书分享：强化学习最新书稿《强化学习导论》（Reinforcement Learning An Introduction）第二版出炉

专知会员服务

118+阅读 · 2019年10月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

小规模训练指南：打造世界级大语言模型的关键方法

无人机编队飞行：复杂环境中作战的策略、挑战与应用

大模型APP，AI时代第一个爆款

从数据中心视角出发的高效大语言模型训练综述

相关资讯

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning

Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning

Arxiv

0+阅读 · 2023年3月2日

Reinforced Labels: Multi-Agent Deep Reinforcement Learning for Point-feature Label Placement

Arxiv

0+阅读 · 2023年3月2日

Masked Distillation with Receptive Tokens

Arxiv

0+阅读 · 2023年3月2日

Expert-Free Online Transfer Learning in Multi-Agent Reinforcement Learning

Arxiv

0+阅读 · 2023年3月2日

Preference Transformer: Modeling Human Preferences using Transformers for RL

Arxiv

0+阅读 · 2023年3月2日

Does Zero-Shot Reinforcement Learning Exist?

Arxiv

0+阅读 · 2023年3月1日

CRC-RL: A Novel Visual Feature Representation Architecture for Unsupervised Reinforcement Learning

Arxiv

0+阅读 · 2023年3月1日

STIR$^2$: Reward Relabelling for combined Reinforcement and Imitation Learning on sparse-reward tasks

Arxiv

0+阅读 · 2023年2月28日

Behavior Prior Representation learning for Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年2月28日

CURL: Contrastive Unsupervised Representations for Reinforcement Learning

Arxiv

17+阅读 · 2020年4月28日

相关基金

Caspases切割ARF-BP1调控p53信号通路的机制

国家自然科学基金

0+阅读 · 2014年12月31日

CD8+CD28- T 细胞在激素非依赖性前列腺癌形成中的调控作用

国家自然科学基金

0+阅读 · 2012年12月31日

PTBP1介导的survivinΔEx3过表达调控胶质母细胞瘤微血管增生的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

基质中成纤维细胞在乳腺癌内分泌耐药中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

Skp2泛素化调控Aurora B的作用与机制

国家自然科学基金

0+阅读 · 2012年12月31日

肿瘤细胞中凋亡抑制蛋白CFLAR乙酰化调控的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

γ#27688;基丁酸通过肿瘤抗原TRAK1(MGb2-Ag)调控胃癌细胞生长的机制

国家自然科学基金

0+阅读 · 2009年12月31日

UGT基因簇进化及调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

去乙酰化转移酶（HDAC)抑制剂MS-275对胃癌细胞的选择性杀伤作用及机制

国家自然科学基金

0+阅读 · 2009年12月31日

食管癌细胞中PI3K/AKT-HIF1α36890;路对糖酵解的影响

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员