学习通过观看YouTube视频驱动:行动-有条件的违反政策预设培训 (Learning to Drive by Watching YouTube Videos: Action-Conditioned Contrastive Policy Pretraining) - 专知论文

会员服务 ·

0

Learning · contrastive · YouTube · Weight · 无监督特征学习 ·

2022 年 7 月 17 日

Learning to Drive by Watching YouTube Videos: Action-Conditioned Contrastive Policy Pretraining

翻译：学习通过观看YouTube视频驱动:行动-有条件的违反政策预设培训

Qihang Zhang,Zhenghao Peng,Bolei Zhou

from arxiv, ECCV accepted paper

Deep visuomotor policy learning, which aims to map raw visual observation to action, achieves promising results in control tasks such as robotic manipulation and autonomous driving. However, it requires a huge number of online interactions with the training environment, which limits its real-world application. Compared to the popular unsupervised feature learning for visual recognition, feature pretraining for visuomotor control tasks is much less explored. In this work, we aim to pretrain policy representations for driving tasks by watching hours-long uncurated YouTube videos. Specifically, we train an inverse dynamic model with a small amount of labeled data and use it to predict action labels for all the YouTube video frames. A new contrastive policy pretraining method is then developed to learn action-conditioned features from the video frames with pseudo action labels. Experiments show that the resulting action-conditioned features obtain substantial improvements for the downstream reinforcement learning and imitation learning tasks, outperforming the weights pretrained from previous unsupervised learning methods and ImageNet pretrained weight. Code, model weights, and data are available at: https://metadriverse.github.io/ACO.

翻译：深潜运动政策学习旨在绘制原始视觉观察结果,从而在机器人操纵和自主驾驶等控制任务中取得有希望的成果。然而,它需要大量与培训环境进行在线互动,这限制了培训环境的实际应用。与普通的未经监督的视觉识别特征学习相比,对生动控制任务的特质培训远没有那么深入探讨。在这项工作中,我们的目标是通过观看未完成的YouTube视频,为驾驶任务预设政策说明。具体地说,我们用少量标签数据来培训反动态模型,并用它来预测所有YouTube视频框架的行动标签。然后开发了新的对比性政策预培训方法,从视频框架中学习带有行动标志的具有行动条件的特征。实验表明,由此产生的具有行动条件的特征大大改进了下游强化学习和仿造学习任务,超过了先前未完成的学习方法和图像网络预先训练的重量。代码、模型重量和数据见:https://metadriverse.github.io/CO。

0

相关内容

Learning

计算机科学课程与视频课件合集，Computer Science courses with video lectures

计算机科学课程与视频课件合集，Computer Science courses with video lectures

专知会员服务

37+阅读 · 2022年1月24日

【伯克利-Pieter Abbeel】深度强化学习基础，附slides与视频

专知会员服务

29+阅读 · 2021年8月26日

不可错过！MILA最新《自监督表示学习》课程，附PPT与视频下载

不可错过！MILA最新《自监督表示学习》课程，附PPT与视频下载

专知会员服务

90+阅读 · 2020年12月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

TRAF3IP3调控T细胞活性与肿瘤免疫的分子机制

国家自然科学基金

0+阅读 · 2016年12月31日

克罗恩病中干预TLE1逆转凋亡介导的肠道粘膜自噬紊乱的策略研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于自供电磁流变阻尼器的斜拉索减振理论与试验研究

国家自然科学基金

0+阅读 · 2013年12月31日

二维过渡金属硫族化合物自旋及能谷电子学研究

国家自然科学基金

0+阅读 · 2013年12月31日

Bi4Ti3O12基Aurivillius化合物设计、结构调控和性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

组蛋白甲基化修饰调控拟南芥冷响应基因TCF1的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Clusterin通过线粒体凋亡通路调节肝细胞肝癌化疗耐受机理的研究

国家自然科学基金

0+阅读 · 2011年12月31日

高含量过渡金属元素的非晶Ge基磁性半导体的微结构、磁性和输运研究

国家自然科学基金

0+阅读 · 2011年12月31日

哮喘中Notch1对T细胞分化作用的调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于多特征情感信息融合的高效率e-Learning关键技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Arxiv

0+阅读 · 2022年9月13日

Adversarial Coreset Selection for Efficient Robust Training

Arxiv

0+阅读 · 2022年9月13日

Video Summarization Based on Video-text Modelling

Arxiv

0+阅读 · 2022年9月13日

An Investigation of Smart Contract for Collaborative Machine Learning Model Training

Arxiv

0+阅读 · 2022年9月12日

Meta-Reinforcement Learning via Language Instructions

Arxiv

1+阅读 · 2022年9月11日

Saliency Guided Adversarial Training for Learning Generalizable Features with Applications to Medical Imaging Classification System

Arxiv

0+阅读 · 2022年9月9日

Condensing Graphs via One-Step Gradient Matching

Arxiv

0+阅读 · 2022年9月8日

ContrastMask: Contrastive Learning to Segment Every Thing

Arxiv

15+阅读 · 2022年3月18日

Cross-Modal Discrete Representation Learning

Arxiv

18+阅读 · 2021年6月10日

A Simple Framework for Contrastive Learning of Visual Representations

Arxiv

21+阅读 · 2020年2月13日

VIP会员

文章信息

相关主题

无监督特征学习

相关VIP内容

计算机科学课程与视频课件合集，Computer Science courses with video lectures

计算机科学课程与视频课件合集，Computer Science courses with video lectures

专知会员服务

37+阅读 · 2022年1月24日

【伯克利-Pieter Abbeel】深度强化学习基础，附slides与视频

专知会员服务

29+阅读 · 2021年8月26日

不可错过！MILA最新《自监督表示学习》课程，附PPT与视频下载

不可错过！MILA最新《自监督表示学习》课程，附PPT与视频下载

专知会员服务

90+阅读 · 2020年12月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

美军“泰坦（TITAN）地面站目标系统”：是颠覆还是一场可预见的军事进步？

美空军指挥参谋学院 · 联合空中作战规划课程介绍（2025年） | 22页

一种基于视觉算法生成三维场景重建的多任务系统 | 2025最新200页

北约第十七届（2025年）网络冲突国际会议论文集 | 272页

相关资讯

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

相关论文

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Arxiv

0+阅读 · 2022年9月13日

Adversarial Coreset Selection for Efficient Robust Training

Arxiv

0+阅读 · 2022年9月13日

Video Summarization Based on Video-text Modelling

Arxiv

0+阅读 · 2022年9月13日

An Investigation of Smart Contract for Collaborative Machine Learning Model Training

Arxiv

0+阅读 · 2022年9月12日

Meta-Reinforcement Learning via Language Instructions

Arxiv

1+阅读 · 2022年9月11日

Saliency Guided Adversarial Training for Learning Generalizable Features with Applications to Medical Imaging Classification System

Arxiv

0+阅读 · 2022年9月9日

Condensing Graphs via One-Step Gradient Matching

Arxiv

0+阅读 · 2022年9月8日

ContrastMask: Contrastive Learning to Segment Every Thing

Arxiv

15+阅读 · 2022年3月18日

Cross-Modal Discrete Representation Learning

Arxiv

18+阅读 · 2021年6月10日

A Simple Framework for Contrastive Learning of Visual Representations

Arxiv

21+阅读 · 2020年2月13日

相关基金

TRAF3IP3调控T细胞活性与肿瘤免疫的分子机制

国家自然科学基金

0+阅读 · 2016年12月31日

克罗恩病中干预TLE1逆转凋亡介导的肠道粘膜自噬紊乱的策略研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于自供电磁流变阻尼器的斜拉索减振理论与试验研究

国家自然科学基金

0+阅读 · 2013年12月31日

二维过渡金属硫族化合物自旋及能谷电子学研究

国家自然科学基金

0+阅读 · 2013年12月31日

Bi4Ti3O12基Aurivillius化合物设计、结构调控和性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

组蛋白甲基化修饰调控拟南芥冷响应基因TCF1的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Clusterin通过线粒体凋亡通路调节肝细胞肝癌化疗耐受机理的研究

国家自然科学基金

0+阅读 · 2011年12月31日

高含量过渡金属元素的非晶Ge基磁性半导体的微结构、磁性和输运研究

国家自然科学基金

0+阅读 · 2011年12月31日

哮喘中Notch1对T细胞分化作用的调控研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于多特征情感信息融合的高效率e-Learning关键技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员