面具自动算术器作为斯波多时学习者 (Masked Autoencoders As Spatiotemporal Learners) - 专知论文

会员服务 ·

0

掩码 · 自编码器 · 掩码自编码MAE · 学成 · 学习器 ·

2022 年 5 月 18 日

Masked Autoencoders As Spatiotemporal Learners

翻译：面具自动算术器作为斯波多时学习者

Christoph Feichtenhofer,Haoqi Fan,Yanghao Li,Kaiming He

from arxiv, Technical report

This paper studies a conceptually simple extension of Masked Autoencoders (MAE) to spatiotemporal representation learning from videos. We randomly mask out spacetime patches in videos and learn an autoencoder to reconstruct them in pixels. Interestingly, we show that our MAE method can learn strong representations with almost no inductive bias on spacetime (only except for patch and positional embeddings), and spacetime-agnostic random masking performs the best. We observe that the optimal masking ratio is as high as 90% (vs. 75% on images), supporting the hypothesis that this ratio is related to information redundancy of the data. A high masking ratio leads to a large speedup, e.g., > 4x in wall-clock time or even more. We report competitive results on several challenging video datasets using vanilla Vision Transformers. We observe that MAE can outperform supervised pre-training by large margins. We further report encouraging results of training on real-world, uncurated Instagram data. Our study suggests that the general framework of masked autoencoding (BERT, MAE, etc.) can be a unified methodology for representation learning with minimal domain knowledge.

翻译：本文从概念上简单而简单地扩展了蒙面自动涂层(MAE), 从视频中学习时空代表。我们随机在视频中隐藏时空补丁, 并学习一个自动编码器, 以像素形式重建它们。有趣的是, 我们的MAE方法可以学到强烈的表层, 在时空( 只能贴补和定位嵌入) 几乎没有感应偏差, 空间时空随机遮罩能发挥最佳效果。我们观察到, 最佳遮蔽率高达90% ( 图像上为75% ), 支持这一比率与数据信息冗余有关的假设。高掩蔽率会导致大超速, 例如, 在墙上或更长时间超过4x 。我们用香草视觉变异器报告一些具有挑战性的视频数据集的竞争结果。我们观察到, MAE 可以在大范围内完成受监督的预培训。我们进一步报告鼓励在现实世界上开展培训的结果, 未加固的Instagram数据。我们的研究显示, 以最小的域化自动代表方法( BER MAE, 等) 可以采用最小化的通用的自动代表方法。

8

相关内容

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

【何恺明组新论文】掩码自编码器作为时空学习器，Masked Autoencoders As Spatiotemporal Learners

【何恺明组新论文】掩码自编码器作为时空学习器，Masked Autoencoders As Spatiotemporal Learners

专知会员服务

39+阅读 · 2022年5月19日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【快讯】ICML 2020论文出炉，1088篇上榜，你的paper中了吗？

【快讯】ICML 2020论文出炉，1088篇上榜，你的paper中了吗？

专知会员服务

52+阅读 · 2020年6月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

【论文推荐】最新七篇图像分割相关论文—域适应深度表示学习、循环残差卷积、二值分割、图像合成、无监督跨模态

【论文推荐】最新七篇图像分割相关论文—域适应深度表示学习、循环残差卷积、二值分割、图像合成、无监督跨模态

专知

19+阅读 · 2018年6月1日

miR-223调控血小板活化在动脉粥样硬化中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

CTLA-4/B7通路在调节性γδ T细胞调控造血干细胞移植后异源反应性T细胞活性中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

共激活蛋白在视网膜感光细胞发育中的分子调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

浅海低频混响中的杂波特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

ATM介导自噬分子Beclin1磷酸化修饰的新功能解析

国家自然科学基金

0+阅读 · 2012年12月31日

血管平滑肌细胞AMPK活性调节在糖尿病并发动脉粥样硬化中作用及其分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

异位表达钙反向转运体基因提高苹果抗斑点落叶病的机理

国家自然科学基金

0+阅读 · 2011年12月31日

蛋白激酶A（PKA）介导自噬基因Beclin1磷酸化修饰的功能解析

国家自然科学基金

0+阅读 · 2011年12月31日

亚硫酸盐氧化气液反应动力学及传质机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于Sparse-Land模型的SAR图像噪声抑制与分割

国家自然科学基金

0+阅读 · 2009年12月31日

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Arxiv

0+阅读 · 2022年7月7日

MaiT: Leverage Attention Masks for More Efficient Image Transformers

Arxiv

0+阅读 · 2022年7月6日

Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation

Arxiv

0+阅读 · 2022年7月6日

Multi-Modal Masked Pre-Training for Monocular Panoramic Depth Completion

Arxiv

0+阅读 · 2022年7月6日

Adversarial Masking for Self-Supervised Learning

Adversarial Masking for Self-Supervised Learning

Arxiv

0+阅读 · 2022年7月6日

Masked Generative Distillation

Arxiv

0+阅读 · 2022年7月5日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

Generative Models as a Data Source for Multiview Representation Learning

Arxiv

16+阅读 · 2021年6月9日

UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

Arxiv

15+阅读 · 2020年2月28日

Pre-Training with Whole Word Masking for Chinese BERT

Arxiv

11+阅读 · 2019年6月19日

VIP会员

文章信息

相关主题

掩码自编码MAE

相关VIP内容

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

【何恺明组新论文】掩码自编码器作为时空学习器，Masked Autoencoders As Spatiotemporal Learners

【何恺明组新论文】掩码自编码器作为时空学习器，Masked Autoencoders As Spatiotemporal Learners

专知会员服务

39+阅读 · 2022年5月19日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【快讯】ICML 2020论文出炉，1088篇上榜，你的paper中了吗？

【快讯】ICML 2020论文出炉，1088篇上榜，你的paper中了吗？

专知会员服务

52+阅读 · 2020年6月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

【论文推荐】最新七篇图像分割相关论文—域适应深度表示学习、循环残差卷积、二值分割、图像合成、无监督跨模态

【论文推荐】最新七篇图像分割相关论文—域适应深度表示学习、循环残差卷积、二值分割、图像合成、无监督跨模态

专知

19+阅读 · 2018年6月1日

相关论文

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Arxiv

0+阅读 · 2022年7月7日

MaiT: Leverage Attention Masks for More Efficient Image Transformers

Arxiv

0+阅读 · 2022年7月6日

Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation

Arxiv

0+阅读 · 2022年7月6日

Multi-Modal Masked Pre-Training for Monocular Panoramic Depth Completion

Arxiv

0+阅读 · 2022年7月6日

Adversarial Masking for Self-Supervised Learning

Adversarial Masking for Self-Supervised Learning

Arxiv

0+阅读 · 2022年7月6日

Masked Generative Distillation

Arxiv

0+阅读 · 2022年7月5日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

Generative Models as a Data Source for Multiview Representation Learning

Arxiv

16+阅读 · 2021年6月9日

UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

Arxiv

15+阅读 · 2020年2月28日

Pre-Training with Whole Word Masking for Chinese BERT

Arxiv

11+阅读 · 2019年6月19日

相关基金

miR-223调控血小板活化在动脉粥样硬化中的作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

CTLA-4/B7通路在调节性γδ T细胞调控造血干细胞移植后异源反应性T细胞活性中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

共激活蛋白在视网膜感光细胞发育中的分子调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

浅海低频混响中的杂波特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

ATM介导自噬分子Beclin1磷酸化修饰的新功能解析

国家自然科学基金

0+阅读 · 2012年12月31日

血管平滑肌细胞AMPK活性调节在糖尿病并发动脉粥样硬化中作用及其分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

异位表达钙反向转运体基因提高苹果抗斑点落叶病的机理

国家自然科学基金

0+阅读 · 2011年12月31日

蛋白激酶A（PKA）介导自噬基因Beclin1磷酸化修饰的功能解析

国家自然科学基金

0+阅读 · 2011年12月31日

亚硫酸盐氧化气液反应动力学及传质机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

基于Sparse-Land模型的SAR图像噪声抑制与分割

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员