ConvMAE: 蒙面革命与蒙面的自动编码器会面 (ConvMAE: Masked Convolution Meets Masked Autoencoders) - 专知论文

会员服务 ·

0

掩码 · 自编码器 · 轮 · 卷积 · Vision ·

2022 年 5 月 8 日

ConvMAE: Masked Convolution Meets Masked Autoencoders

翻译：ConvMAE: 蒙面革命与蒙面的自动编码器会面

Peng Gao,Teli Ma,Hongsheng Li,Jifeng Dai,Yu Qiao

from arxiv, 10 pages

Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer architectures can further unleash the potentials of ViT, leading to state-of-the-art performances on image classification, detection and semantic segmentation. In this paper, our ConvMAE framework demonstrates that multi-scale hybrid convolution-transformer can learn more discriminative representations via the mask auto-encoding scheme. However, directly using the original masking strategy leads to the heavy computational cost and pretraining-finetuning discrepancy. To tackle the issue, we adopt the masked convolution to prevent information leakage in the convolution blocks. A simple block-wise masking strategy is proposed to ensure computational efficiency. We also propose to more directly supervise the multi-scale features of the encoder to boost multi-scale features. Based on our pretrained ConvMAE models, ConvMAE-Base improves ImageNet-1K finetuning accuracy by 1.4% compared with MAE-Base. On object detection, ConvMAE-Base finetuned for only 25 epochs surpasses MAE-Base fined-tuned for 100 epochs by 2.9% box AP and 2.2% mask AP respectively. Code and pretrained models are available at https://github.com/Alpha-VL/ConvMAE.

翻译：视觉变异器( VIT) 被广泛采用, 成为各种视觉任务的建筑。蒙面自动编码, 用于特别训练前和多规模混合革命- 变异的复合结构, 可以进一步释放VIT的潜力, 导致图像分类、检测和语义分割方面的最先进的表现。在本文中, 我们的ConvMAE框架表明, 多规模混合革命- 变异器可以通过蒙面自动编码计划学习更具有歧视性的表现形式。但是, 直接使用原始遮罩战略, 导致计算成本过高和预训练- 调整差异。为了解决这个问题, 我们采用蒙面变异式变异器, 以防止在变异区出现信息泄漏。提议了一个简单的整块化掩码战略, 以确保计算效率。我们还提议更直接监督编码前的多尺度特性, 以我们经过预先培训的 ConvMAE 模型为基础, ConvMAE- Basebase改进了图像Net-1K 精确度, 比MAE- base- base- basy- basyal- basyal- basional- pretrament on.

0

相关内容

何恺明最新论文！用于计算机视觉的可扩展自监督学习方案Masked AutoEncoders

何恺明最新论文！用于计算机视觉的可扩展自监督学习方案Masked AutoEncoders

专知会员服务

30+阅读 · 2021年11月13日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

【伯克利】自回归模型的局部掩卷积，Locally Masked Convolution for Autoregressive Models

【伯克利】自回归模型的局部掩卷积，Locally Masked Convolution for Autoregressive Models

专知会员服务

20+阅读 · 2020年6月23日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

专知

14+阅读 · 2018年3月30日

MoCoGAN 分解运动和内容的视频生成

MoCoGAN 分解运动和内容的视频生成

CreateAMind

18+阅读 · 2017年10月21日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

间充质干细胞调控Treg/Th17平衡诱导肝移植免疫耐受的作用机制

国家自然科学基金

0+阅读 · 2014年12月31日

新型法氏囊活性肽调节免疫反应和B细胞分化的作用机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

病毒感染中肺局部的B细胞免疫反应及其调控机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

以脂滴结构蛋白为核心- - 白藜芦醇对泡沫细胞形成的影响和调控机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

柴油机排气颗粒物与氮氧化物氧化反应的基础研究

国家自然科学基金

0+阅读 · 2012年12月31日

缺陷和非磁性元素掺杂诱导的氮化镓纳米结构磁性机理及调控研究

国家自然科学基金

0+阅读 · 2011年12月31日

DNA甲基化介导的CLDN6表达沉默机制及其对人乳腺癌细胞转移表型的影响

国家自然科学基金

0+阅读 · 2011年12月31日

融合地形结构的DEM质量分析研究

国家自然科学基金

0+阅读 · 2009年12月31日

新BRCA1剪接异构体在乳腺癌细胞中的功能研究

国家自然科学基金

0+阅读 · 2008年12月31日

新型诊断和治疗性AuNP-Cy5-ASODN分子探针体系的建立及其对恶性肿瘤分子成像与治疗评价的研究

国家自然科学基金

0+阅读 · 2008年12月31日

LViT: Language meets Vision Transformer in Medical Image Segmentation

Arxiv

0+阅读 · 2022年6月29日

vMFNet: Compositionality Meets Domain-generalised Segmentation

vMFNet: Compositionality Meets Domain-generalised Segmentation

Arxiv

0+阅读 · 2022年6月29日

Learning Gait Representation from Massive Unlabelled Walking Videos: A Benchmark

Arxiv

0+阅读 · 2022年6月28日

Training Your Sparse Neural Network Better with Any Mask

Arxiv

0+阅读 · 2022年6月28日

DeepFry: Identifying Vocal Fry Using Deep Neural Networks

Arxiv

0+阅读 · 2022年6月26日

Voxel-MAE: Masked Autoencoders for Pre-training Large-scale Point Clouds

Arxiv

0+阅读 · 2022年6月24日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

Arxiv

15+阅读 · 2020年2月28日

End-to-End Dense Video Captioning with Masked Transformer

Arxiv

14+阅读 · 2018年4月3日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

VIP会员

文章信息

相关主题

相关VIP内容

何恺明最新论文！用于计算机视觉的可扩展自监督学习方案Masked AutoEncoders

何恺明最新论文！用于计算机视觉的可扩展自监督学习方案Masked AutoEncoders

专知会员服务

30+阅读 · 2021年11月13日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

【伯克利】自回归模型的局部掩卷积，Locally Masked Convolution for Autoregressive Models

【伯克利】自回归模型的局部掩卷积，Locally Masked Convolution for Autoregressive Models

专知会员服务

20+阅读 · 2020年6月23日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

《物联网（IoT）中的无人机通信高效控制》135页

《在GNSS信号降级环境中利用共识实现无人机集群稳健协调》

中程单向攻击无人机的战略意义：俄乌战争启示

《面向无人机集群的避障动态传感器覆盖算法》最新38页

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

【论文推荐】最新八篇网络节点表示相关论文—可扩展嵌入、对抗自编码器、图划分、异构信息、显式矩阵分解、深度高斯、图、随机游走

专知

14+阅读 · 2018年3月30日

MoCoGAN 分解运动和内容的视频生成

MoCoGAN 分解运动和内容的视频生成

CreateAMind

18+阅读 · 2017年10月21日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

相关论文

LViT: Language meets Vision Transformer in Medical Image Segmentation

Arxiv

0+阅读 · 2022年6月29日

vMFNet: Compositionality Meets Domain-generalised Segmentation

vMFNet: Compositionality Meets Domain-generalised Segmentation

Arxiv

0+阅读 · 2022年6月29日

Learning Gait Representation from Massive Unlabelled Walking Videos: A Benchmark

Arxiv

0+阅读 · 2022年6月28日

Training Your Sparse Neural Network Better with Any Mask

Arxiv

0+阅读 · 2022年6月28日

DeepFry: Identifying Vocal Fry Using Deep Neural Networks

Arxiv

0+阅读 · 2022年6月26日

Voxel-MAE: Masked Autoencoders for Pre-training Large-scale Point Clouds

Arxiv

0+阅读 · 2022年6月24日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

Arxiv

15+阅读 · 2020年2月28日

End-to-End Dense Video Captioning with Masked Transformer

Arxiv

14+阅读 · 2018年4月3日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

相关基金

间充质干细胞调控Treg/Th17平衡诱导肝移植免疫耐受的作用机制

国家自然科学基金

0+阅读 · 2014年12月31日

新型法氏囊活性肽调节免疫反应和B细胞分化的作用机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

病毒感染中肺局部的B细胞免疫反应及其调控机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

以脂滴结构蛋白为核心- - 白藜芦醇对泡沫细胞形成的影响和调控机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

柴油机排气颗粒物与氮氧化物氧化反应的基础研究

国家自然科学基金

0+阅读 · 2012年12月31日

缺陷和非磁性元素掺杂诱导的氮化镓纳米结构磁性机理及调控研究

国家自然科学基金

0+阅读 · 2011年12月31日

DNA甲基化介导的CLDN6表达沉默机制及其对人乳腺癌细胞转移表型的影响

国家自然科学基金

0+阅读 · 2011年12月31日

融合地形结构的DEM质量分析研究

国家自然科学基金

0+阅读 · 2009年12月31日

新BRCA1剪接异构体在乳腺癌细胞中的功能研究

国家自然科学基金

0+阅读 · 2008年12月31日

新型诊断和治疗性AuNP-Cy5-ASODN分子探针体系的建立及其对恶性肿瘤分子成像与治疗评价的研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员