自我淡化的愿景转换器 (Self-slimmed Vision Transformer) - 专知论文

会员服务 ·

0

词元分析器 · Vision · 变换 · SLIM · Unstructured ·

2022 年 7 月 26 日

Self-slimmed Vision Transformer

翻译：自我淡化的愿景转换器

Zhuofan Zong,Kunchang Li,Guanglu Song,Yali Wang,Yu Qiao,Biao Leng,Yu Liu

from arxiv, Accepted by ECCV 2022. Code is available at https://github.com/Sense-X/SiT

Vision transformers (ViTs) have become the popular structures and outperformed convolutional neural networks (CNNs) on various vision tasks. However, such powerful transformers bring a huge computation burden, because of the exhausting token-to-token comparison. The previous works focus on dropping insignificant tokens to reduce the computational cost of ViTs. But when the dropping ratio increases, this hard manner will inevitably discard the vital tokens, which limits its efficiency. To solve the issue, we propose a generic self-slimmed learning approach for vanilla ViTs, namely SiT. Specifically, we first design a novel Token Slimming Module (TSM), which can boost the inference efficiency of ViTs by dynamic token aggregation. As a general method of token hard dropping, our TSM softly integrates redundant tokens into fewer informative ones. It can dynamically zoom visual attention without cutting off discriminative token relations in the images, even with a high slimming ratio. Furthermore, we introduce a concise Feature Recalibration Distillation (FRD) framework, wherein we design a reverse version of TSM (RTSM) to recalibrate the unstructured token in a flexible auto-encoder manner. Due to the similar structure between teacher and student, our FRD can effectively leverage structure knowledge for better convergence. Finally, we conduct extensive experiments to evaluate our SiT. It demonstrates that our method can speed up ViTs by 1.7x with negligible accuracy drop, and even speed up ViTs by 3.6x while maintaining 97% of their performance. Surprisingly, by simply arming LV-ViT with our SiT, we achieve new state-of-the-art performance on ImageNet. Code is available at https://github.com/Sense-X/SiT.

翻译：视觉变异器(ViTs)已成为流行结构,在各种视觉任务上,其速度速度网络比变异性神经网络(CNNs)要快。然而,如此强大的变异器带来了巨大的计算负担,因为光耗的象征性比对。先前的工作重点是丢弃微不足道的象征物,以降低ViTs的计算成本。但是,当下降比率上升时,这种艰难的方式将不可避免地丢弃重要象征物,这限制了它的效率。为了解决这个问题,我们提议对VanXlavla ViTs(即SiT)采用一种通用的自我滑动的学习方法。具体地说,我们首次设计了一个新型的 Token 攀升模块(TM),它能通过动态的象征性聚合来提高 ViXslimmmus的推断效率。作为一般的方法,我们的TSIM将多余的象征物软化地整合为更少的信息。它可以动态地放大视觉关注,而不会在图像中消除歧视象征物的关系,即使它是一个高的微缩率。此外,我们也可以在 Vireal Redial Redial Dial Dust (FRD) lax) 框架中以更灵活地保持它们的自我稳定化的图像结构,我们可以在SMSMSMSDral Stal Staldestral Stal-st-st-deal Stal Stal Stal Stal Stal) 结构中进行一个反演算。

0

相关内容

词元分析器

词元分析器

20篇「ICCV2021 Oral」最新论文抢先看！看当下计算机视觉在研究什么？

20篇「ICCV2021 Oral」最新论文抢先看！看当下计算机视觉在研究什么？

专知会员服务

62+阅读 · 2021年7月30日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

325+阅读 · 2020年11月26日

【陈天奇】TVM：端到端自动深度学习编译器，244页ppt

【陈天奇】TVM：端到端自动深度学习编译器，244页ppt

专知会员服务

87+阅读 · 2020年5月11日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

ExBert — 可视化分析Transformer学到的表示

ExBert — 可视化分析Transformer学到的表示

专知会员服务

32+阅读 · 2019年10月16日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

ABCB6基因在眼组织缺损中的功能研究

国家自然科学基金

0+阅读 · 2014年12月31日

非ABA依赖型SnRK2激酶调控马铃薯响应干旱胁迫的机制解析

国家自然科学基金

0+阅读 · 2014年12月31日

蓝绿菌诱导碳酸盐矿物形成机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

硅量子点纳米结构薄膜材料及其太阳电池的制备研究

国家自然科学基金

0+阅读 · 2012年12月31日

铁电材料对FePt薄膜垂直磁各向异性调控机理的研究

国家自然科学基金

0+阅读 · 2012年12月31日

全基因组水平植物信使RNA选择性多聚腺苷化调控研究

国家自然科学基金

0+阅读 · 2011年12月31日

利用斑马鱼模型研究多环芳烃的视觉发育毒性和机制

国家自然科学基金

0+阅读 · 2011年12月31日

HGF诱导NSCLC细胞对EGFR-TKIs耐药机制的研究。

国家自然科学基金

0+阅读 · 2011年12月31日

全基因组甲基化CpG岛扩增技术的建立及在食管癌早期诊断中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

基于表面增强共振拉曼光谱技术的脱氧核酶构效关系研究

国家自然科学基金

0+阅读 · 2008年12月31日

NIERT: Accurate Numerical Interpolation through Unifying Scattered Data Representations using Transformer Encoder

Arxiv

0+阅读 · 2022年9月19日

Fine-Grained Semantically Aligned Vision-Language Pre-Training

Arxiv

0+阅读 · 2022年9月19日

Panoramic Vision Transformer for Saliency Detection in 360° Videos

Arxiv

0+阅读 · 2022年9月19日

SurroundDepth: Entangling Surrounding Views for Self-Supervised Multi-Camera Depth Estimation

Arxiv

0+阅读 · 2022年9月19日

Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with Knowledge

Arxiv

0+阅读 · 2022年9月15日

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

Arxiv

0+阅读 · 2022年9月15日

Graph Self-Supervised Learning: A Survey

Arxiv

15+阅读 · 2021年8月5日

SiT: Self-supervised vIsion Transformer

Arxiv

19+阅读 · 2021年4月8日

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Arxiv

11+阅读 · 2020年7月31日

Text Generation from Knowledge Graphs with Graph Transformers

Arxiv

35+阅读 · 2019年4月4日

VIP会员

文章信息

相关主题

词元分析器

相关VIP内容

20篇「ICCV2021 Oral」最新论文抢先看！看当下计算机视觉在研究什么？

20篇「ICCV2021 Oral」最新论文抢先看！看当下计算机视觉在研究什么？

专知会员服务

62+阅读 · 2021年7月30日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

325+阅读 · 2020年11月26日

【陈天奇】TVM：端到端自动深度学习编译器，244页ppt

【陈天奇】TVM：端到端自动深度学习编译器，244页ppt

专知会员服务

87+阅读 · 2020年5月11日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

ExBert — 可视化分析Transformer学到的表示

ExBert — 可视化分析Transformer学到的表示

专知会员服务

32+阅读 · 2019年10月16日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《为多域数字战场变革装甲力量》报告

《多域训练：利用开放标准将太空与网络域同陆、海、空域训练相整合》报告

面向城市战：欧美徒步作战新装备

《人工智能增强监视分析：利用跨网络、陆地、空中及海上领域的威胁向量实时建模》

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

相关论文

NIERT: Accurate Numerical Interpolation through Unifying Scattered Data Representations using Transformer Encoder

Arxiv

0+阅读 · 2022年9月19日

Fine-Grained Semantically Aligned Vision-Language Pre-Training

Arxiv

0+阅读 · 2022年9月19日

Panoramic Vision Transformer for Saliency Detection in 360° Videos

Arxiv

0+阅读 · 2022年9月19日

SurroundDepth: Entangling Surrounding Views for Self-Supervised Multi-Camera Depth Estimation

Arxiv

0+阅读 · 2022年9月19日

Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with Knowledge

Arxiv

0+阅读 · 2022年9月15日

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

Arxiv

0+阅读 · 2022年9月15日

Graph Self-Supervised Learning: A Survey

Arxiv

15+阅读 · 2021年8月5日

SiT: Self-supervised vIsion Transformer

Arxiv

19+阅读 · 2021年4月8日

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Arxiv

11+阅读 · 2020年7月31日

Text Generation from Knowledge Graphs with Graph Transformers

Arxiv

35+阅读 · 2019年4月4日

相关基金

ABCB6基因在眼组织缺损中的功能研究

国家自然科学基金

0+阅读 · 2014年12月31日

非ABA依赖型SnRK2激酶调控马铃薯响应干旱胁迫的机制解析

国家自然科学基金

0+阅读 · 2014年12月31日

蓝绿菌诱导碳酸盐矿物形成机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

硅量子点纳米结构薄膜材料及其太阳电池的制备研究

国家自然科学基金

0+阅读 · 2012年12月31日

铁电材料对FePt薄膜垂直磁各向异性调控机理的研究

国家自然科学基金

0+阅读 · 2012年12月31日

全基因组水平植物信使RNA选择性多聚腺苷化调控研究

国家自然科学基金

0+阅读 · 2011年12月31日

利用斑马鱼模型研究多环芳烃的视觉发育毒性和机制

国家自然科学基金

0+阅读 · 2011年12月31日

HGF诱导NSCLC细胞对EGFR-TKIs耐药机制的研究。

国家自然科学基金

0+阅读 · 2011年12月31日

全基因组甲基化CpG岛扩增技术的建立及在食管癌早期诊断中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

基于表面增强共振拉曼光谱技术的脱氧核酶构效关系研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员