VSA: 在愿景变换器中学习不同程度的窗口关注 (VSA: Learning Varied-Size Window Attention in Vision Transformers) - 专知论文

会员服务 ·

0

Microsoft Windows · 注意力机制 · Performer · Vision · 变换 ·

2022 年 4 月 18 日

VSA: Learning Varied-Size Window Attention in Vision Transformers

翻译：VSA: 在愿景变换器中学习不同程度的窗口关注

Qiming Zhang,Yufei Xu,Jing Zhang,Dacheng Tao

from arxiv, 23 pages, 13 tables, and 5 figures

Attention within windows has been widely explored in vision transformers to balance the performance, computation complexity, and memory footprint. However, current models adopt a hand-crafted fixed-size window design, which restricts their capacity of modeling long-term dependencies and adapting to objects of different sizes. To address this drawback, we propose \textbf{V}aried-\textbf{S}ize Window \textbf{A}ttention (VSA) to learn adaptive window configurations from data. Specifically, based on the tokens within each default window, VSA employs a window regression module to predict the size and location of the target window, i.e., the attention area where the key and value tokens are sampled. By adopting VSA independently for each attention head, it can model long-term dependencies, capture rich context from diverse windows, and promote information exchange among overlapped windows. VSA is an easy-to-implement module that can replace the window attention in state-of-the-art representative models with minor modifications and negligible extra computational cost while improving their performance by a large margin, e.g., 1.1\% for Swin-T on ImageNet classification. In addition, the performance gain increases when using larger images for training and test. Experimental results on more downstream tasks, including object detection, instance segmentation, and semantic segmentation, further demonstrate the superiority of VSA over the vanilla window attention in dealing with objects of different sizes. The code will be released https://github.com/ViTAE-Transformer/ViTAE-VSA.

翻译：在视觉变压器中广泛探索了窗口内部的注意,以平衡性能、计算复杂性和记忆足迹。然而,当前模型采用手工制作的固定规模窗口设计,限制了其模拟长期依赖性和适应不同大小对象的能力。为解决这一退步,我们提议采用以下方法:Textbf{V}aried-textbf{S} 将窗口的上下文从数据中学习适应性窗口配置。具体地说,基于每个默认窗口的标志, VSA使用一个窗口回归模块来预测目标窗口的大小和位置,即关键和值符号抽样的注意区域。通过对每个关注头独立采用 VSA 来模拟长期依赖性,从不同的窗口中捕捉丰富的背景,促进重叠窗口之间的信息交流。 VSA是一个容易执行的模块,可以取代州级代表模型中的窗口关注度,有轻微的修改和微不足道的超值计算成本,同时通过大边距的S-Vial-dealalal-dealal 测试性能提升其性能,在更大范围内的S-delive-deal-deal-dealal realal exalal exal exalation lagistraudeal 上,在图像中将显示上,在更大性变压值/e-tailal-tailal-traal-tailal-tailal-tailal-tamental-tailal-tamentaltractionalisalisalisalisaltractionxal delisalisalisalxxaltraaltraalxxxxxxalation delvialation 。

0

相关内容

Microsoft Windows

Microsoft Windows

Microsoft Windows（视窗操作系统）是微软公司推出的一系列操作系统。它问世于1985年，当时是DOS之下的操作环境，而后其后续版本作逐渐发展成为个人电脑和服务器用户设计的操作系统。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

321+阅读 · 2020年11月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

专知

19+阅读 · 2018年5月31日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

新型表面增强拉曼-荧光双编码微球的制备及其对DNA的检测应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

概率和平均框架下一系列Sobolev空间中的函数逼近与恢复

国家自然科学基金

1+阅读 · 2015年12月31日

多控磁性核-壳微球Fe3O4@MOFs/GO的构筑及载药性能的研究

国家自然科学基金

0+阅读 · 2013年12月31日

荒漠化矿区土壤湿度多分辨率时空演变机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

不同尺度海洋溢油风险分区方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于稀疏优化的空时分布密集多径信号估计方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

X射线真彩色CT图像重建研究

国家自然科学基金

0+阅读 · 2013年12月31日

裂纹在沿晶氧化膜内形核的应力腐蚀新机理

国家自然科学基金

0+阅读 · 2011年12月31日

南海深海沉积物来源真菌活性次级代谢产物研究

国家自然科学基金

0+阅读 · 2009年12月31日

不同类型强心苷抗肿瘤活性的研究

国家自然科学基金

0+阅读 · 2009年12月31日

Separable Self-attention for Mobile Vision Transformers

Arxiv

1+阅读 · 2022年6月6日

Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning

Arxiv

0+阅读 · 2022年6月6日

NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

Arxiv

1+阅读 · 2022年6月6日

Anomaly detection in surveillance videos using transformer based attention model

Anomaly detection in surveillance videos using transformer based attention model

Arxiv

0+阅读 · 2022年6月3日

A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing Images

Arxiv

0+阅读 · 2022年6月3日

A Survey on Vision Transformer

Arxiv

17+阅读 · 2022年2月23日

TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classication

Arxiv

17+阅读 · 2021年6月2日

SiT: Self-supervised vIsion Transformer

Arxiv

19+阅读 · 2021年4月8日

Transformer Tracking

Arxiv

17+阅读 · 2021年3月29日

Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling

Arxiv

16+阅读 · 2018年1月31日

VIP会员

文章信息

相关主题

Microsoft Windows

注意力机制

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

最新《Transformers模型》教程，64页ppt

最新《Transformers模型》教程，64页ppt

专知会员服务

321+阅读 · 2020年11月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

专知

19+阅读 · 2018年5月31日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

相关论文

Separable Self-attention for Mobile Vision Transformers

Arxiv

1+阅读 · 2022年6月6日

Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning

Arxiv

0+阅读 · 2022年6月6日

NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

Arxiv

1+阅读 · 2022年6月6日

Anomaly detection in surveillance videos using transformer based attention model

Anomaly detection in surveillance videos using transformer based attention model

Arxiv

0+阅读 · 2022年6月3日

A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing Images

Arxiv

0+阅读 · 2022年6月3日

A Survey on Vision Transformer

Arxiv

17+阅读 · 2022年2月23日

TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classication

Arxiv

17+阅读 · 2021年6月2日

SiT: Self-supervised vIsion Transformer

Arxiv

19+阅读 · 2021年4月8日

Transformer Tracking

Arxiv

17+阅读 · 2021年3月29日

Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling

Arxiv

16+阅读 · 2018年1月31日

相关基金

新型表面增强拉曼-荧光双编码微球的制备及其对DNA的检测应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

概率和平均框架下一系列Sobolev空间中的函数逼近与恢复

国家自然科学基金

1+阅读 · 2015年12月31日

多控磁性核-壳微球Fe3O4@MOFs/GO的构筑及载药性能的研究

国家自然科学基金

0+阅读 · 2013年12月31日

荒漠化矿区土壤湿度多分辨率时空演变机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

不同尺度海洋溢油风险分区方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于稀疏优化的空时分布密集多径信号估计方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

X射线真彩色CT图像重建研究

国家自然科学基金

0+阅读 · 2013年12月31日

裂纹在沿晶氧化膜内形核的应力腐蚀新机理

国家自然科学基金

0+阅读 · 2011年12月31日

南海深海沉积物来源真菌活性次级代谢产物研究

国家自然科学基金

0+阅读 · 2009年12月31日

不同类型强心苷抗肿瘤活性的研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员