动力变形器:弥合自我关注与其线性化之间的绩效差距 (Momentum Transformer: Closing the Performance Gap Between Self-attention and Its Linearization) - 专知论文

会员服务 ·

0

动量 · 线性的 · Attention · Performer · 变换 ·

2022 年 8 月 1 日

Momentum Transformer: Closing the Performance Gap Between Self-attention and Its Linearization

翻译：动力变形器:弥合自我关注与其线性化之间的绩效差距

Tan Nguyen,Richard G. Baraniuk,Robert M. Kirby,Stanley J. Osher,Bao Wang

from arxiv, 22 pages, 5 figures. arXiv admin note: substantial text overlap with arXiv:2110.07034

Transformers have achieved remarkable success in sequence modeling and beyond but suffer from quadratic computational and memory complexities with respect to the length of the input sequence. Leveraging techniques include sparse and linear attention and hashing tricks; efficient transformers have been proposed to reduce the quadratic complexity of transformers but significantly degrade the accuracy. In response, we first interpret the linear attention and residual connections in computing the attention map as gradient descent steps. We then introduce momentum into these components and propose the \emph{momentum transformer}, which utilizes momentum to improve the accuracy of linear transformers while maintaining linear memory and computational complexities. Furthermore, we develop an adaptive strategy to compute the momentum value for our model based on the optimal momentum for quadratic optimization. This adaptive momentum eliminates the need to search for the optimal momentum value and further enhances the performance of the momentum transformer. A range of experiments on both autoregressive and non-autoregressive tasks, including image generation and machine translation, demonstrate that the momentum transformer outperforms popular linear transformers in training efficiency and accuracy.

翻译：作为回应,我们首先将计算关注图中的线性关注和剩余连接解读为梯度下降步骤。然后,我们对这些组件引入动力,并提议“emph{momentum变压器 ”,它利用动力提高线性变压器的准确性,同时保持线性内存和计算复杂性。此外,我们制定了适应性战略,根据四面优化的最佳势头计算模型的动力值。这种适应性动力消除了寻找最佳动力值和进一步提高动力变压器性能的需要。关于自动递增和非倾斜性任务的一系列实验,包括图像生成和机器翻译,表明动力变压器在培训效率和准确性方面超越了流行的线性线性变压器。

0

相关内容

动量方法 (Polyak, 1964) 旨在加速学习，特别是处理高曲率、小但一致的梯度，或是带噪声的梯度。动量算法积累了之前梯度指数级衰减的移动平均，并且继续沿该方向移动。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【ICLR2020论文】自我注意力与卷积层的关系，On the Relationship between Self-Attention and Convolutional Layers

【ICLR2020论文】自我注意力与卷积层的关系，On the Relationship between Self-Attention and Convolutional Layers

专知会员服务

37+阅读 · 2020年1月12日

【Python Tricks新书】The book: A Buffet of Awesome Python Features，299页pdf

【Python Tricks新书】The book: A Buffet of Awesome Python Features，299页pdf

专知会员服务

45+阅读 · 2020年1月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Multi-Task Learning的几篇综述文章

Multi-Task Learning的几篇综述文章

深度学习自然语言处理

15+阅读 · 2020年6月15日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

生物自发光触发高荧光碳量子点基复合光敏剂的合成及其光动力治疗性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

压电式六维力/力矩传感器静动态特性建模及主动优化设计

国家自然科学基金

0+阅读 · 2014年12月31日

支撑二元过渡金属团簇磁各向异性能的调控研究

国家自然科学基金

0+阅读 · 2014年12月31日

大场景集成成像三维显示系统建模与光场转换研究

国家自然科学基金

0+阅读 · 2013年12月31日

珠江口甲烷生物氧化途径与环境特征关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

类石墨烯表面结构和性能预测

国家自然科学基金

0+阅读 · 2011年12月31日

金属/稀土复合胶体纳米发光材料的光学性质研究

国家自然科学基金

0+阅读 · 2011年12月31日

X射线相衬成像技术研究柴油机喷嘴近场初级雾化机理

国家自然科学基金

0+阅读 · 2009年12月31日

生物可降解性多模态纳米微粒构建与TIMP-2、Endostatin联合靶向转运抑制动脉粥样硬化易损斑块血管发生的研究

国家自然科学基金

0+阅读 · 2009年12月31日

壳聚糖-聚乳酸接枝共聚物制备生物可吸收水凝胶药物缓释体系

国家自然科学基金

0+阅读 · 2008年12月31日

Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning

Arxiv

0+阅读 · 2022年9月27日

Improving Image Clustering through Sample Ranking and Its Application to remote--sensing images

Arxiv

0+阅读 · 2022年9月26日

Mega: Moving Average Equipped Gated Attention

Arxiv

0+阅读 · 2022年9月26日

A Deep Investigation of RNN and Self-attention for the Cyrillic-Traditional Mongolian Bidirectional Conversion

Arxiv

0+阅读 · 2022年9月24日

FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks

Arxiv

0+阅读 · 2022年9月24日

A Survey on Vision Transformer

Arxiv

17+阅读 · 2022年2月23日

A Survey on Visual Transformer

Arxiv

19+阅读 · 2020年12月23日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

12+阅读 · 2020年6月23日

Learning in the Frequency Domain

Learning in the Frequency Domain

Arxiv

11+阅读 · 2020年3月12日

Learning to Propagate for Graph Meta-Learning

Arxiv

14+阅读 · 2019年9月11日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【ICLR2020论文】自我注意力与卷积层的关系，On the Relationship between Self-Attention and Convolutional Layers

【ICLR2020论文】自我注意力与卷积层的关系，On the Relationship between Self-Attention and Convolutional Layers

专知会员服务

37+阅读 · 2020年1月12日

【Python Tricks新书】The book: A Buffet of Awesome Python Features，299页pdf

【Python Tricks新书】The book: A Buffet of Awesome Python Features，299页pdf

专知会员服务

45+阅读 · 2020年1月1日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】数据驱动决策中的激励、信息与不确定性

DGP双粒度提示框架：图增强大模型助力欺诈检测

【ICCV2025】ESSENTIAL：用于视频类增量学习的情景记忆与语义记忆整合

唯快不破：大型语言模型高效架构综述

相关资讯

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Multi-Task Learning的几篇综述文章

Multi-Task Learning的几篇综述文章

深度学习自然语言处理

15+阅读 · 2020年6月15日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

相关论文

Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning

Arxiv

0+阅读 · 2022年9月27日

Improving Image Clustering through Sample Ranking and Its Application to remote--sensing images

Arxiv

0+阅读 · 2022年9月26日

Mega: Moving Average Equipped Gated Attention

Arxiv

0+阅读 · 2022年9月26日

A Deep Investigation of RNN and Self-attention for the Cyrillic-Traditional Mongolian Bidirectional Conversion

Arxiv

0+阅读 · 2022年9月24日

FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks

Arxiv

0+阅读 · 2022年9月24日

A Survey on Vision Transformer

Arxiv

17+阅读 · 2022年2月23日

A Survey on Visual Transformer

Arxiv

19+阅读 · 2020年12月23日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

12+阅读 · 2020年6月23日

Learning in the Frequency Domain

Learning in the Frequency Domain

Arxiv

11+阅读 · 2020年3月12日

Learning to Propagate for Graph Meta-Learning

Arxiv

14+阅读 · 2019年9月11日

相关基金

生物自发光触发高荧光碳量子点基复合光敏剂的合成及其光动力治疗性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

压电式六维力/力矩传感器静动态特性建模及主动优化设计

国家自然科学基金

0+阅读 · 2014年12月31日

支撑二元过渡金属团簇磁各向异性能的调控研究

国家自然科学基金

0+阅读 · 2014年12月31日

大场景集成成像三维显示系统建模与光场转换研究

国家自然科学基金

0+阅读 · 2013年12月31日

珠江口甲烷生物氧化途径与环境特征关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

类石墨烯表面结构和性能预测

国家自然科学基金

0+阅读 · 2011年12月31日

金属/稀土复合胶体纳米发光材料的光学性质研究

国家自然科学基金

0+阅读 · 2011年12月31日

X射线相衬成像技术研究柴油机喷嘴近场初级雾化机理

国家自然科学基金

0+阅读 · 2009年12月31日

生物可降解性多模态纳米微粒构建与TIMP-2、Endostatin联合靶向转运抑制动脉粥样硬化易损斑块血管发生的研究

国家自然科学基金

0+阅读 · 2009年12月31日

壳聚糖-聚乳酸接枝共聚物制备生物可吸收水凝胶药物缓释体系

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员