H- Transferent- 1D: 序列快速一维级级注意 (H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences)

We describe an efficient hierarchical method to compute attention in the Transformer architecture. The proposed attention mechanism exploits a matrix structure similar to the Hierarchical Matrix (H-Matrix) developed by the numerical analysis community, and has linear run time and memory complexity. We perform extensive experiments to show that the inductive bias embodied by our hierarchical attention is effective in capturing the hierarchical structure in the sequences typical for natural language and vision tasks. Our method is superior to alternative sub-quadratic proposals by over +6 points on average on the Long Range Arena benchmark. It also sets a new SOTA test perplexity on One-Billion Word dataset with 5x fewer model parameters than that of the previous-best Transformer-based models.

翻译：我们描述了一种高效的等级方法,用以计算变异器结构中的注意程度。提议的注意机制利用了一个与数字分析界开发的等级矩阵(H-Matrix)相似的矩阵结构,并具有线性运行时间和记忆复杂性。我们进行了广泛的实验,以表明我们分级关注所表现的感应偏差能够有效地捕捉自然语言和视觉任务典型序列中的等级结构。我们的方法优于替代的次水道结构,在长距离阿雷纳基准中平均超过+6点。它还为一亿字数据集设定了新的SOTA测试,其模型参数比以前最佳变异器模型少5x倍。

相关内容

注意力机制

关注 120

Attention机制最早是在视觉图像领域提出来的，但是真正火起来应该算是google mind团队的这篇论文《Recurrent Models of Visual Attention》[14]，他们在RNN模型上使用了attention机制来进行图像分类。随后，Bahdanau等人在论文《Neural Machine Translation by Jointly Learning to Align and Translate》 [1]中，使用类似attention的机制在机器翻译任务上将翻译和对齐同时进行，他们的工作算是是第一个提出attention机制应用到NLP领域中。接着类似的基于attention机制的RNN模型扩展开始应用到各种NLP任务中。最近，如何在CNN中使用attention机制也成为了大家的研究热点。下图表示了attention研究进展的大概趋势。

【ICML2021】PoolingFormer：具有池化注意力机制的长序列输入模型

专知会员服务

35+阅读 · 2021年7月25日

最新《Transformers模型》教程，64页ppt

专知会员服务

320+阅读 · 2020年11月26日

替换Transformer！谷歌提出 Performer 模型，全面提升注意力机制！

专知会员服务

43+阅读 · 2020年10月29日

【商汤科技】可变形Transformers端到端对象检测，Deformable DETR

专知会员服务

33+阅读 · 2020年10月11日