FLAT: 用于缓解注意力瓶颈的优化数据流 (FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks) - 专知论文

会员服务 ·

0

Attention · state-of-the-art · 最优化 · Processing（编程语言） · 优化器 ·

2022 年 9 月 24 日

FLAT: An Optimized Dataflow for Mitigating Attention Bottlenecks

翻译：FLAT: 用于缓解注意力瓶颈的优化数据流

Sheng-Chun Kao,Suvinay Subramanian,Gaurav Agrawal,Amir Yazdanbakhsh,Tushar Krishna

Attention mechanisms, primarily designed to capture pairwise correlations between words, have become the backbone of machine learning, expanding beyond natural language processing into other domains. This growth in adaptation comes at the cost of prohibitively large memory requirements and computational complexity, especially at higher number of input elements. This limitation is due to inherently limited data reuse opportunities and quadratic growth in memory footprints, leading to severe memory-boundedness and limited scalability of input elements. This work addresses these challenges by devising a tailored dataflow optimization, called FLAT, for attention mechanisms without altering their functionality. This dataflow processes costly attention operations through a unique fusion mechanism, transforming the memory footprint quadratic growth to merely a linear one. To realize the full potential of this bespoke mechanism, we propose a tiling approach to enhance the data reuse across attention operations. Our method both mitigates the off-chip bandwidth bottleneck as well as reduces the on-chip memory requirement. FLAT delivers 1.94x (1.76x) speedup and 49% and (42%) of energy savings compared to the state-of-the-art Edge (Cloud) accelerators with no customized dataflow optimization. When on-chip resources are scarce (20 KB-200 KB), FLAT yields, on average, 1.5x end-to-end latency reduction across a diverse range of conventional attention-based models with input sequence lengths ranging from 512-token to 64K-token. Our evaluations demonstrate that state-of-the-art DNN dataflow applied to attention operations reach the efficiency limit for inputs above 512 elements. In contrast, FLAT unblocks transformer models for inputs with up to 64K elements

翻译：关注机制,主要是为了捕捉言词之间的对等关系,已经成为机器学习的主干,超越自然语言处理,扩展到其他领域。适应的增加是以惊人的庞大记忆要求和计算复杂性的代价,特别是投入元素数量较多。这一限制是由于数据再利用机会内在有限,记忆足迹四度增长,导致严重的记忆限制和输入元素缩缩放。这项工作通过设计定制的数据流优化(称为FLAT)来应对这些挑战,用于不改变功能的注意机制。这一数据流通过一个独特的聚合机制处理昂贵的注意力操作,将记忆足部二次增长转变为仅仅是线性增长。为了实现这一表达机制的全部潜力,我们建议采取平铺式方法,提高数据在关注操作中的再利用。我们的方法既可以缓解离子带宽带宽的瓶颈,也可以降低在芯片内存储要求。FLAT为1.94x(1.76x)的快速增长和49%的节能节流,而与电量的电量评估(Clooverate State State)相比,将存储足足部的平级增长值增长(KLAxx) 数据流将数据从5Klistrax 递减为数据。

0

相关内容

Attention

【牛津大学博士论文】流形的几何优化与深度学习的应用，154页pdf，Geometric Optimisation on Manifolds with Applications to Deep Learning

【牛津大学博士论文】流形的几何优化与深度学习的应用，154页pdf，Geometric Optimisation on Manifolds with Applications to Deep Learning

专知会员服务

22+阅读 · 2022年3月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

专知会员服务

104+阅读 · 2022年2月10日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

转录激活蛋白YLGat1介导氮饥饿与油脂合成偶联的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

不确定条件下基于分群策略的柔性Flow Shop调度问题研究

国家自然科学基金

0+阅读 · 2013年12月31日

Kronheimer-Nakajima quiver 模空间与有理曲面

国家自然科学基金

1+阅读 · 2013年12月31日

转录因子Ste12调控玉米大斑病菌侵染过程的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

柽柳Dof转录因子的耐盐调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

柑橘绿霉病菌对DMI杀菌剂抗性的调控机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

准周期薛定谔算子中的动力系统理论

国家自然科学基金

0+阅读 · 2012年12月31日

PI-IBS中TMEM16A介导IL-4对Cajal细胞损伤的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

葡萄叶片和果实白藜芦醇合成、转化的相互影响及其酶学和分子机制的研究

国家自然科学基金

0+阅读 · 2008年12月31日

Noise in the Clouds: Influence of Network Performance Variability on Application Scalability

Arxiv

0+阅读 · 2022年11月1日

Compressed Gastric Image Generation Based on Soft-Label Dataset Distillation for Medical Data Sharing

Arxiv

0+阅读 · 2022年11月1日

Empowering Data Centers for Next Generation Trusted Computing

Arxiv

0+阅读 · 2022年11月1日

SOLAR: A Highly Optimized Data Loading Framework for Distributed Training of CNN-based Scientific Surrogates

Arxiv

0+阅读 · 2022年11月1日

AdaVITS: Tiny VITS for Low Computing Resource Speaker Adaptation

Arxiv

0+阅读 · 2022年10月31日

GNN at the Edge: Cost-Efficient Graph Neural Network Processing over Distributed Edge Servers

Arxiv

0+阅读 · 2022年10月31日

Symmetries, flat minima, and the conserved quantities of gradient flow

Arxiv

0+阅读 · 2022年10月31日

CATs++: Boosting Cost Aggregation with Convolutions and Transformers

Arxiv

0+阅读 · 2022年10月30日

Do Pre-trained Models Benefit Equally in Continual Learning?

Arxiv

0+阅读 · 2022年10月27日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

VIP会员

文章信息

相关主题

state-of-the-art

Processing（编程语言）

相关VIP内容

【牛津大学博士论文】流形的几何优化与深度学习的应用，154页pdf，Geometric Optimisation on Manifolds with Applications to Deep Learning

【牛津大学博士论文】流形的几何优化与深度学习的应用，154页pdf，Geometric Optimisation on Manifolds with Applications to Deep Learning

专知会员服务

22+阅读 · 2022年3月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

【干货书】深度学习合成数据，354页pdf，Synthetic Data for Deep Learning

专知会员服务

104+阅读 · 2022年2月10日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《使用量化测量将传感器节点关联到融合中心的算法设计》171页

军事前沿模型

提升军事训练能力的最佳人工智能模拟工具

《社交媒体信息作战》最新48页技术报告

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Noise in the Clouds: Influence of Network Performance Variability on Application Scalability

Arxiv

0+阅读 · 2022年11月1日

Compressed Gastric Image Generation Based on Soft-Label Dataset Distillation for Medical Data Sharing

Arxiv

0+阅读 · 2022年11月1日

Empowering Data Centers for Next Generation Trusted Computing

Arxiv

0+阅读 · 2022年11月1日

SOLAR: A Highly Optimized Data Loading Framework for Distributed Training of CNN-based Scientific Surrogates

Arxiv

0+阅读 · 2022年11月1日

AdaVITS: Tiny VITS for Low Computing Resource Speaker Adaptation

Arxiv

0+阅读 · 2022年10月31日

GNN at the Edge: Cost-Efficient Graph Neural Network Processing over Distributed Edge Servers

Arxiv

0+阅读 · 2022年10月31日

Symmetries, flat minima, and the conserved quantities of gradient flow

Arxiv

0+阅读 · 2022年10月31日

CATs++: Boosting Cost Aggregation with Convolutions and Transformers

Arxiv

0+阅读 · 2022年10月30日

Do Pre-trained Models Benefit Equally in Continual Learning?

Arxiv

0+阅读 · 2022年10月27日

End-to-End Multi-Task Learning with Attention

Arxiv

19+阅读 · 2018年3月28日

相关基金

转录激活蛋白YLGat1介导氮饥饿与油脂合成偶联的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

不确定条件下基于分群策略的柔性Flow Shop调度问题研究

国家自然科学基金

0+阅读 · 2013年12月31日

Kronheimer-Nakajima quiver 模空间与有理曲面

国家自然科学基金

1+阅读 · 2013年12月31日

转录因子Ste12调控玉米大斑病菌侵染过程的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

柽柳Dof转录因子的耐盐调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

柑橘绿霉病菌对DMI杀菌剂抗性的调控机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

准周期薛定谔算子中的动力系统理论

国家自然科学基金

0+阅读 · 2012年12月31日

PI-IBS中TMEM16A介导IL-4对Cajal细胞损伤的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

葡萄叶片和果实白藜芦醇合成、转化的相互影响及其酶学和分子机制的研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员