优化GPU上动态平行主义的汇编器框架 (A Compiler Framework for Optimizing Dynamic Parallelism on GPUs) - 专知论文

会员服务 ·

0

编译器 · Performer · 优化器 · 块 · 阈值 ·

2022 年 1 月 8 日

A Compiler Framework for Optimizing Dynamic Parallelism on GPUs

翻译：优化GPU上动态平行主义的汇编器框架

Mhd Ghaith Olabi,Juan Gómez Luna,Onur Mutlu,Wen-mei Hwu,Izzat El Hajj

Dynamic parallelism on GPUs allows GPU threads to dynamically launch other GPU threads. It is useful in applications with nested parallelism, particularly where the amount of nested parallelism is irregular and cannot be predicted beforehand. However, prior works have shown that dynamic parallelism may impose a high performance penalty when a large number of small grids are launched. The large number of launches results in high launch latency due to congestion, and the small grid sizes result in hardware underutilization. To address this issue, we propose a compiler framework for optimizing the use of dynamic parallelism in applications with nested parallelism. The framework features three key optimizations: thresholding, coarsening, and aggregation. Thresholding involves launching a grid dynamically only if the number of child threads exceeds some threshold, and serializing the child threads in the parent thread otherwise. Coarsening involves executing the work of multiple thread blocks by a single coarsened block to amortize the common work across them. Aggregation involves combining multiple child grids into a single aggregated grid. Our evaluation shows that our compiler framework improves the performance of applications with nested parallelism by a geometric mean of 43.0x over applications that use dynamic parallelism, 8.7x over applications that do not use dynamic parallelism, and 3.6x over applications that use dynamic parallelism with aggregation alone as proposed in prior work.

翻译：在 GPU 上的动态平行关系使 GPU 线索能够动态地启动其他 GPU 线索。它在嵌入平行关系的应用中非常有用, 特别是在嵌入平行关系的数量不固定且无法事先预测的情况下。但是, 先前的工作表明, 动态平行关系可能会在大量小网格启动时施加很高的性能处罚。大量发射导致由于拥堵导致的发射延迟, 以及小网格大小导致硬件利用不足。为了解决这个问题, 我们提议了一个编译框架, 优化在与嵌入平行关系的应用应用中使用动态平行的动态平行关系。我们的评估显示, 我们的编译框架本身改进了动态平行关系应用的性能, 而不是以动态关系平行关系平行关系, 而不是以动态关系平行关系的方式, 而不是以动态关系平行关系进行。

0

相关内容

编译器

编译器（Compiler），是一种计算机程序，它会将用某种编程语言写成的源代码（原始语言），转换成另一种编程语言（目标语言）。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

实践教程 | 如何设置CUDA Kernel中的grid_size和block_size？

实践教程 | 如何设置CUDA Kernel中的grid_size和block_size？

极市平台

0+阅读 · 2022年1月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

基于GPU的CSAMT三维正演的并行外推多网格法研究

国家自然科学基金

0+阅读 · 2014年12月31日

体系结构级GPU功耗建模及软件低功耗优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

物流配送云资源虚拟化管理与服务组合优化研究

国家自然科学基金

0+阅读 · 2013年12月31日

Vlasov-Poisson-Boltzmann方程研究

国家自然科学基金

0+阅读 · 2013年12月31日

功耗自适应视频编码与多核处理器架构优化研究

国家自然科学基金

0+阅读 · 2013年12月31日

面向GPU的电力系统电磁暂态并行计算方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

CUDA、OpenMP和MPI混合加速的隐式粒子模拟算法与框架研究

国家自然科学基金

1+阅读 · 2012年12月31日

基于几何代数的多维统一空间关系计算模型及并行化方法

国家自然科学基金

1+阅读 · 2011年12月31日

基于NURBS曲面的弹跳射线法的GPU加速

国家自然科学基金

0+阅读 · 2008年12月31日

图的正则性和胞腔代数

国家自然科学基金

0+阅读 · 2008年12月31日

Theoretical analysis of edit distance algorithms: an applied perspective

Arxiv

0+阅读 · 2022年4月20日

Adaptive Non-linear Filtering Technique for Image Restoration

Arxiv

1+阅读 · 2022年4月20日

Scalable Motif Counting for Large-scale Temporal Graphs

Arxiv

0+阅读 · 2022年4月20日

SnapFuzz: An Efficient Fuzzing Framework for Network Applications

Arxiv

0+阅读 · 2022年4月19日

CPU- and GPU-based Distributed Sampling in Dirichlet Process Mixtures for Large-scale Analysis

CPU- and GPU-based Distributed Sampling in Dirichlet Process Mixtures for Large-scale Analysis

Arxiv

0+阅读 · 2022年4月19日

Suffix tree-based linear algorithms for multiple prefixes, single suffix counting and listing problems

Suffix tree-based linear algorithms for multiple prefixes, single suffix counting and listing problems

Arxiv

0+阅读 · 2022年4月18日

Dynamic Approximate Maximum Independent Set on Massive Graphs

Arxiv

0+阅读 · 2022年4月18日

Twin-width can be exponential in treewidth

Arxiv

0+阅读 · 2022年4月15日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

Arxiv

16+阅读 · 2021年5月2日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《使用量化测量将传感器节点关联到融合中心的算法设计》171页

军事前沿模型

提升军事训练能力的最佳人工智能模拟工具

《社交媒体信息作战》最新48页技术报告

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

实践教程 | 如何设置CUDA Kernel中的grid_size和block_size？

实践教程 | 如何设置CUDA Kernel中的grid_size和block_size？

极市平台

0+阅读 · 2022年1月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

相关论文

Theoretical analysis of edit distance algorithms: an applied perspective

Arxiv

0+阅读 · 2022年4月20日

Adaptive Non-linear Filtering Technique for Image Restoration

Arxiv

1+阅读 · 2022年4月20日

Scalable Motif Counting for Large-scale Temporal Graphs

Arxiv

0+阅读 · 2022年4月20日

SnapFuzz: An Efficient Fuzzing Framework for Network Applications

Arxiv

0+阅读 · 2022年4月19日

CPU- and GPU-based Distributed Sampling in Dirichlet Process Mixtures for Large-scale Analysis

CPU- and GPU-based Distributed Sampling in Dirichlet Process Mixtures for Large-scale Analysis

Arxiv

0+阅读 · 2022年4月19日

Suffix tree-based linear algorithms for multiple prefixes, single suffix counting and listing problems

Suffix tree-based linear algorithms for multiple prefixes, single suffix counting and listing problems

Arxiv

0+阅读 · 2022年4月18日

Dynamic Approximate Maximum Independent Set on Massive Graphs

Arxiv

0+阅读 · 2022年4月18日

Twin-width can be exponential in treewidth

Arxiv

0+阅读 · 2022年4月15日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

Arxiv

16+阅读 · 2021年5月2日

相关基金

基于GPU的CSAMT三维正演的并行外推多网格法研究

国家自然科学基金

0+阅读 · 2014年12月31日

体系结构级GPU功耗建模及软件低功耗优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

物流配送云资源虚拟化管理与服务组合优化研究

国家自然科学基金

0+阅读 · 2013年12月31日

Vlasov-Poisson-Boltzmann方程研究

国家自然科学基金

0+阅读 · 2013年12月31日

功耗自适应视频编码与多核处理器架构优化研究

国家自然科学基金

0+阅读 · 2013年12月31日

面向GPU的电力系统电磁暂态并行计算方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

CUDA、OpenMP和MPI混合加速的隐式粒子模拟算法与框架研究

国家自然科学基金

1+阅读 · 2012年12月31日

基于几何代数的多维统一空间关系计算模型及并行化方法

国家自然科学基金

1+阅读 · 2011年12月31日

基于NURBS曲面的弹跳射线法的GPU加速

国家自然科学基金

0+阅读 · 2008年12月31日

图的正则性和胞腔代数

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员