改进 GPU- Aware 自动同步任务可缩放性 (Improving Scalability with GPU-Aware Asynchronous Tasks) - 专知论文

会员服务 ·

0

Performer · 可约的 · 缩放 · GPU · Integration ·

2022 年 2 月 23 日

Improving Scalability with GPU-Aware Asynchronous Tasks

翻译：改进 GPU- Aware 自动同步任务可缩放性

Jaemin Choi,David F. Richards,Laxmikant V. Kale

from arxiv, 10 pages, 9 figures, submitted to HIPS 2022 workshop

Asynchronous tasks, when created with overdecomposition, enable automatic computation-communication overlap which can substantially improve performance and scalability. This is not only applicable to traditional CPU-based systems, but also to modern GPU-accelerated platforms. While the ability to hide communication behind computation can be highly effective in weak scaling scenarios, performance begins to suffer with smaller problem sizes or in strong scaling due to fine-grained overheads and reduced room for overlap. In this work, we integrate GPU-aware communication into asynchronous tasks in addition to computation-communication overlap, with the goal of reducing time spent in communication and further increasing GPU utilization. We demonstrate the performance impact of our approach using a proxy application that performs the Jacobi iterative method on GPUs, Jacobi3D. In addition to optimizations for minimizing host-device synchronization and increasing the concurrency of GPU operations, we explore techniques such as kernel fusion and CUDA Graphs to combat fine-grained overheads at scale.

翻译：自动计算通信重叠,可以大大改善功能和可缩放性。这不仅适用于传统的基于CPU的系统,而且适用于现代的GPU加速平台。虽然在微弱的缩放假设中,将通信隐藏在计算背后的能力会非常有效,但由于细微缩小的间接费用和减少重叠的空间,业绩开始遇到较小的问题,或大大缩小。在这项工作中,我们除了计算通信重叠之外,还将GPU-觉悟通信纳入非同步任务,目的是减少通信时间,并进一步提高GPU的利用率。我们用一个代理应用程序展示我们方法的绩效影响,该应用程序在GPU(Jacobi3D)上执行Jacobi迭代方法。除了最大限度地减少主控设备同步和增加GPU业务的调值外,我们还探索诸如内核聚和CUDA图等技术,以打击规模的微缩放间接费用。

0

相关内容

Performer

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

Call for Nominations: 2022 Multimedia Prize Paper Award

Call for Nominations: 2022 Multimedia Prize Paper Award

CCF多媒体专委会

0+阅读 · 2022年2月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

基于GPU的脉冲星宽带观测的相干消色散研究

国家自然科学基金

0+阅读 · 2013年12月31日

多GPU并行的热/化学反应非平衡N-S方程求解算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

考虑长期监测应力时序的钢箱梁桥疲劳评估方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

电磁场特征值问题的间断 Galerkin 算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

随机与动态环境下物流配送区域划分与配送路径集成优化问题研究

国家自然科学基金

0+阅读 · 2012年12月31日

几类新型目标罚函数理论与算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

浸入边界法的高效稳定数值格式

国家自然科学基金

0+阅读 · 2012年12月31日

基于高维全局分叉的水下航行器空间运动稳定性数值分析与自航模试验研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于深度变分模型及分散控制理论的机器人三维环境建模新方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于NURBS曲面的弹跳射线法的GPU加速

国家自然科学基金

0+阅读 · 2008年12月31日

Multi-Level Interaction Reranking with User Behavior History

Arxiv

0+阅读 · 2022年4月20日

Expert Finding in Legal Community Question Answering

Arxiv

0+阅读 · 2022年4月19日

SnapFuzz: An Efficient Fuzzing Framework for Network Applications

Arxiv

0+阅读 · 2022年4月19日

Reliable Actors with Retry Orchestration

Arxiv

0+阅读 · 2022年4月18日

Characterizing and Understanding Distributed GNN Training on GPUs

Arxiv

1+阅读 · 2022年4月18日

A Distributed and Elastic Aggregation Service for Scalable Federated Learning Systems

Arxiv

0+阅读 · 2022年4月16日

Few-shot Instruction Prompts for Pretrained Language Models to Detect Social Biases

Arxiv

0+阅读 · 2022年4月15日

Effects of Multi-Aspect Online Reviews with Unobserved Confounders: Estimation and Implication

Effects of Multi-Aspect Online Reviews with Unobserved Confounders: Estimation and Implication

Arxiv

0+阅读 · 2022年4月15日

Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions

Arxiv

31+阅读 · 2021年9月27日

mvn2vec: Preservation and Collaboration in Multi-View Network Embedding

Arxiv

10+阅读 · 2018年1月19日

VIP会员

文章信息

相关主题

相关VIP内容

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

热门VIP内容

开通专知VIP会员享更多权益服务

【博士论文】扩展可扩展会话推荐的边界

别想太多：高效 R1 风格大型推理模型综述

【ACMMM2025】EvoVLMA: 进化式视觉-语言模型自适应

智能体网络：用AI智能体编织下一代网络

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

Call for Nominations: 2022 Multimedia Prize Paper Award

Call for Nominations: 2022 Multimedia Prize Paper Award

CCF多媒体专委会

0+阅读 · 2022年2月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

相关论文

Multi-Level Interaction Reranking with User Behavior History

Arxiv

0+阅读 · 2022年4月20日

Expert Finding in Legal Community Question Answering

Arxiv

0+阅读 · 2022年4月19日

SnapFuzz: An Efficient Fuzzing Framework for Network Applications

Arxiv

0+阅读 · 2022年4月19日

Reliable Actors with Retry Orchestration

Arxiv

0+阅读 · 2022年4月18日

Characterizing and Understanding Distributed GNN Training on GPUs

Arxiv

1+阅读 · 2022年4月18日

A Distributed and Elastic Aggregation Service for Scalable Federated Learning Systems

Arxiv

0+阅读 · 2022年4月16日

Few-shot Instruction Prompts for Pretrained Language Models to Detect Social Biases

Arxiv

0+阅读 · 2022年4月15日

Effects of Multi-Aspect Online Reviews with Unobserved Confounders: Estimation and Implication

Effects of Multi-Aspect Online Reviews with Unobserved Confounders: Estimation and Implication

Arxiv

0+阅读 · 2022年4月15日

Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions

Arxiv

31+阅读 · 2021年9月27日

mvn2vec: Preservation and Collaboration in Multi-View Network Embedding

Arxiv

10+阅读 · 2018年1月19日

相关基金

基于GPU的脉冲星宽带观测的相干消色散研究

国家自然科学基金

0+阅读 · 2013年12月31日

多GPU并行的热/化学反应非平衡N-S方程求解算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

考虑长期监测应力时序的钢箱梁桥疲劳评估方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

电磁场特征值问题的间断 Galerkin 算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

随机与动态环境下物流配送区域划分与配送路径集成优化问题研究

国家自然科学基金

0+阅读 · 2012年12月31日

几类新型目标罚函数理论与算法研究

国家自然科学基金

0+阅读 · 2012年12月31日

浸入边界法的高效稳定数值格式

国家自然科学基金

0+阅读 · 2012年12月31日

基于高维全局分叉的水下航行器空间运动稳定性数值分析与自航模试验研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于深度变分模型及分散控制理论的机器人三维环境建模新方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于NURBS曲面的弹跳射线法的GPU加速

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员