优化BHAC的混合平行化 (Optimizing the hybrid parallelization of BHAC) - 专知论文

会员服务 ·

0

Performer · 优化器 · 可辨认的 · 缩放 · Guidance ·

2021 年 8 月 27 日

Optimizing the hybrid parallelization of BHAC

翻译：优化BHAC的混合平行化

Salvatore Cielo,Oliver Porth,Luigi Iapichino,Anupam Karmakar,Hector Olivares,Chun Xia

from arxiv, 10 pages, 9 figures, 1 table; in review

We present our experience with the modernization on the GR-MHD code BHAC, aimed at improving its novel hybrid (MPI+OpenMP) parallelization scheme. In doing so, we showcase the use of performance profiling tools usable on x86 (Intel-based) architectures. Our performance characterization and threading analysis provided guidance in improving the concurrency and thus the efficiency of the OpenMP parallel regions. We assess scaling and communication patterns in order to identify and alleviate MPI bottlenecks, with both runtime switches and precise code interventions. The performance of optimized version of BHAC improved by $\sim28\%$, making it viable for scaling on several hundreds of supercomputer nodes. We finally test whether porting such optimizations to different hardware is likewise beneficial on the new architecture by running on ARM A64FX vector nodes.

翻译：我们介绍了我们关于GR-MHD代码BHAC现代化的经验,其目的是改进其新型混合(MPI+OpenMP)平行计划。在这样做的过程中,我们展示了在x86(基于 Intel)结构中可用的性能特征分析工具的使用情况。我们的性能特征和线性分析为改进同值货币从而提高OpenMP平行区域的效率提供了指导。我们评估了规模和通信模式,以便查明和缓解MPI瓶颈,包括运行时间开关和精确的代码干预。BHAC的优化版本的性能提高了$\sim28 ⁇ $,使之可以推广到数百个超级计算机节点。我们最后通过运行ARM A64FX矢量节点来测试将这种优化移植到不同的硬件是否同样有益于新架构。

0

相关内容

Performer

【ICML2021】异质风险最小化，Heterogeneous Risk Minimization

专知会员服务

16+阅读 · 2021年5月21日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

54+阅读 · 2020年9月7日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

【IJCAI 2019】基于时间的规划:理论与实践（Timeline-based Planning: Theory and Practice），Nicola Gigante，Angelo Montanari

【IJCAI 2019】基于时间的规划:理论与实践（Timeline-based Planning: Theory and Practice），Nicola Gigante，Angelo Montanari

专知会员服务

9+阅读 · 2019年8月10日

CCF推荐 | 国际会议信息8条

CCF推荐 | 国际会议信息8条

Call4Papers

9+阅读 · 2019年5月23日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Call for Participation: Shared Tasks in NLPCC 2019

Call for Participation: Shared Tasks in NLPCC 2019

中国计算机学会

5+阅读 · 2019年3月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

逆强化学习几篇论文笔记

逆强化学习几篇论文笔记

CreateAMind

9+阅读 · 2018年12月13日

条件GAN重大改进！cGANs with Projection Discriminator

条件GAN重大改进！cGANs with Projection Discriminator

CreateAMind

8+阅读 · 2018年2月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

Adversarial Variational Bayes: Unifying VAE and GAN 代码

Adversarial Variational Bayes: Unifying VAE and GAN 代码

CreateAMind

7+阅读 · 2017年10月4日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

Learning Stochastic Majority Votes by Minimizing a PAC-Bayes Generalization Bound

Arxiv

0+阅读 · 2021年10月19日

Correct Probabilistic Model Checking with Floating-Point Arithmetic

Arxiv

0+阅读 · 2021年10月17日

Minimal Conditions for Beneficial Local Search

Arxiv

0+阅读 · 2021年10月17日

Noise-Augmented Privacy-Preserving Empirical Risk Minimization with Dual-purpose Regularizer and Privacy Budget Retrieval and Recycling

Arxiv

0+阅读 · 2021年10月16日

Complexity of optimizing over the integers

Arxiv

0+阅读 · 2021年10月15日

BPPChecker: An SMT-based Model Checker on Basic Parallel Processes

Arxiv

0+阅读 · 2021年10月15日

A "Proof" of $P\neq NP$

Arxiv

0+阅读 · 2021年10月14日

ZARTS: On Zero-order Optimization for Neural Architecture Search

Arxiv

0+阅读 · 2021年10月10日

How Powerful are Graph Neural Networks?

Arxiv

23+阅读 · 2018年10月1日

MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning

Arxiv

4+阅读 · 2018年1月11日

VIP会员

文章信息

相关主题

相关VIP内容

【ICML2021】异质风险最小化，Heterogeneous Risk Minimization

专知会员服务

16+阅读 · 2021年5月21日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

54+阅读 · 2020年9月7日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

【IJCAI 2019】基于时间的规划:理论与实践（Timeline-based Planning: Theory and Practice），Nicola Gigante，Angelo Montanari

【IJCAI 2019】基于时间的规划:理论与实践（Timeline-based Planning: Theory and Practice），Nicola Gigante，Angelo Montanari

专知会员服务

9+阅读 · 2019年8月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《俄乌战争背景下俄罗斯的战略性海军分析（2022-2025年）》最新100页报告

【斯坦福博士论文】数据、决策与依赖：构建可信人工智能的挑战

人工智能时代背景下的未来海战

接触战中的无人机优势：美军旅级部队面临的小型无人机系统挑战与调整

相关资讯

CCF推荐 | 国际会议信息8条

CCF推荐 | 国际会议信息8条

Call4Papers

9+阅读 · 2019年5月23日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Call for Participation: Shared Tasks in NLPCC 2019

Call for Participation: Shared Tasks in NLPCC 2019

中国计算机学会

5+阅读 · 2019年3月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

逆强化学习几篇论文笔记

逆强化学习几篇论文笔记

CreateAMind

9+阅读 · 2018年12月13日

条件GAN重大改进！cGANs with Projection Discriminator

条件GAN重大改进！cGANs with Projection Discriminator

CreateAMind

8+阅读 · 2018年2月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

Adversarial Variational Bayes: Unifying VAE and GAN 代码

Adversarial Variational Bayes: Unifying VAE and GAN 代码

CreateAMind

7+阅读 · 2017年10月4日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

相关论文

Learning Stochastic Majority Votes by Minimizing a PAC-Bayes Generalization Bound

Arxiv

0+阅读 · 2021年10月19日

Correct Probabilistic Model Checking with Floating-Point Arithmetic

Arxiv

0+阅读 · 2021年10月17日

Minimal Conditions for Beneficial Local Search

Arxiv

0+阅读 · 2021年10月17日

Noise-Augmented Privacy-Preserving Empirical Risk Minimization with Dual-purpose Regularizer and Privacy Budget Retrieval and Recycling

Arxiv

0+阅读 · 2021年10月16日

Complexity of optimizing over the integers

Arxiv

0+阅读 · 2021年10月15日

BPPChecker: An SMT-based Model Checker on Basic Parallel Processes

Arxiv

0+阅读 · 2021年10月15日

A "Proof" of $P\neq NP$

Arxiv

0+阅读 · 2021年10月14日

ZARTS: On Zero-order Optimization for Neural Architecture Search

Arxiv

0+阅读 · 2021年10月10日

How Powerful are Graph Neural Networks?

Arxiv

23+阅读 · 2018年10月1日

MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning

Arxiv

4+阅读 · 2018年1月11日

微信扫码咨询专知VIP会员