缩略图: 实际为 I/O 最佳多线性代数 (Deinsum: Practically I/O Optimal Multilinear Algebra) - 专知论文

会员服务 ·

0

Tensor · Performer · 优化器 · state-of-the-art · 核化 ·

2022 年 6 月 16 日

Deinsum: Practically I/O Optimal Multilinear Algebra

翻译：缩略图: 实际为 I/O 最佳多线性代数

Alexandros Nikolaos Ziogas,Grzegorz Kwasniewski,Tal Ben-Nun,Timo Schneider,Torsten Hoefler

Multilinear algebra kernel performance on modern massively-parallel systems is determined mainly by data movement. However, deriving data movement-optimal distributed schedules for programs with many high-dimensional inputs is a notoriously hard problem. State-of-the-art libraries rely on heuristics and often fall back to suboptimal tensor folding and BLAS calls. We present Deinsum, an automated framework for distributed multilinear algebra computations expressed in Einstein notation, based on rigorous mathematical tools to address this problem. Our framework automatically derives data movement-optimal tiling and generates corresponding distributed schedules, further optimizing the performance of local computations by increasing their arithmetic intensity. To show the benefits of our approach, we test it on two important tensor kernel classes: Matricized Tensor Times Khatri-Rao Products and Tensor Times Matrix chains. We show performance results and scaling on the Piz Daint supercomputer, with up to 19x speedup over state-of-the-art solutions on 512 nodes.

翻译：多线性代数内核在现代大规模平行系统中的多线性能主要由数据流动决定。然而,为含有许多高维投入的程式生成数据移动最优化分布时间表是一个臭名昭著的难题。最先进的图书馆依赖于超光学,往往会退回到亚优的极点折叠和 BLAS 调用。我们介绍Deinsum, 这是一个基于严格数学工具的爱因斯坦符号表达的分布式多线性代数计算自动框架,用以解决这一问题。我们的框架自动生成数据移动-最优化的饱和并生成相应的分布时间表,通过提高算术强度进一步优化本地计算绩效。为了展示我们的方法的好处,我们测试了两个重要的格子内核级:数学泰森泰姆时的Khatri-Rao产品和Tensor Tensor Tens 矩阵链。我们展示了业绩结果,并在Piz Daint 超级计算机上提升了规模,在512个节点上超越了状态的解决方案,达到19x速度。

0

相关内容

Tensor

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

经典书《机器学习：概率视角》（Machine Learning: a Probabilistic Perspective）第二版Python代码，附1098页pdf下载

经典书《机器学习：概率视角》（Machine Learning: a Probabilistic Perspective）第二版Python代码，附1098页pdf下载

专知会员服务

275+阅读 · 2019年10月25日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

基于对称识别方法的贝叶斯probit模型稳健性研究

国家自然科学基金

3+阅读 · 2015年12月31日

Schr？dinger-Poisson方程守恒DDG方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

量子级联激光种子注入高重频窄脉冲CO2激光再生放大技术

国家自然科学基金

0+阅读 · 2014年12月31日

三维椭圆方程Cauchy问题的正则化方法

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

茶树ERF类转录因子家族的克隆及功能鉴定

国家自然科学基金

0+阅读 · 2012年12月31日

益气活血方通过"DAMPs-PRRs-巨噬细胞"途径影响动脉粥样硬化斑块易损性的机制

国家自然科学基金

0+阅读 · 2011年12月31日

燃煤火焰中间产物的在线监测与污染气体排放预测研究

国家自然科学基金

0+阅读 · 2009年12月31日

Accelerating the Sinkhorn algorithm for sparse multi-marginal optimal transport by fast Fourier transforms

Arxiv

0+阅读 · 2022年8月5日

A Gaze into the Internal Logic of Graph Neural Networks, with Logic

Arxiv

0+阅读 · 2022年8月5日

Learning to Re-weight Examples with Optimal Transport for Imbalanced Classification

Arxiv

0+阅读 · 2022年8月5日

Bayesian Optimization For Multi-Objective Mixed-Variable Problems

Arxiv

0+阅读 · 2022年8月4日

Efficiently Generating Independent Samples Directly from the Posterior Distribution for a Large Class of Bayesian Generalized Linear Mixed Effects Models

Arxiv

0+阅读 · 2022年8月4日

On the Identity Problem and the Group Problem for subsemigroups of unipotent matrix groups

Arxiv

0+阅读 · 2022年8月4日

A Simple and Tighter Derivation of Achievability for Classical Communication over Quantum Channels

A Simple and Tighter Derivation of Achievability for Classical Communication over Quantum Channels

Arxiv

0+阅读 · 2022年8月3日

Estimation of sub-Gaussian random vectors using the method of moments

Arxiv

0+阅读 · 2022年8月3日

An Algorithm for Ennola's Second Theorem and Counting Smooth Numbers in Practice

Arxiv

0+阅读 · 2022年8月2日

Class-Balanced Loss Based on Effective Number of Samples

Arxiv

12+阅读 · 2019年1月16日

VIP会员

文章信息

相关主题

state-of-the-art

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

经典书《机器学习：概率视角》（Machine Learning: a Probabilistic Perspective）第二版Python代码，附1098页pdf下载

经典书《机器学习：概率视角》（Machine Learning: a Probabilistic Perspective）第二版Python代码，附1098页pdf下载

专知会员服务

275+阅读 · 2019年10月25日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《多智能体不确定环境追逃博弈研究》216页

美智库最新发布《解放军"人机编组协同作战"发展路径：理论与实践》53页

现代战争"杀伤区"理论：空间尺度与结构特征、控制手段与毁伤机制、生存策略与战线转移

《俄军无人机创新技术或已在乌克兰达成"战场空中封锁"作战效果》最新18页报告

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Accelerating the Sinkhorn algorithm for sparse multi-marginal optimal transport by fast Fourier transforms

Arxiv

0+阅读 · 2022年8月5日

A Gaze into the Internal Logic of Graph Neural Networks, with Logic

Arxiv

0+阅读 · 2022年8月5日

Learning to Re-weight Examples with Optimal Transport for Imbalanced Classification

Arxiv

0+阅读 · 2022年8月5日

Bayesian Optimization For Multi-Objective Mixed-Variable Problems

Arxiv

0+阅读 · 2022年8月4日

Efficiently Generating Independent Samples Directly from the Posterior Distribution for a Large Class of Bayesian Generalized Linear Mixed Effects Models

Arxiv

0+阅读 · 2022年8月4日

On the Identity Problem and the Group Problem for subsemigroups of unipotent matrix groups

Arxiv

0+阅读 · 2022年8月4日

A Simple and Tighter Derivation of Achievability for Classical Communication over Quantum Channels

A Simple and Tighter Derivation of Achievability for Classical Communication over Quantum Channels

Arxiv

0+阅读 · 2022年8月3日

Estimation of sub-Gaussian random vectors using the method of moments

Arxiv

0+阅读 · 2022年8月3日

An Algorithm for Ennola's Second Theorem and Counting Smooth Numbers in Practice

Arxiv

0+阅读 · 2022年8月2日

Class-Balanced Loss Based on Effective Number of Samples

Arxiv

12+阅读 · 2019年1月16日

相关基金

Hamilton-Jacibi方程的弱KAM理论

国家自然科学基金

2+阅读 · 2017年12月31日

基于对称识别方法的贝叶斯probit模型稳健性研究

国家自然科学基金

3+阅读 · 2015年12月31日

Schr？dinger-Poisson方程守恒DDG方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

量子级联激光种子注入高重频窄脉冲CO2激光再生放大技术

国家自然科学基金

0+阅读 · 2014年12月31日

三维椭圆方程Cauchy问题的正则化方法

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

茶树ERF类转录因子家族的克隆及功能鉴定

国家自然科学基金

0+阅读 · 2012年12月31日

益气活血方通过"DAMPs-PRRs-巨噬细胞"途径影响动脉粥样硬化斑块易损性的机制

国家自然科学基金

0+阅读 · 2011年12月31日

燃煤火焰中间产物的在线监测与污染气体排放预测研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员