使用 Risc-V 矢量指示优化 SpGEMM</s> (Optimization of SpGEMM with Risc-V vector instructions) - 专知论文

会员服务 ·

0

稀疏 · 列 · 向量化 · 哈希学习 · 块 ·

2023 年 3 月 4 日

Optimization of SpGEMM with Risc-V vector instructions

翻译：使用 Risc-V 矢量指示优化 SpGEMM

Valentin Le Fèvre,Marc Casas

The Sparse GEneral Matrix-Matrix multiplication (SpGEMM) $C = A \times B$ is a fundamental routine extensively used in domains like machine learning or graph analytics. Despite its relevance, the efficient execution of SpGEMM on vector architectures is a relatively unexplored topic. The most recent algorithm to run SpGEMM on these architectures is based on the SParse Accumulator (SPA) approach, and it is relatively efficient for sparse matrices featuring several tens of non-zero coefficients per column as it computes C columns one by one. However, when dealing with matrices containing just a few non-zero coefficients per column, the state-of-the-art algorithm is not able to fully exploit long vector architectures when computing the SpGEMM kernel. To overcome this issue we propose the SPA paRallel with Sorting (SPARS) algorithm, which computes in parallel several C columns among other optimizations, and the HASH algorithm, which uses dynamically sized hash tables to store intermediate output values. To combine the efficiency of SPA for relatively dense matrix blocks with the high performance that SPARS and HASH deliver for very sparse matrix blocks we propose H-SPA(t) and H-HASH(t), which dynamically switch between different algorithms. H-SPA(t) and H-HASH(t) obtain 1.24$\times$ and 1.57$\times$ average speed-ups with respect to SPA respectively, over a set of 40 sparse matrices obtained from the SuiteSparse Matrix Collection. For the 22 most sparse matrices, H-SPA(t) and H-HASH(t) deliver 1.42$\times$ and 1.99$\times$ average speed-ups respectively.

翻译：Sparse General 矩阵- Matrix 乘法 (SpGEMM) $42 = A = A 计数 C 列乘以 1 乘以 C 列时, 以零系数计数。但是, 处理仅包含几部非零的机器学习或图形分析等域的基本常规 B$ 。尽管相关, SpGEM 在矢量结构中高效执行 SpGEM 是一个相对未探索的专题。运行 SpGEMM 在这些结构中运行 SpGEM 的最近算法基于 SParse Across 累积( SPARM ) 方法, 而对于以几部非零系数计算每列数个非零系数的稀释矩阵来说相对有效。然而, 当处理仅包含几部非零位的 IMFlexcal 的矩阵时, 状态算法不能完全利用长期矢量的 SVAS- sal- sal- sal- sal- sal- sal lavedal 和 SH- sal- h- h- sal- sal- sal- h- sal- sal- sal- sal- sal- h- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- 和和和和和和和和和和 sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- sal- s- sal- s- s- sal- sal- sal- sal- sal- sal- sal- sal- s</s>

0

相关内容

【干货书】开放数据结构，Open Data Structures，337页pdf

【干货书】开放数据结构，Open Data Structures，337页pdf

专知会员服务

17+阅读 · 2021年9月17日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

11篇ICLR2020满分文章，来看看他们都在做什么？

11篇ICLR2020满分文章，来看看他们都在做什么？

专知

18+阅读 · 2019年11月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇自注意力机制(Self-attention)相关论文—结构化自注意力、相对位置、混合、句子表达、文本向量

【论文推荐】最新七篇自注意力机制(Self-attention)相关论文—结构化自注意力、相对位置、混合、句子表达、文本向量

专知

29+阅读 · 2018年3月12日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

《数学学报》期刊

国家自然科学基金

5+阅读 · 2015年12月31日

Toll样受体在中药成分保护肠黏膜微血管内皮细胞免受细菌毒素损伤中的作用研究

国家自然科学基金

0+阅读 · 2014年12月31日

有机半导体/无机纳晶杂化材料的界面控制及光电性质研究

国家自然科学基金

0+阅读 · 2013年12月31日

海洋弧菌菌群感应信号分子N-acyl homoserine lactones对NK细胞的调控作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

新型双极性给受体共聚物半导体的设计，合成与光电性质研究

国家自然科学基金

0+阅读 · 2012年12月31日

语音识别中的稀疏性深度学习

国家自然科学基金

11+阅读 · 2012年12月31日

Galectin-7在哮喘发病中的调控以及作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

基于Junction tree推理的多运动平台分散式协同导航算法研究

国家自然科学基金

2+阅读 · 2012年12月31日

SiO2复合材料表面CNTs生长及与TC4钛合金的复合反应钎焊机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

斑马鱼心脏发育

国家自然科学基金

0+阅读 · 2009年12月31日

LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions

Arxiv

0+阅读 · 2023年4月27日

Pushing the Boundaries of Tractable Multiperspective Reasoning: A Deduction Calculus for Standpoint EL+

Arxiv

0+阅读 · 2023年4月27日

Mixtures of Gaussian process experts based on kernel stick-breaking processes

Arxiv

0+阅读 · 2023年4月26日

An accelerated proximal gradient method for multiobjective optimization

Arxiv

0+阅读 · 2023年4月26日

SCV-GNN: Sparse Compressed Vector-based Graph Neural Network Aggregation

Arxiv

0+阅读 · 2023年4月26日

Splitting physics-informed neural networks for inferring the dynamics of integer- and fractional-order neuron models

Arxiv

0+阅读 · 2023年4月26日

BO-ICP: Initialization of Iterative Closest Point Based on Bayesian Optimization

Arxiv

0+阅读 · 2023年4月25日

Sequential Attention for Feature Selection

Arxiv

0+阅读 · 2023年4月25日

Alternating Local Enumeration (TnALE): Solving Tensor Network Structure Search with Fewer Evaluations

Arxiv

0+阅读 · 2023年4月25日

Determination of the effective cointegration rank in high-dimensional time-series predictive regressions

Arxiv

0+阅读 · 2023年4月25日

VIP会员

文章信息

相关主题

相关VIP内容

【干货书】开放数据结构，Open Data Structures，337页pdf

【干货书】开放数据结构，Open Data Structures，337页pdf

专知会员服务

17+阅读 · 2021年9月17日

ICLR 2021杰出论文奖出炉，8篇论文上榜！

专知会员服务

26+阅读 · 2021年4月2日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

操作系统智能体：基于多模态大模型（MLLM）的通用计算设备智能体综述

《美国太空军系统全生命周期建模、仿真与分析效能提升方案》最新84页报告

【博士论文】推进数据高效的深度学习：非参数 Transformer、主动测试与上下文学习

自主人工智能：未来战争是否将是自主化的？

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

11篇ICLR2020满分文章，来看看他们都在做什么？

11篇ICLR2020满分文章，来看看他们都在做什么？

专知

18+阅读 · 2019年11月7日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇自注意力机制(Self-attention)相关论文—结构化自注意力、相对位置、混合、句子表达、文本向量

【论文推荐】最新七篇自注意力机制(Self-attention)相关论文—结构化自注意力、相对位置、混合、句子表达、文本向量

专知

29+阅读 · 2018年3月12日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

相关论文

LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions

Arxiv

0+阅读 · 2023年4月27日

Pushing the Boundaries of Tractable Multiperspective Reasoning: A Deduction Calculus for Standpoint EL+

Arxiv

0+阅读 · 2023年4月27日

Mixtures of Gaussian process experts based on kernel stick-breaking processes

Arxiv

0+阅读 · 2023年4月26日

An accelerated proximal gradient method for multiobjective optimization

Arxiv

0+阅读 · 2023年4月26日

SCV-GNN: Sparse Compressed Vector-based Graph Neural Network Aggregation

Arxiv

0+阅读 · 2023年4月26日

Splitting physics-informed neural networks for inferring the dynamics of integer- and fractional-order neuron models

Arxiv

0+阅读 · 2023年4月26日

BO-ICP: Initialization of Iterative Closest Point Based on Bayesian Optimization

Arxiv

0+阅读 · 2023年4月25日

Sequential Attention for Feature Selection

Arxiv

0+阅读 · 2023年4月25日

Alternating Local Enumeration (TnALE): Solving Tensor Network Structure Search with Fewer Evaluations

Arxiv

0+阅读 · 2023年4月25日

Determination of the effective cointegration rank in high-dimensional time-series predictive regressions

Arxiv

0+阅读 · 2023年4月25日

相关基金

《数学学报》期刊

国家自然科学基金

5+阅读 · 2015年12月31日

Toll样受体在中药成分保护肠黏膜微血管内皮细胞免受细菌毒素损伤中的作用研究

国家自然科学基金

0+阅读 · 2014年12月31日

有机半导体/无机纳晶杂化材料的界面控制及光电性质研究

国家自然科学基金

0+阅读 · 2013年12月31日

海洋弧菌菌群感应信号分子N-acyl homoserine lactones对NK细胞的调控作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

新型双极性给受体共聚物半导体的设计，合成与光电性质研究

国家自然科学基金

0+阅读 · 2012年12月31日

语音识别中的稀疏性深度学习

国家自然科学基金

11+阅读 · 2012年12月31日

Galectin-7在哮喘发病中的调控以及作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

基于Junction tree推理的多运动平台分散式协同导航算法研究

国家自然科学基金

2+阅读 · 2012年12月31日

SiO2复合材料表面CNTs生长及与TC4钛合金的复合反应钎焊机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

斑马鱼心脏发育

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员