LoROCAT: Spark SQL 应用软件的低管理在线配置自动自动调试 (LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications) - 专知论文

会员服务 ·

0

Spark SQL · 簇 · Spark · Performer · 优化器 ·

2022 年 5 月 15 日

LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications

翻译：LoROCAT: Spark SQL 应用软件的低管理在线配置自动自动调试

Jinhan Xin,Kai Hwang,Zhibin Yu

from arxiv, 16 pages, 21 figures. This arxiv version is an extended version of the SIGMOD'2022 paper with same title, allowed by conference chairs

Spark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem, but suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others. To address these issues, we propose a novel Bayesian Optimization (BO) based approach named LOCAT to automatically tune the configurations of Spark SQL applications online. LOCAT innovates three techniques. The first technique, named QCSA, eliminates the configuration-insensitive queries by Query Configuration Sensitivity Analysis (QCSA) when collecting training samples. The second technique, dubbed DAGP, is a Datasize-Aware Gaussian Process (DAGP) which models the performance of an application as a distribution of functions of configuration parameters as well as input data size. The third technique, called IICP, Identifies Important Configuration Parameters (IICP) with respect to performance and only tunes the important ones. As such, LOCAT can tune the configurations of a Spark SQL application with low overhead and adapt to different input data sizes. We employ Spark SQL applications from benchmark suites TPC-DS, TPC-H, and HiBench running on two significantly different clusters, a four-node ARM cluster and an eight-node x86 cluster, to evaluate LOCAT. The experimental results on the ARM cluster show that LOCAT accelerates the optimization procedures of the state-of-the-art approaches by at least 4.1x and up to 9.7x; moreover, LOCAT improves the application performance by at least 1.9x and up to 2.4x. On the x86 cluster, LOCAT shows similar results to those on the ARM cluster.

翻译：Spark SQL 已经在行业中广泛部署 SQL 。最近的研究试图利用机器学习(ML) 解决这个问题,但有两个缺点。首先, 收集培训样本需要很长的时间( 高管理) 。第二, 同一应用程序的一个输入数据大小的最佳配置可能不是其他应用程序的最佳配置。为了解决这些问题, 我们建议采用名为 LOCAT (BO) 的新型Bayesian Optim化(BO) 方法, 自动调整 Spark SQL 应用程序的配置。 LOCAT 创新了三种技术。第一种技术, 名为 QCSA (QCSA), 消除了Query 配置敏感度分析(QCSA) 在收集培训样本时的配置不敏感度查询。第二种技术, 调制DGP( ), 是一个数据缩略图- Award Gauss 进程(DGP), 将应用程序的性能作为配置参数的分布以及输入数据大小。第三个技术, 名为 IICP, 识别重要配置参数(IICP), 有关业绩, 只标定了TROC- RodL 程序, 运行S- RDS 。

0

相关内容

Spark SQL

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

54+阅读 · 2021年1月20日

2020数据工程师成长路线图

专知会员服务

19+阅读 · 2020年9月6日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

蓖麻矮化相关RcDof基因功能分析及调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

二亚硝基哌嗪（DNP）介导Clusterin表达参与鼻咽癌转移的分子机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

含有缺失值的纵向数据回归模型的稳健推断

国家自然科学基金

3+阅读 · 2012年12月31日

铂族金属纳米颗粒的形貌与其不对称催化氢化性能的构效关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

金属/有机骨架化合物与多孔金属复合材料的制备与性能

国家自然科学基金

0+阅读 · 2012年12月31日

TFPI-2对巨噬细胞胆固醇流入/流出通路的作用及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

软骨细胞生长相关miRNA、C/EBPβ和Runx2循环协控促进梯度材料复合ADSCs成软骨分化的机制

国家自然科学基金

0+阅读 · 2011年12月31日

手性多孔有机无机杂化配位聚合物材料的离子热合成与性能研究

国家自然科学基金

0+阅读 · 2009年12月31日

SDF-1/CXCR4信号通路的干预及调节关节软骨退变的研究

国家自然科学基金

0+阅读 · 2008年12月31日

A Literature Review on Serverless Computing

A Literature Review on Serverless Computing

Arxiv

0+阅读 · 2022年7月6日

Predicting Out-of-Domain Generalization with Local Manifold Smoothness

Predicting Out-of-Domain Generalization with Local Manifold Smoothness

Arxiv

0+阅读 · 2022年7月5日

Best Subset Selection with Efficient Primal-Dual Algorithm

Arxiv

0+阅读 · 2022年7月5日

Application of multilayer perceptron with data augmentation in nuclear physics

Arxiv

0+阅读 · 2022年7月5日

QuPeD: Quantized Personalization via Distillation with Applications to Federated Learning

Arxiv

0+阅读 · 2022年7月5日

Effect of boundary conditions on a high-performance isolation hexapod platform

Arxiv

0+阅读 · 2022年7月5日

Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Arxiv

0+阅读 · 2022年7月1日

Prioritized training on points that are learnable, worth learning, and not yet learned (workshop version)

Arxiv

0+阅读 · 2022年7月1日

A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP

Arxiv

12+阅读 · 2021年8月30日

Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview

Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview

Arxiv

11+阅读 · 2019年9月19日

VIP会员

文章信息

相关主题

相关VIP内容

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

54+阅读 · 2021年1月20日

2020数据工程师成长路线图

专知会员服务

19+阅读 · 2020年9月6日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

视觉-语言-动作模型解析：从模块构成到里程碑与挑战

《解析陆域作战方向：一个概念性框架》报告

【博士论文】基于多模态基础模型的上下文学习

追寻真正的AI自主性：从遗留思维到战场优势

相关资讯

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

A Literature Review on Serverless Computing

A Literature Review on Serverless Computing

Arxiv

0+阅读 · 2022年7月6日

Predicting Out-of-Domain Generalization with Local Manifold Smoothness

Predicting Out-of-Domain Generalization with Local Manifold Smoothness

Arxiv

0+阅读 · 2022年7月5日

Best Subset Selection with Efficient Primal-Dual Algorithm

Arxiv

0+阅读 · 2022年7月5日

Application of multilayer perceptron with data augmentation in nuclear physics

Arxiv

0+阅读 · 2022年7月5日

QuPeD: Quantized Personalization via Distillation with Applications to Federated Learning

Arxiv

0+阅读 · 2022年7月5日

Effect of boundary conditions on a high-performance isolation hexapod platform

Arxiv

0+阅读 · 2022年7月5日

Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

Arxiv

0+阅读 · 2022年7月1日

Prioritized training on points that are learnable, worth learning, and not yet learned (workshop version)

Arxiv

0+阅读 · 2022年7月1日

A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP

Arxiv

12+阅读 · 2021年8月30日

Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview

Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview

Arxiv

11+阅读 · 2019年9月19日

相关基金

蓖麻矮化相关RcDof基因功能分析及调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

二亚硝基哌嗪（DNP）介导Clusterin表达参与鼻咽癌转移的分子机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

含有缺失值的纵向数据回归模型的稳健推断

国家自然科学基金

3+阅读 · 2012年12月31日

铂族金属纳米颗粒的形貌与其不对称催化氢化性能的构效关系研究

国家自然科学基金

0+阅读 · 2012年12月31日

金属/有机骨架化合物与多孔金属复合材料的制备与性能

国家自然科学基金

0+阅读 · 2012年12月31日

TFPI-2对巨噬细胞胆固醇流入/流出通路的作用及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

软骨细胞生长相关miRNA、C/EBPβ和Runx2循环协控促进梯度材料复合ADSCs成软骨分化的机制

国家自然科学基金

0+阅读 · 2011年12月31日

手性多孔有机无机杂化配位聚合物材料的离子热合成与性能研究

国家自然科学基金

0+阅读 · 2009年12月31日

SDF-1/CXCR4信号通路的干预及调节关节软骨退变的研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员