《改进学习业绩守则》编辑 (Learning Performance-Improving Code Edits) - 专知论文

会员服务 ·

0

Performer · Continuity · MoDELS · 语言模型化 · 代码 ·

2023 年 2 月 15 日

Learning Performance-Improving Code Edits

翻译：《改进学习业绩守则》编辑

Aman Madaan,Alexander Shypula,Uri Alon,Milad Hashemi,Parthasarathy Ranganathan,Yiming Yang,Graham Neubig,Amir Yazdanbakhsh

from arxiv, https://pie4perf.com/

The waning of Moore's Law has shifted the focus of the tech industry towards alternative methods for continued performance gains. While optimizing compilers are a standard tool to help increase program efficiency, programmers continue to shoulder much responsibility in crafting and refactoring code with better performance characteristics. In this paper, we investigate the ability of large language models (LLMs) to suggest functionally correct, performance improving code edits. We hypothesize that language models can suggest such edits in ways that would be impractical for static analysis alone. We investigate these questions by curating a large-scale dataset of Performance-Improving Edits, PIE. PIE contains trajectories of programs, where a programmer begins with an initial, slower version and iteratively makes changes to improve the program's performance. We use PIE to evaluate and improve the capacity of large language models. Specifically, use examples from PIE to fine-tune multiple variants of CODEGEN, a billion-scale Transformer-decoder model. Additionally, we use examples from PIE to prompt OpenAI's CODEX using a few-shot prompting. By leveraging PIE, we find that both CODEX and CODEGEN can generate performance-improving edits, with speedups of more than 2.5x for over 25% of the programs, for C++ and Python, even after the C++ programs were compiled using the O3 optimization level. Crucially, we show that PIE allows CODEGEN, an open-sourced and 10x smaller model than CODEX, to match the performance of CODEX on this challenging task. Overall, this work opens new doors for creating systems and methods that can help programmers write efficient code.

翻译：Moore Law 的衰减使技术产业的重点转向了持续绩效增益的替代方法。虽然优化编译者是帮助提高程序效率的一个标准工具, 但程序员在编译和重构功能特点更好的代码方面继续承担着很大的责任。在本文中, 我们调查大型语言模型(LLIMs) 的能力, 以建议功能正确、性能改进代码编辑。我们假设语言模型可以建议这类编辑方式, 仅对静态分析来说是不切实际的。我们通过整理绩效改进编辑的大型数据集( PIE ), PIE 是一个标准工具, 帮助提高程序的效率, 在编译程序初始、较慢的版本和迭代修改程序方面, 我们用PIE 来评估和微调的多变异模式, 10亿级的变异模式, 更小的变异模式。此外, 我们从 PIEE 到快速的 OODI 的 CODEX, 使用微缩略的 CODUD, 和 CIE 快速化程序, 生成了更具有挑战性的业绩, 25 GEN 的 CIEODOD,,, 的CODODD, 和我们发现我们用新的系统可以比 CIERDODODODDDDD 更快的系统, 更更更更更新的C- CODODD 的系统, 更新了和 25 更新的系统。

0

相关内容

Performer

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

专知会员服务

35+阅读 · 2022年3月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

pytorch-pretrained-BERT：BERT PyTorch实现，可加载Google BERT预训练模型

pytorch-pretrained-BERT：BERT PyTorch实现，可加载Google BERT预训练模型

AINLP

35+阅读 · 2018年11月6日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

CRISPR/Cas9介导的基因组进化构建固态发酵耐热酵母及机理研究

国家自然科学基金

0+阅读 · 2016年12月31日

高光谱光学近场显微成像方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

哺乳动物体细胞克隆胚胎抗氧化机理的研究

国家自然科学基金

0+阅读 · 2012年12月31日

Fibulin-5/β1-integrin 信号通路在醛固酮诱导血管平滑肌细胞凋亡中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

基于OC-seislet变换的三维叠前复杂地震波场迭代数据插值方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于稀疏反演算法的地震波场重建

国家自然科学基金

0+阅读 · 2012年12月31日

可见及近红外宽光谱响应的高效固态量子点敏化太阳能电池

国家自然科学基金

0+阅读 · 2012年12月31日

土壤拟除虫菊酯低浓度长期暴露的毒性效应及致毒机理

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米及纳米颗粒两相湍流的研究

国家自然科学基金

0+阅读 · 2011年12月31日

胶体晶体模板法制备SOFC有序纳米阴极及机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

EZClone: Improving DNN Model Extraction Attack via Shape Distillation from GPU Execution Profiles

Arxiv

0+阅读 · 2023年4月6日

Diff-Font: Diffusion Model for Robust One-Shot Font Generation

Arxiv

0+阅读 · 2023年4月6日

VindLU: A Recipe for Effective Video-and-Language Pretraining

Arxiv

0+阅读 · 2023年4月5日

A Call to Reflect on Evaluation Practices for Failure Detection in Image Classification

Arxiv

0+阅读 · 2023年4月5日

oBERTa: Improving Sparse Transfer Learning via improved initialization, distillation, and pruning regimes

Arxiv

0+阅读 · 2023年4月4日

Q2ATransformer: Improving Medical VQA via an Answer Querying Decoder

Arxiv

1+阅读 · 2023年4月4日

An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning

Arxiv

0+阅读 · 2023年4月3日

Improving Passage Retrieval with Zero-Shot Question Generation

Arxiv

0+阅读 · 2023年4月3日

Improving RF-DNA Fingerprinting Performance in an Indoor Multipath Environment Using Semi-Supervised Learning

Arxiv

0+阅读 · 2023年4月2日

Reasoning in Dialog: Improving Response Generation by Context Reading Comprehension

Arxiv

12+阅读 · 2020年12月14日

VIP会员

文章信息

相关主题

语言模型化

相关VIP内容

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

专知会员服务

35+阅读 · 2022年3月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

面向性能、成本效益、云边隐私与可信性的大小语言模型协作综述

乌克兰太空研究（2022-2024年） | 176页

【CMU博士论文】大型语言模型的隐性特性

国防领域人工智能走向何方？

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

pytorch-pretrained-BERT：BERT PyTorch实现，可加载Google BERT预训练模型

pytorch-pretrained-BERT：BERT PyTorch实现，可加载Google BERT预训练模型

AINLP

35+阅读 · 2018年11月6日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

相关论文

EZClone: Improving DNN Model Extraction Attack via Shape Distillation from GPU Execution Profiles

Arxiv

0+阅读 · 2023年4月6日

Diff-Font: Diffusion Model for Robust One-Shot Font Generation

Arxiv

0+阅读 · 2023年4月6日

VindLU: A Recipe for Effective Video-and-Language Pretraining

Arxiv

0+阅读 · 2023年4月5日

A Call to Reflect on Evaluation Practices for Failure Detection in Image Classification

Arxiv

0+阅读 · 2023年4月5日

oBERTa: Improving Sparse Transfer Learning via improved initialization, distillation, and pruning regimes

Arxiv

0+阅读 · 2023年4月4日

Q2ATransformer: Improving Medical VQA via an Answer Querying Decoder

Arxiv

1+阅读 · 2023年4月4日

An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning

Arxiv

0+阅读 · 2023年4月3日

Improving Passage Retrieval with Zero-Shot Question Generation

Arxiv

0+阅读 · 2023年4月3日

Improving RF-DNA Fingerprinting Performance in an Indoor Multipath Environment Using Semi-Supervised Learning

Arxiv

0+阅读 · 2023年4月2日

Reasoning in Dialog: Improving Response Generation by Context Reading Comprehension

Arxiv

12+阅读 · 2020年12月14日

相关基金

CRISPR/Cas9介导的基因组进化构建固态发酵耐热酵母及机理研究

国家自然科学基金

0+阅读 · 2016年12月31日

高光谱光学近场显微成像方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

哺乳动物体细胞克隆胚胎抗氧化机理的研究

国家自然科学基金

0+阅读 · 2012年12月31日

Fibulin-5/β1-integrin 信号通路在醛固酮诱导血管平滑肌细胞凋亡中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

基于OC-seislet变换的三维叠前复杂地震波场迭代数据插值方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于稀疏反演算法的地震波场重建

国家自然科学基金

0+阅读 · 2012年12月31日

可见及近红外宽光谱响应的高效固态量子点敏化太阳能电池

国家自然科学基金

0+阅读 · 2012年12月31日

土壤拟除虫菊酯低浓度长期暴露的毒性效应及致毒机理

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米及纳米颗粒两相湍流的研究

国家自然科学基金

0+阅读 · 2011年12月31日

胶体晶体模板法制备SOFC有序纳米阴极及机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员