标题： (Towards Generating Functionally Correct Code Edits from Natural Language Issue Descriptions) - 专知论文

会员服务 ·

0

自然语言描述 · 代码 · 基准 · 数据集 · 基准测试 ·

2023 年 4 月 7 日

Towards Generating Functionally Correct Code Edits from Natural Language Issue Descriptions

翻译：标题：

Sarah Fakhoury,Saikat Chakraborty,Madan Musuvathi,Shuvendu K. Lahiri

Large language models (LLMs), such as OpenAI's Codex, have demonstrated their potential to generate code from natural language descriptions across a wide range of programming tasks. Several benchmarks have recently emerged to evaluate the ability of LLMs to generate functionally correct code from natural language intent with respect to a set of hidden test cases. This has enabled the research community to identify significant and reproducible advancements in LLM capabilities. However, there is currently a lack of benchmark datasets for assessing the ability of LLMs to generate functionally correct code edits based on natural language descriptions of intended changes. This paper aims to address this gap by motivating the problem NL2Fix of translating natural language descriptions of code changes (namely bug fixes described in Issue reports in repositories) into correct code fixes. To this end, we introduce Defects4J-NL2Fix, a dataset of 283 Java programs from the popular Defects4J dataset augmented with high-level descriptions of bug fixes, and empirically evaluate the performance of several state-of-the-art LLMs for the this task. Results show that these LLMS together are capable of generating plausible fixes for 64.6% of the bugs, and the best LLM-based technique can achieve up to 21.20% top-1 and 35.68% top-5 accuracy on this benchmark.

翻译：从自然语言问题描述中生成功能正确的代码编辑摘要：大型语言模型（LLM）如OpenAI的Codex已经展示了它们在广泛的编程任务中从自然语言描述生成代码的潜力。最近出现了几个基准测试，以评估LLM从自然语言意图生成功能正确代码的能力，涉及一组隐藏测试用例。这使得研究社区能够确定LLM性能中显着且可重复的进展。但是，当前缺乏基准数据集，用于评估LLM根据旨在改变的自然语言描述生成功能正确的代码编辑的能力。本文旨在通过提出将自然语言代码更改的描述（即存储库中的问题报告中描述的错误修复）翻译为正确代码修复的NL2Fix问题，来填补这一空白。为此，我们引入了Defects4J-NL2Fix数据集，该数据集包含283个Java程序，这些程序来自流行的Defects4J数据集，其中加入了修复错误的高级描述，并对该任务的几种最先进的LLM进行了经验性评估。结果显示，这些LLMS的组合能够对64.6％的错误生成合理的修复，而最佳的基于LLM的技术可以在这个基准测试上达到21.20％的top-1和35.68％的top-5准确度。

0

相关内容

自然语言描述

自然语言描述

【EMNLP2020】自然语言生成，Neural Language Generation

【EMNLP2020】自然语言生成，Neural Language Generation

专知会员服务

39+阅读 · 2020年11月20日

【ACL2020】对抗性文本生成，Improving Adversarial Text Generation

专知会员服务

52+阅读 · 2020年5月5日

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

专知会员服务

33+阅读 · 2020年5月2日

【ACL2020】生成事实验证解释，Generating Fact Checking Explanations

【ACL2020】生成事实验证解释，Generating Fact Checking Explanations

专知会员服务

17+阅读 · 2020年4月15日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

论文浅尝 | Language Models (Mostly) Know What They Know

论文浅尝 | Language Models (Mostly) Know What They Know

开放知识图谱

2+阅读 · 2022年11月18日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

论文浅尝 | Global Relation Embedding for Relation Extraction

论文浅尝 | Global Relation Embedding for Relation Extraction

开放知识图谱

12+阅读 · 2019年3月3日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

谷歌发表的史上最强NLP模型BERT的官方代码和预训练模型可以下载了

谷歌发表的史上最强NLP模型BERT的官方代码和预训练模型可以下载了

AINLP

12+阅读 · 2018年11月1日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

ABCB6基因在眼组织缺损中的功能研究

国家自然科学基金

0+阅读 · 2014年12月31日

Hippo信号通路调控间充质干细胞向ARDS肺泡上皮细胞分化的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

测地流的动力学研究

国家自然科学基金

0+阅读 · 2013年12月31日

表面等离子体耦合宽频上转换增强TiO2光催化性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

microRNA-708及其调控Notch信号通路介导的血管生成在新生儿支气管肺发育不良中的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

动态面孔语音情绪的整合加工及神经生理机制

国家自然科学基金

0+阅读 · 2013年12月31日

面向协作生成服务的社交搜索研究

国家自然科学基金

0+阅读 · 2013年12月31日

利用异种脱细胞软骨膜复合静电纺可降解纳米膜构建组织工程气管的实验研究

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

心肌灌注与心脏再同步化治疗疗效相关性的超声研究

国家自然科学基金

0+阅读 · 2009年12月31日

Text-to-SQL Error Correction with Language Models of Code

Arxiv

0+阅读 · 2023年5月28日

MVP: Multi-task Supervised Pre-training for Natural Language Generation

Arxiv

0+阅读 · 2023年5月28日

Define, Evaluate, and Improve Task-Oriented Cognitive Capabilities for Instruction Generation Models

Arxiv

0+阅读 · 2023年5月28日

Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model

Arxiv

1+阅读 · 2023年5月26日

RFiD: Towards Rational Fusion-in-Decoder for Open-Domain Question Answering

Arxiv

0+阅读 · 2023年5月26日

On Evaluating Adversarial Robustness of Large Vision-Language Models

Arxiv

0+阅读 · 2023年5月26日

DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions

Arxiv

0+阅读 · 2023年5月26日

Multimedia Generative Script Learning for Task Planning

Arxiv

0+阅读 · 2023年5月26日

Abstractive Summary Generation for the Urdu Language

Arxiv

0+阅读 · 2023年5月25日

Deep Learning for UAV-based Object Detection and Tracking: A Survey

Arxiv

62+阅读 · 2021年10月25日

VIP会员

文章信息

相关主题

自然语言描述

相关VIP内容

【EMNLP2020】自然语言生成，Neural Language Generation

【EMNLP2020】自然语言生成，Neural Language Generation

专知会员服务

39+阅读 · 2020年11月20日

【ACL2020】对抗性文本生成，Improving Adversarial Text Generation

专知会员服务

52+阅读 · 2020年5月5日

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

专知会员服务

33+阅读 · 2020年5月2日

【ACL2020】生成事实验证解释，Generating Fact Checking Explanations

【ACL2020】生成事实验证解释，Generating Fact Checking Explanations

专知会员服务

17+阅读 · 2020年4月15日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

热门VIP内容

开通专知VIP会员享更多权益服务

GPT-5如何对齐？从硬性拒绝到安全完成：走向以输出为中心的安全训练

【伯克利博士论文】超越人类监督的视觉智能

【ICCV2025】SO(3) 上连续非保守动力系统的预测

2025年中国数据要素行业发展研究报告

相关资讯

论文浅尝 | Language Models (Mostly) Know What They Know

论文浅尝 | Language Models (Mostly) Know What They Know

开放知识图谱

2+阅读 · 2022年11月18日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

论文浅尝 | Global Relation Embedding for Relation Extraction

论文浅尝 | Global Relation Embedding for Relation Extraction

开放知识图谱

12+阅读 · 2019年3月3日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

谷歌发表的史上最强NLP模型BERT的官方代码和预训练模型可以下载了

谷歌发表的史上最强NLP模型BERT的官方代码和预训练模型可以下载了

AINLP

12+阅读 · 2018年11月1日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

相关论文

Text-to-SQL Error Correction with Language Models of Code

Arxiv

0+阅读 · 2023年5月28日

MVP: Multi-task Supervised Pre-training for Natural Language Generation

Arxiv

0+阅读 · 2023年5月28日

Define, Evaluate, and Improve Task-Oriented Cognitive Capabilities for Instruction Generation Models

Arxiv

0+阅读 · 2023年5月28日

Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model

Arxiv

1+阅读 · 2023年5月26日

RFiD: Towards Rational Fusion-in-Decoder for Open-Domain Question Answering

Arxiv

0+阅读 · 2023年5月26日

On Evaluating Adversarial Robustness of Large Vision-Language Models

Arxiv

0+阅读 · 2023年5月26日

DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions

Arxiv

0+阅读 · 2023年5月26日

Multimedia Generative Script Learning for Task Planning

Arxiv

0+阅读 · 2023年5月26日

Abstractive Summary Generation for the Urdu Language

Arxiv

0+阅读 · 2023年5月25日

Deep Learning for UAV-based Object Detection and Tracking: A Survey

Arxiv

62+阅读 · 2021年10月25日

相关基金

ABCB6基因在眼组织缺损中的功能研究

国家自然科学基金

0+阅读 · 2014年12月31日

Hippo信号通路调控间充质干细胞向ARDS肺泡上皮细胞分化的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

测地流的动力学研究

国家自然科学基金

0+阅读 · 2013年12月31日

表面等离子体耦合宽频上转换增强TiO2光催化性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

microRNA-708及其调控Notch信号通路介导的血管生成在新生儿支气管肺发育不良中的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

动态面孔语音情绪的整合加工及神经生理机制

国家自然科学基金

0+阅读 · 2013年12月31日

面向协作生成服务的社交搜索研究

国家自然科学基金

0+阅读 · 2013年12月31日

利用异种脱细胞软骨膜复合静电纺可降解纳米膜构建组织工程气管的实验研究

国家自然科学基金

0+阅读 · 2012年12月31日

去酰基化ghrelin改善脂肪组织炎症所致胰岛素抵抗的机制- - 调节性T细胞的作用

国家自然科学基金

0+阅读 · 2011年12月31日

心肌灌注与心脏再同步化治疗疗效相关性的超声研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员