未受监督地学习用于改变守则的通用嵌套 (Unsupervised Learning of General-Purpose Embeddings for Code Changes)

A lot of problems in the field of software engineering - bug fixing, commit message generation, etc. - require analyzing not only the code itself but specifically code changes. Applying machine learning models to these tasks requires us to create numerical representations of the changes, i.e. embeddings. Recent studies demonstrate that the best way to obtain these embeddings is to pre-train a deep neural network in an unsupervised manner on a large volume of unlabeled data and then further fine-tune it for a specific task. In this work, we propose an approach for obtaining such embeddings of code changes during pre-training and evaluate them on two different downstream tasks - applying changes to code and commit message generation. The pre-training consists of the model learning to apply the given change (an edit sequence) to the code in a correct way, and therefore requires only the code change itself. To increase the quality of the obtained embeddings, we only consider the changed tokens in the edit sequence. In the task of applying code changes, our model outperforms the model that uses full edit sequences by 5.9 percentage points in accuracy. As for the commit message generation, our model demonstrated the same results as supervised models trained for this specific task, which indicates that it can encode code changes well and can be improved in the future by pre-training on a larger dataset of easily gathered code changes.

翻译：软件工程 — 错误修正、承诺生成信息等等 — 领域的许多问题不仅需要分析代码本身, 还需要分析代码本身, 还需要具体代码修改。应用机器学习模型来完成这些任务, 需要我们创建这些变化的数字表示, 即嵌入。最近的研究显示, 获得这些嵌入的最佳方式是, 以不受监督的方式对大量无标签数据进行深层神经网络前导, 然后对它进行进一步微调, 具体任务。在这项工作中, 我们建议了一种方法, 在培训前获得这种代码修改的嵌入, 并且对两个不同的下游任务进行评估 : 将修改应用到代码并承诺生成信息。预培训包括模型学习将给定的修改( 编辑序列) 正确应用到代码中, 因此只需要代码本身改变。为了提高获得的嵌入数据的质量, 我们只在编辑序列中考虑修改的符号。在应用代码修改的过程中, 我们的模型比模型要优于模型, 使用完全编辑序列的模型, 在精确度上应用5.9 百分比点。对于将修改的模型来说, 将用于将改进修改改进修改的修改的将修改用于修改修改修改的修改的的将用于修改修改修改用于修改的修改的用于修改的修改的的用于的修改修改修改修改修改修改的的的修改修改修改修改修改修改修改的修改的修改修改修改修改修改的修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改的修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改修改的修改的的的的的的修改修改的的的的的的修改修改修改修改修改修改修改修改的的

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【干货书】Python程序员编程，810页pdf，Python® for Programmers

专知会员服务

62+阅读 · 2020年8月6日

【微软亚洲研究院】无监督词嵌入对齐的几何感知域自适应，Geometry-aware Domain Adaptation for Unsupervised Alignment of Word Embeddings

专知会员服务

23+阅读 · 2020年4月21日

【ACL2020-Facebook AI】跨语言表示学习，Unsupervised Cross-lingual Representation Learning at Scale

专知会员服务

27+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日