关于利用深学习模式闻闻软件守则的可预测性的经验研究 (An Empirical Study on Predictability of Software Code Smell Using Deep Learning Models)

from arxiv, 12 pages, 6 Figures, 3 Tables, Accepted in the 35th International Conference on Advanced Information Networking and Applications (AINA-2021)

Code Smell, similar to a bad smell, is a surface indication of something tainted but in terms of software writing practices. This metric is an indication of a deeper problem lies within the code and is associated with an issue which is prominent to experienced software developers with acceptable coding practices. Recent studies have often observed that codes having code smells are often prone to a higher probability of change in the software development cycle. In this paper, we developed code smell prediction models with the help of features extracted from source code to predict eight types of code smell. Our work also presents the application of data sampling techniques to handle class imbalance problem and feature selection techniques to find relevant feature sets. Previous studies had made use of techniques such as Naive - Bayes and Random forest but had not explored deep learning methods to predict code smell. A total of 576 distinct Deep Learning models were trained using the features and datasets mentioned above. The study concluded that the deep learning models which used data from Synthetic Minority Oversampling Technique gave better results in terms of accuracy, AUC with the accuracy of some models improving from 88.47 to 96.84.

翻译：代码的嗅觉类似于一种臭味,它代表着一种被污染的事物的表面,但从软件写作做法的角度来看,它表明代码中存在一个更深的问题,它与一个对有经验的软件开发者具有可接受编码做法的突出问题有关。最近的研究经常发现,代码的嗅觉往往容易在软件开发周期中发生更大的变化。在这份文件中,我们在从源代码提取的特征的帮助下开发了代码的嗅觉预测模型,以预测8种代码的嗅觉。我们的工作还介绍了数据取样技术的应用,以处理阶级不平衡问题和特征选择技术,以找到相关的特征组。以前的研究利用了Naive-Bayes和随机森林等技术,但没有探索过预测代码嗅觉的深度学习方法。共有576个不同的深学习模型利用上述特征和数据集接受了培训。研究的结论是,使用合成少数群体过量抽样技术的数据的深学习模型在准确性方面产生了更好的结果。AUC在一些模型的精确性改进了88.47至96.84。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【干货书】机器学习Primer，122页pdf

专知会员服务

109+阅读 · 2020年10月5日

回顾机器学习公平的数学框架，Review of Mathematical frameworks for Fairness in Machine Learning

专知会员服务

38+阅读 · 2020年5月30日

【微众银行】联邦学习白皮书_v2.0，48页pdf，

专知会员服务

170+阅读 · 2020年4月26日

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

专知会员服务

59+阅读 · 2020年1月25日