综合研究以学习为基础的PE 恶意恶意家庭分类方法 (A Comprehensive Study on Learning-Based PE Malware Family Classification Methods)

Driven by the high profit, Portable Executable (PE) malware has been consistently evolving in terms of both volume and sophistication. PE malware family classification has gained great attention and a large number of approaches have been proposed. With the rapid development of machine learning techniques and the exciting results they achieved on various tasks, machine learning algorithms have also gained popularity in the PE malware family classification task. Three mainstream approaches that use learning based algorithms, as categorized by the input format the methods take, are image-based, binary-based and disassembly-based approaches. Although a large number of approaches are published, there is no consistent comparisons on those approaches, especially from the practical industry adoption perspective. Moreover, there is no comparison in the scenario of concept drift, which is a fact for the malware classification task due to the fast evolving nature of malware. In this work, we conduct a thorough empirical study on learning-based PE malware classification approaches on 4 different datasets and consistent experiment settings. Based on the experiment results and an interview with our industry partners, we find that (1) there is no individual class of methods that significantly outperforms the others; (2) All classes of methods show performance degradation on concept drift (by an average F1-score of 32.23%); and (3) the prediction time and high memory consumption hinder existing approaches from being adopted for industry usage.

翻译：在高利润驱动下,可移植执行软件恶意软件在数量和复杂程度方面不断演变。PE恶意软件家庭分类引起了极大关注,并提出了大量的办法。随着机器学习技术的迅速发展及其在各种任务上取得的令人兴奋的结果,机器学习算法在PE恶意软件家庭分类任务中也越来越受欢迎。三种使用学习算法的主流方法,按所用输入格式分类,是基于图像、基于二进制和基于拆解的方法。虽然公布了大量的方法,但没有对这些方法进行一致的比较,特别是从实际行业采用的角度来看。此外,由于机器学习技术的迅速发展及其在各种任务上取得的令人振奋人心的结果,机器学习算法在PE恶意软件的家庭分类任务中也越来越受欢迎。在这项工作中,我们根据4种不同输入格式分类法和一致的实验环境,对基于学习的PE软件软件分类方法进行了彻底的经验研究。根据实验结果和与我们行业伙伴的访谈,我们发现(1) 没有一种明显超出现有消费使用率的个别方法,尤其是从实际行业采用的方法;(2) 所有流化和流化方法都显示现有流化方法。

相关内容

Machine Learning

关注 2245

机器学习（Machine Learning）是一个研究计算学习方法的国际论坛。该杂志发表文章，报告广泛的学习方法应用于各种学习问题的实质性结果。该杂志的特色论文描述研究的问题和方法，应用研究和研究方法的问题。有关学习问题或方法的论文通过实证研究、理论分析或与心理现象的比较提供了坚实的支持。应用论文展示了如何应用学习方法来解决重要的应用问题。研究方法论文改进了机器学习的研究方法。所有的论文都以其他研究人员可以验证或复制的方式描述了支持证据。论文还详细说明了学习的组成部分，并讨论了关于知识表示和性能任务的假设。官网地址：http://dblp.uni-trier.de/db/journals/ml/

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

【基于模型的强化学习的博弈论框架】A Game Theoretic Framework for Model Based Reinforcement Learning

专知会员服务

131+阅读 · 2020年4月19日