jtrans: 二进制代码相似性的跳跃软件变换器 (jTrans: Jump-Aware Transformer for Binary Code Similarity) - 专知论文

会员服务 ·

0

binary · 代码 · 相似度 · SOTA · 变换 ·

2022 年 5 月 25 日

jTrans: Jump-Aware Transformer for Binary Code Similarity

翻译：jtrans: 二进制代码相似性的跳跃软件变换器

Hao Wang,Wenjie Qu,Gilad Katz,Wenyu Zhu,Zeyu Gao,Han Qiu,Jianwei Zhuge,Chao Zhang

from arxiv, In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2022

Binary code similarity detection (BCSD) has important applications in various fields such as vulnerability detection, software component analysis, and reverse engineering. Recent studies have shown that deep neural networks (DNNs) can comprehend instructions or control-flow graphs (CFG) of binary code and support BCSD. In this study, we propose a novel Transformer-based approach, namely jTrans, to learn representations of binary code. It is the first solution that embeds control flow information of binary code into Transformer-based language models, by using a novel jump-aware representation of the analyzed binaries and a newly-designed pre-training task. Additionally, we release to the community a newly-created large dataset of binaries, BinaryCorp, which is the most diverse to date. Evaluation results show that jTrans outperforms state-of-the-art (SOTA) approaches on this more challenging dataset by 30.5% (i.e., from 32.0% to 62.5%). In a real-world task of known vulnerability searching, jTrans achieves a recall that is 2X higher than existing SOTA baselines.

翻译：二进制代码检测(BCSD)在脆弱性检测、软件元件分析和反向工程等各个领域都有重要的应用。最近的研究显示,深神经网络(DNNs)能够理解二进制代码的指示或控制流图(CFG)并支持BCSD。在这个研究中,我们提出了一种新的基于变异器的方法,即jTrans,以学习二进制代码的表达方式。这是第一个将二进制代码的控制流信息嵌入基于变异器的语言模型的解决办法,方法是利用分析的二进制和新设计的预培训任务的新跳入觉。此外,我们向社区发放了新创建的二进制二进制(Binary Corp)的大型数据集,这是迄今为止最多样化的。评估结果表明,在这种更具挑战性的数据中,30.5%(即从32.0%到62.5%)的基变异语言模型中,将控制流信息嵌入到变异化器语言模型中。在现实世界已知的脆弱性搜索工作中,jTransforms 实现的回收量比现有的SOTA基线高出2X。

0

相关内容

binary

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

对偶Auslander转置及其诱导模类的同调性质研究

国家自然科学基金

0+阅读 · 2015年12月31日

β2-AR/PKA通路在内皮祖细胞修复急性肾损伤中的作用及机制

国家自然科学基金

0+阅读 · 2013年12月31日

多项式代数上自同构的结构研究

国家自然科学基金

0+阅读 · 2013年12月31日

不同途径移植HUCB-MSCs治疗脑血管病大鼠microPET-CT评价及其治疗机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

顺序激活RARα及TrkB受体介导信号的协同机制对成体急性脊髓损伤神经再生修复的实验研究

国家自然科学基金

0+阅读 · 2012年12月31日

阵列天线3D-SAR的DEM生成技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

Eulerian bond-cubic 模型渗流性质的数值研究

国家自然科学基金

0+阅读 · 2012年12月31日

有机分子半导体的非局域电声子耦合：声子色散与二阶电声子相互作用的影响

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

多物理场小周期结构体双尺度有限元分析

国家自然科学基金

0+阅读 · 2008年12月31日

Modelling Evolutionary and Stationary User Preferences for Temporal Sets Prediction

Arxiv

0+阅读 · 2022年7月13日

AVA-AVD: Audio-Visual Speaker Diarization in the Wild

Arxiv

0+阅读 · 2022年7月13日

EAGAN: Efficient Two-stage Evolutionary Architecture Search for GANs

Arxiv

0+阅读 · 2022年7月12日

Are We Building on the Rock? On the Importance of Data Preprocessing for Code Summarization

Are We Building on the Rock? On the Importance of Data Preprocessing for Code Summarization

Arxiv

0+阅读 · 2022年7月12日

Language-specific Characteristic Assistance for Code-switching Speech Recognition

Arxiv

0+阅读 · 2022年7月12日

Improved Soft-aided Decoding of Product Codes with Dynamic Reliability Scores

Arxiv

0+阅读 · 2022年7月11日

Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

Arxiv

0+阅读 · 2022年7月11日

Few-shot training LLMs for project-specific code-summarization

Arxiv

0+阅读 · 2022年7月9日

Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder

Arxiv

0+阅读 · 2022年7月9日

Time-Series Event Prediction with Evolutionary State Graph

Arxiv

14+阅读 · 2020年11月25日

VIP会员

文章信息

相关主题

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津博士论文】零样本强化学习综述

《美军条令：陆军指挥官与规划人员地理空间指南》60页

战术边缘指挥控制：防务面临的核心挑战

迈向开放世界检测：综述

相关资讯

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

Modelling Evolutionary and Stationary User Preferences for Temporal Sets Prediction

Arxiv

0+阅读 · 2022年7月13日

AVA-AVD: Audio-Visual Speaker Diarization in the Wild

Arxiv

0+阅读 · 2022年7月13日

EAGAN: Efficient Two-stage Evolutionary Architecture Search for GANs

Arxiv

0+阅读 · 2022年7月12日

Are We Building on the Rock? On the Importance of Data Preprocessing for Code Summarization

Are We Building on the Rock? On the Importance of Data Preprocessing for Code Summarization

Arxiv

0+阅读 · 2022年7月12日

Language-specific Characteristic Assistance for Code-switching Speech Recognition

Arxiv

0+阅读 · 2022年7月12日

Improved Soft-aided Decoding of Product Codes with Dynamic Reliability Scores

Arxiv

0+阅读 · 2022年7月11日

Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

Arxiv

0+阅读 · 2022年7月11日

Few-shot training LLMs for project-specific code-summarization

Arxiv

0+阅读 · 2022年7月9日

Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder

Arxiv

0+阅读 · 2022年7月9日

Time-Series Event Prediction with Evolutionary State Graph

Arxiv

14+阅读 · 2020年11月25日

相关基金

对偶Auslander转置及其诱导模类的同调性质研究

国家自然科学基金

0+阅读 · 2015年12月31日

β2-AR/PKA通路在内皮祖细胞修复急性肾损伤中的作用及机制

国家自然科学基金

0+阅读 · 2013年12月31日

多项式代数上自同构的结构研究

国家自然科学基金

0+阅读 · 2013年12月31日

不同途径移植HUCB-MSCs治疗脑血管病大鼠microPET-CT评价及其治疗机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

顺序激活RARα及TrkB受体介导信号的协同机制对成体急性脊髓损伤神经再生修复的实验研究

国家自然科学基金

0+阅读 · 2012年12月31日

阵列天线3D-SAR的DEM生成技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

Eulerian bond-cubic 模型渗流性质的数值研究

国家自然科学基金

0+阅读 · 2012年12月31日

有机分子半导体的非局域电声子耦合：声子色散与二阶电声子相互作用的影响

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

多物理场小周期结构体双尺度有限元分析

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员