DIRE: A Neural Approach to Decompiled Identifier Naming

DIRE: A Neural Approach to Decompiled Identifier Naming
复制标题

DOI:
10.1109/ase.2019.00064
复制
发表时间:
2019-09
期刊:
2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Jeremy Lacomis;Pengcheng Yin;Edward J. Schwartz;Miltiadis Allamanis;Claire Le Goues;Graham Neubig;Bogdan Vasilescu
Jeremy Lacomis;Pengcheng Yin;Edward J. Schwartz;Miltiadis Allamanis;Claire Le Goues;Graham Neubig;Bogdan Vasilescu
中科院分区:
其他
文献类型:
--
作者:
Jeremy Lacomis;Pengcheng Yin;Edward J. Schwartz;Miltiadis Allamanis;Claire Le Goues;Graham Neubig;Bogdan Vasilescu

文献摘要

被引文献

相似文献

分解器是检查二进制文件的最常见工具之一,而无需相应的源代码。它将二进制文件转换为高级代码,从而扭转了编译过程。分解器可以重建在编译过程中丢失的许多信息(例如结构和类型信息)。不幸的是,它们没有重建具有语义上有意义的变量名称,这些变量名称已知会增加代码可理解性。我们提出了分解标识符重命名引擎(DIRE),这是一种用于可变名称恢复的新型概率技术,使用了分解器恢复的词汇和结构信息。我们还提出了一种用于生成适合培训和评估倒编码重命名模型的语料库的技术,我们用来创建从Github开采的C项目生成的164,632个唯一X86-64二进制文件的语料库。我们的结果表明,在这个语料库上,可预测的可变名称与原始源代码中的名称相同,最高74.3%。
The decompiler is one of the most common tools for examining binaries without corresponding source code. It transforms binaries into high-level code, reversing the compilation process. Decompilers can reconstruct much of the information that is lost during the compilation process (e.g., structure and type information). Unfortunately, they do not reconstruct semantically meaningful variable names, which are known to increase code understandability. We propose the Decompiled Identifier Renaming Engine (DIRE), a novel probabilistic technique for variable name recovery that uses both lexical and structural information recovered by the decompiler. We also present a technique for generating corpora suitable for training and evaluating models of decompiled code renaming, which we use to create a corpus of 164,632 unique x86-64 binaries generated from C projects mined from GitHub. Our results show that on this corpus DIRE can predict variable names identical to the names in the original source code up to 74.3% of the time.