Evaluation of Large Language Models on Code Obfuscation (Student Abstract)

Evaluation of Large Language Models on Code Obfuscation (Student Abstract)
复制标题

DOI:
10.1609/aaai.v38i21.30517
复制
发表时间:
2024-03
期刊:
--
影响因子:
--
通讯作者:
Adrian Swindle;Derrick McNealy;Giri Krishnan;Ramyaa Ramyaa-Ramyaa
Adrian Swindle;Derrick McNealy;Giri Krishnan;Ramyaa Ramyaa-Ramyaa
中科院分区:
其他
文献类型:
--
作者:
Adrian Swindle;Derrick McNealy;Giri Krishnan;Ramyaa Ramyaa-Ramyaa

文献摘要

相似文献

混淆旨在降低代码的可解释性和代码行为的识别性。大型语言模型(LLM)已经被提出用于代码合成和代码分析。本文试图了解LLM如何分析代码和识别代码行为。具体来说,本文系统地评估了几个LLM的能力,以检测混淆代码和识别行为的各种混淆技术与不同程度的复杂性。与涉及代码插入的混淆(未使用的变量,以及用计算为这些常量的表达式替换常量的变量)相比,LLM被证明在检测改变标识符的混淆方面更好,甚至是误导性的混淆。最难检测到的是将多个简单转换分层的模糊处理。对于这些,只有20-40%的LLM的回答是正确的。添加误导性文件也成功地误导了LLM。我们在https://github.com/SwindleA/LLMCodeObfuscation上提供了所有用于复制结果的代码。总的来说,我们的结果表明LLM理解代码的能力存在差距。
Obfuscation intends to decrease interpretability of code and identification of code behavior. Large Language Models(LLMs) have been proposed for code synthesis and code analysis. This paper attempts to understand how well LLMs can analyse code and identify code behavior. Specifically, this paper systematically evaluates several LLMs’ capabilities to detect obfuscated code and identify behavior across a variety of obfuscation techniques with varying levels of complexity. LLMs proved to be better at detecting obfuscations that changed identifiers, even to misleading ones, compared to obfuscations involving code insertions (unused variables, as well as variables that replace constants with expressions that evaluate to those constants). Hardest to detect were obfuscations that layered multiple simple transformations. For these, only 20-40% of the LLMs’ responses were correct. Adding misleading documentation was also successful in misleading LLMs. We provide all our code to replicate results at https://github.com/SwindleA/LLMCodeObfuscation. Overall, our results suggest a gap in LLMs’ ability to understand code.