Generating Fake Cyber Threat Intelligence Using Transformer-Based Models

Generating Fake Cyber Threat Intelligence Using Transformer-Based Models
复制标题

DOI:
10.1109/ijcnn52387.2021.9534192
复制
发表时间:
2021-02
期刊:
2021 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
P. Ranade;Aritran Piplai;Sudip Mittal;A. Joshi;Tim Finin-
P. Ranade;Aritran Piplai;Sudip Mittal;A. Joshi;Tim Finin-
中科院分区:
其他
文献类型:
--
作者:
P. Ranade;Aritran Piplai;Sudip Mittal;A. Joshi;Tim Finin-

文献摘要

被引文献

相似文献

网络防御系统正在开发中,以自动摄取包含半结构化数据和/或文本的网络威胁情报(CTI),以填充知识图谱。潜在的风险是,伪造的CTI可以通过开源情报(OSINT)社区或Web生成和传播,从而对这些系统进行数据中毒攻击。攻击者可以使用虚假的CTI示例作为训练输入来破坏网络防御系统,迫使他们的模型学习错误的输入来满足攻击者的恶意需求。在本文中,我们展示了如何使用变压器自动生成假CTI文本描述。给定一个初始提示句子,像GPT-2这样经过微调的公共语言模型可以生成可信的CTI文本,从而误导网络防御系统。我们使用生成的假CTI文本对网络安全知识图(CKG)和网络安全语料库执行数据中毒攻击。攻击带来了不利影响,例如返回错误的推理输出,表示中毒以及其他依赖的基于人工智能的网络防御系统的破坏。我们使用传统方法进行评估,并与网络安全专业人员和威胁猎人一起进行人类评估研究。根据这项研究,专业的威胁猎人同样有可能认为我们生成的虚假CTI和真实CTI是真实的。
Cyber-defense systems are being developed to automatically ingest Cyber Threat Intelligence (CTI) that contains semi-structured data and/or text to populate knowledge graphs. A potential risk is that fake CTI can be generated and spread through Open-Source Intelligence (OSINT) communities or on the Web to effect a data poisoning attack on these systems. Adversaries can use fake CTI examples as training input to subvert cyber defense systems, forcing their models to learn incorrect inputs to serve the attackers' malicious needs. In this paper, we show how to automatically generate fake CTI text descriptions using transformers. Given an initial prompt sentence, a public language model like GPT-2 with fine-tuning can generate plausible CTI text that can mislead cyber-defense systems. We use the generated fake CTI text to perform a data poisoning attack on a Cybersecurity Knowledge Graph (CKG) and a cybersecurity corpus. The attack introduced adverse impacts such as returning incorrect reasoning outputs, representation poisoning, and corruption of other dependent AI-based cyber defense systems. We evaluate with traditional approaches and conduct a human evaluation study with cyber-security professionals and threat hunters. Based on the study, professional threat hunters were equally likely to consider our fake generated CTI and authentic CTI as true.