CyBERT: Contextualized Embeddings for the Cybersecurity Domain

CyBERT: Contextualized Embeddings for the Cybersecurity Domain
复制标题

DOI:
10.1109/bigdata52589.2021.9671824
复制
发表时间:
2021-12
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
P. Ranade;Aritran Piplai;A. Joshi;Timothy W. Finin
P. Ranade;Aritran Piplai;A. Joshi;Timothy W. Finin
中科院分区:
其他
文献类型:
--
作者:
P. Ranade;Aritran Piplai;A. Joshi;Timothy W. Finin

文献摘要

被引文献

相似文献

我们提出了CyBERT,一种来自Transformers的领域特定的双向编码器表示(BERT)模型,使用大量文本网络安全数据进行了微调。最先进的自然语言模型可以处理密集、细粒度的文本威胁、攻击和漏洞信息,可以为网络安全社区带来许多好处。本文的主要贡献是为安全社区提供了一个初始的微调的BERT模型,该模型可以高精度地执行各种网络安全特定的下游任务,并有效地利用资源。我们从开源的非结构化和半非结构化的网络威胁情报(CTI)数据创建网络安全语料库,并使用它通过掩蔽语言建模(MLM)微调基本BERT模型,以识别专门的网络安全实体。我们使用可使现代安全运营中心(SOC)受益的各种下游任务来评估该模型。在特定领域的传销评估中,微调的CyBERT模型优于基本的BERT模型。我们还提供了CyBERT在基于网络安全的下游任务中的应用用例。
We present CyBERT, a domain-specific Bidirectional Encoder Representations from Transformers (BERT) model, fine-tuned with a large corpus of textual cybersecurity data. State-of-the-art natural language models that can process dense, fine-grained textual threat, attack, and vulnerability information can provide numerous benefits to the cybersecurity community. The primary contribution of this paper is providing the security community with an initial fine-tuned BERT model that can perform a variety of cybersecurity-specific downstream tasks with high accuracy and efficient use of resources. We create a cybersecurity corpus from open-source unstructured and semi-unstructured Cyber Threat Intelligence (CTI) data and use it to fine-tune a base BERT model with Masked Language Modeling (MLM) to recognize specialized cybersecurity entities. We evaluate the model using various downstream tasks that can benefit modern Security Operations Centers (SOCs). The fine-tuned CyBERT model outperforms the base BERT model in the domain-specific MLM evaluation. We also provide use-cases of CyBERT application in cybersecurity based downstream tasks.