Malicious and Benign Webpages Dataset.

Malicious and Benign Webpages Dataset.
复制标题

DOI:
10.1016/j.dib.2020.106304
复制
发表时间:
2020-10
期刊:
影响因子:
1.2
通讯作者:
Singh AK
Singh AK
中科院分区:
其他
文献类型:
--
作者:
Singh AK

文献摘要

被引文献

相似文献

网络安全是一项具有挑战性的任务,在互联网上不断上升的威胁。随着互联网上数十亿个网站的活跃,以及黑客不断发展新的技术来诱捕网络用户,机器学习提供了有前途的技术来检测恶意网站。本文中描述的数据集旨在用于基于机器学习的恶意和良性网页分析。这些数据是使用专门的网络爬虫MalCrawler从互联网上收集的。数据集包括各种提取的属性,以及包括JavaScript代码的原始网页内容。它支持监督学习和无监督学习。对于监督学习,恶意和良性网页的类标签已使用Google安全浏览API添加到数据集中。1范围内最相关的属性已被提取并包含在此数据集中。然而,如果需要的话,包括包含在该数据集中的JavaScript代码的原始web内容支持进一步的属性提取。此外,这些原始内容和代码可以用作基于文本的分析的非结构化数据输入。该数据集由大约150万个网页的数据组成,这使得它适合深度学习算法。本文还提供了用于数据提取及其分析的代码片段。
Web Security is a challenging task amidst ever rising threats on the Internet. With billions of websites active on Internet, and hackers evolving newer techniques to trap web users, machine learning offers promising techniques to detect malicious websites. The dataset described in this manuscript is meant for such machine learning based analysis of malicious and benign webpages. The data has been collected from Internet using a specialized focused web crawler named MalCrawler. The dataset comprises of various extracted attributes, and also raw webpage content including JavaScript code. It supports both supervised and unsupervised learning. For supervised learning, class labels for malicious and benign webpages have been added to the dataset using the Google Safe Browsing API.1 The most relevant attributes within the scope have already been extracted and included in this dataset. However, the raw web content, including JavaScript code included in this dataset supports further attribute extraction, if so desired. Also, this raw content and code can be used as unstructured data input for text-based analytics. This dataset consists of data from approximately 1.5 million webpages, which makes it suitable for deep learning algorithms. This article also provides code snippets used for data extraction and its analysis.