A Close Look at a Daily Dataset of Malware Samples

A Close Look at a Daily Dataset of Malware Samples
复制标题

DOI:
10.1145/3291061
复制
发表时间:
2019-01-01
影响因子:
2.3
通讯作者:
Balzarotti, Davide
Balzarotti, Davide
中科院分区:
计算机科学4区
文献类型:
--
作者:
Ugarte-Pedrero, Xabier;Graziano, Mariano;Balzarotti, Davide

文献摘要

被引文献

相似文献

独特恶意软件样本的数量正在失去控制。多年来,安全公司设计并部署了复杂的基础设施来收集和分析大量的样本。因此,一家安全公司每天仅从不同的信息源中就能收集100多万份独特文件。这些信息被自动存储和处理,以从静态和动态分析中提取可操作的信息。然而,这些数据中只有一小部分是安全研究人员感兴趣的,并吸引了人类专家的兴趣。据我们所知,没有人系统地剖析过这些数据集,以准确地理解它们真正包含的内容。由于所谓的无趣样本普遍存在,安全社区通常抛弃了这个问题。在本文中,我们将引导读者逐步分析一天内从这些提要收集的数十万个Windows可执行文件。我们的目标是展示一家公司如何使用现有的最先进的技术来自动处理这些样本,然后执行手动实验来理解和记录这个庞大数据集的真实内容。我们介绍了过滤步骤,并详细讨论了如何根据样本的行为将它们分组在一起以支持手动验证。最后,我们使用这个测量实验的结果来提供一个粗略的估计,即需要的人力和计算机资源,以达到一天的捕获的底部。
The number of unique malware samples is growing out of control. Over the years, security companies have designed and deployed complex infrastructures to collect and analyze this overwhelming number of samples. As a result, a security company can collect more than 1M unique files per day only from its different feeds. These are automatically stored and processed to extract actionable information derived from static and dynamic analysis. However, only a tiny amount of this data is interesting for security researchers and attracts the interest of a human expert.To the best of our knowledge, nobody has systematically dissected these datasets to precisely understand what they really contain. The security community generally discards the problem because of the alleged prevalence of uninteresting samples.In this article, we guide the reader through a step-by-step analysis of the hundreds of thousands Windows executables collected in one day from these feeds. Our goal is to show how a company can employ existing state-of-the-art techniques to automatically process these samples and then perform manual experiments to understand and document what is the real content of this gigantic dataset. We present the filtering steps, and we discuss in detail how samples can be grouped together according to their behavior to support manual verification. Finally, we use the results of this measurement experiment to provide a rough estimate of both the human and computer resources that are required to get to the bottom of the catch of the day.