We need to talk about antiviruses: challenges & pitfalls of AV evaluations

We need to talk about antiviruses: challenges & pitfalls of AV evaluations
复制标题

DOI:
10.1016/j.cose.2020.101859
复制
发表时间:
2020-08
期刊:
Comput. Secur.
影响因子:
--
通讯作者:
Marcus Botacin;Fabrício Ceschin;P. Geus;A. Grégio
Marcus Botacin;Fabrício Ceschin;P. Geus;A. Grégio
中科院分区:
其他
文献类型:
--
作者:
Marcus Botacin;Fabrício Ceschin;P. Geus;A. Grégio

文献摘要

被引文献

相似文献

安全评估是一项基本任务,用于确定运行系统中完成的保护级别,或帮助为每个特定场景选择更好的解决方案。虽然反病毒软件(AV)是大多数终端用户和公司的主要防御解决方案之一,但反病毒软件的评估是由少数组织进行的,通常仅限于比较检测率。此外,自动驾驶汽车运行模式的其他重要因素(如响应时间和检测回归)通常被低估。忽略这些因素会对自动驾驶汽车在实际场景中的有效性产生“理解差距”,我们的目标是通过对当前自动驾驶汽车的操作模式进行更广泛的描述来弥合这一差距。在我们的描述中,我们考虑了不同的文件类型、操作系统、数据集和时间框架。为此,我们每天从两个不同的、具有代表性的恶意软件来源收集样本,并连续30天将它们提交给VirusTotal (VT)服务。我们总共考虑了28,875个独特的恶意软件样本。每天,我们检索提交的样品的检测率并分配标签,结果总共有超过1M个不同的VT提交。我们的实验结果表明:(i)网络钓鱼上下文对所有自动驾驶汽车来说都是一个挑战,使恶意网页检测器不如恶意文件检测器有效;(ii)通用程序不足以确保广泛的检测覆盖率,导致特定数据集(例如,特定国家)的检出率低于全球收集样本的检出率;(iii)检测率不稳定,因为所有自动驾驶汽车在使用相同的数据集在不同的时间框架内扫描后呈现检测回归效果;(iv)自动驾驶汽车在提供新签名/启发式方面的长响应时间在我们首次识别恶意二进制文件后的前30天内创造了一个重要的攻击机会窗口。为了解决我们的研究结果的影响,我们提出了六个新的指标来评估影响自动驾驶汽车有效性的多个方面。有了它们,我们希望评估企业(和国内)用户,以更好地评估更充分满足其需求的解决方案。
Security evaluation is an essential task to identify the level of protection accomplished in running systems or to aid in choosing better solutions for each specific scenario. Although antiviruses (AVs) are one of the main defensive solutions for most end-users and corporations, AV’s evaluations are conducted by few organizations and often limited to compare detection rates. Moreover, other important factors of AVs’ operating mode (e.g., response time and detection regression) are usually underestimated. Ignoring such factors create an “understanding gap” on the effectiveness of AVs in actual scenarios, which we aim to bridge by presenting a broader characterization of current AVs’ modes of operation. In our characterization, we consider distinct file types, operating systems, datasets, and time frames. To do so, we daily collected samples from two distinct, representative malware sources and submitted them to the VirusTotal (VT) service for 30 consecutive days. In total, we considered 28,875 unique malware samples. For each day, we retrieved the submitted samples’ detection rates and assigned labels, resulting in more than 1M distinct VT submissions overall. Our experimental results show that: (i) phishing contexts are a challenge for all AVs, turning malicious Web pages detectors less effective than malicious files detectors; (ii) generic procedures are insufficient to ensure broad detection coverage, incurring in lower detection rates for particular datasets (e.g., country-specific) than for those with world-wide collected samples; (iii) detection rates are unstable since all AVs presented detection regression effects after scans in different time frames using the same dataset and (iv) AVs’ long response times in delivering new signatures/heuristics create a significant attack opportunity window within the first 30 days after we first identified a malicious binary. To address the effects of our findings, we propose six new metrics to evaluate the multiple aspects that impact the effectiveness of AVs. With them, we hope to assess corporate (and domestic) users to better evaluate the solutions that fit their needs more adequately.