Mis-shapes, Mistakes, Misfits: An Analysis of Domain Classification Services

Mis-shapes, Mistakes, Misfits: An Analysis of Domain Classification Services
复制标题

DOI:
10.1145/3419394.3423660
复制
发表时间:
2020-10
期刊:
Proceedings of the ACM Internet Measurement Conference
影响因子:
--
通讯作者:
Pelayo Vallina;Victor Le Pochat;Álvaro Feal;Marius Paraschiv;Julien Gamba;Tim Burke;O. Hohlfeld;J. Tapiador;Narseo Vallina-Rodriguez
Pelayo Vallina;Victor Le Pochat;Álvaro Feal;Marius Paraschiv;Julien Gamba;Tim Burke;O. Hohlfeld;J. Tapiador;Narseo Vallina-Rodriguez
中科院分区:
其他
文献类型:
--
作者:
Pelayo Vallina;Victor Le Pochat;Álvaro Feal;Marius Paraschiv;Julien Gamba;Tim Burke;O. Hohlfeld;J. Tapiador;Narseo Vallina-Rodriguez

文献摘要

被引文献

相似文献

域分类服务在多个领域都有应用,包括网络安全、内容拦截和定向广告。然而,这些服务在分类领域的方法方面往往是一个黑盒子,这使得很难评估它们的优势、对特定应用程序的适用性和局限性。在这项工作中,我们对超过4.4万个主机名上的13个流行域名分类服务进行了大规模分析。我们的研究从经验上探讨了它们的方法、可扩展性限制、标签星座,以及它们对学术研究和其他实际应用(如内容过滤)的适用性。我们发现,各个医疗机构的覆盖率差异很大,从90%以上到1%以下不等。所有服务都偏离了它们的文档分类法,妨碍了对研究的合理使用。此外,提供商之间的标签高度不一致,他们在域上几乎没有一致,这使得比较或组合这些服务变得困难。我们还展示了众包工作的动态如何受到可扩展性和覆盖方面的阻碍,以及人类标注者之间的主观分歧。最后,通过案例研究,我们展示了大多数服务不适合检测用于研究或内容屏蔽目的的专门内容。最后,我们根据我们的经验见解和经验,对它们的使用提出了可行的建议。我们特别关注用户应该如何处理在技术解决方案和研究中观察到的不同服务之间的显著差异。
Domain classification services have applications in multiple areas, including cybersecurity, content blocking, and targeted advertising. Yet, these services are often a black box in terms of their methodology to classifying domains, which makes it difficult to assess their strengths, aptness for specific applications, and limitations. In this work, we perform a large-scale analysis of 13 popular domain classification services on more than 4.4M hostnames. Our study empirically explores their methodologies, scalability limitations, label constellations, and their suitability to academic research as well as other practical applications such as content filtering. We find that the coverage varies enormously across providers, ranging from over 90% to below 1%. All services deviate from their documented taxonomy, hampering sound usage for research. Further, labels are highly inconsistent across providers, who show little agreement over domains, making it difficult to compare or combine these services. We also show how the dynamics of crowd-sourced efforts may be obstructed by scalability and coverage aspects as well as subjective disagreements among human labelers. Finally, through case studies, we showcase that most services are not fit for detecting specialized content for research or content-blocking purposes. We conclude with actionable recommendations on their usage based on our empirical insights and experience. Particularly, we focus on how users should handle the significant disparities observed across services both in technical solutions and in research.