Towards Provable Network Traffic Measurement and Analysis via Semi-Labeled Trace Datasets

Towards Provable Network Traffic Measurement and Analysis via Semi-Labeled Trace Datasets
复制标题

DOI:
10.23919/tma.2018.8506498
复制
发表时间:
2018-06
期刊:
2018 Network Traffic Measurement and Analysis Conference (TMA)
影响因子:
--
通讯作者:
Milan Cermák;Tomás Jirsík;P. Velan;Jana Komárková;Stanislav Špaček;Martin Drasar;Tomáš Plesník
Milan Cermák;Tomás Jirsík;P. Velan;Jana Komárková;Stanislav Špaček;Martin Drasar;Tomáš Plesník
中科院分区:
其他
文献类型:
--
作者:
Milan Cermák;Tomás Jirsík;P. Velan;Jana Komárková;Stanislav Špaček;Martin Drasar;Tomáš Plesník

文献摘要

被引文献

相似文献

网络流量测量和分析的研究是一个长期的领域,科学家和业界的兴趣越来越大。然而,即使过了这么多年,成果的复制、批评和审查仍然很少。我们不仅面临缺乏研究标准的问题,而且也无法获得可用于方法开发和评估的适当数据集。因此,许多潜在的高质量研究无法得到验证,也没有被行业或社区采用。本文的目的是克服这一争议的一个独特的解决方案的基础上提出的不同的方法相结合的其他研究工作。与这些研究不同,我们关注的是整个问题,涵盖数据匿名化、真实性、新近性、公开性及其在研究可证明性中的使用等所有领域。我们相信,这些挑战可以通过利用由真实世界的网络流量和注释单元组成的半标记数据集来解决,这些数据集只包含与兴趣相关的数据包跟踪。在本文中,我们概述了该方法的基本思想,从单位迹收集和半标记数据集创建其用于研究评估。我们努力使这一建议开始讨论的方法,并帮助克服今天的研究所面临的一些挑战。
Research in network traffic measurement and analysis is a long-lasting field with growing interest from both scientists and the industry. However, even after so many years, results replication, criticism, and review are still rare. We face not only a lack of research standards, but also inaccessibility of appropriate datasets that can be used for methods development and evaluation. Therefore, a lot of potentially high-quality research cannot be verified and is not adopted by the industry or the community. The aim of this paper is to overcome this controversy with a unique solution based on a combination of distinct approaches proposed by other research works. Unlike these studies, we focus on the whole issue covering all areas of data anonymization, authenticity, recency, publicity, and their usage for research provability. We believe that these challenges can be solved by utilization of semi-labeled datasets composed of real-world network traffic and annotated units with interest-related packet traces only. In this paper, we outline the basic ideas of the methodology from unit trace collection and semi-labeled dataset creation to its usage for research evaluation. We strive for this proposal to start a discussion of the approach and help to overcome some of the challenges the research faces today.