Toward Unsupervised Protocol Feature Word Extraction

Toward Unsupervised Protocol Feature Word Extraction
复制标题

走向无监督协议特征词提取

DOI:
10.1109/jsac.2014.2358857
复制
发表时间:
2014-09
影响因子:
16.4
通讯作者:
Gaogang Xie
Gaogang Xie
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhibin Zhang;Patrick P. C. Lee;Yunjie Liu;Gaogang Xie

文献摘要

参考文献

被引文献

相似文献

协议特征字是流量有效载荷中的字节子序列,可以区分应用协议,它们构成了网络管理、测量和安全系统中许多深度数据包分析规则构造的构建块。然而,如何系统、高效地从网络流量中提取协议特征词仍然是一个具有挑战性的问题。现有的方法,如基于n元语法或公共字符串(CS)的方法,只是将有效负载分成等长的片段或试图找到频繁项集,在捕获有效负载内容的隐藏统计结构方面效率低下。本文提出了一种从流量跟踪中提取协议特征词的无监督方法ProWord。ProWord建立在两个重要的算法之上。首先,我们提出了一种基于改进的投票专家算法的无监督切分算法,该算法根据信息量信息将有效载荷分解为候选词,并提供比现有的n元语法和CS方法更准确的切分。其次,我们提出了一种融合了不同类型的已知特征词检索启发式算法的排序算法,使得我们可以在候选词上建立一个有序的结构,并选择排序最高的作为协议特征词。我们通过对真实世界的交通轨迹进行评估,比较了ProWord和现有的先前方法。我们表明,ProWord能够更准确地捕获真正的协议特征词,并且执行速度显著加快。
Protocol feature words are byte subsequences within traffic payload that can distinguish application protocols, and they form the building blocks of many constructions of deep packet analysis rules in network management, measurement, and security systems. However, how to systematically and efficiently extract protocol feature words from network traffic remains a challenging issue. Existing approaches like those based on n-gram or Common String (CS), which simply breaks payload into equal-length pieces or attempts to find a frequent itemset, are ineffective in capturing the hidden statistical structure of the payload content. In this paper, we propose ProWord, an unsupervised approach that extracts protocol feature words from traffic traces. ProWord builds on two nontrivial algorithms. First, we propose an unsupervised segmentation algorithm based on the modified Voting Experts algorithm, such that we break payload into candidate words according to entropy information and provide more accurate segmentation than existing n-gram and CS approaches. Second, we propose a ranking algorithm that incorporates different types of well-known feature word retrieval heuristics, such that we can build an ordered structure on the candidate words and select the highest ranked ones as protocol feature words. We compare ProWord and existing prior approaches via evaluation on real-world traffic traces. We show that ProWord captures true protocol feature words more accurately and performs significantly faster.
DOI: --
发表时间: 2007-08
期刊: --
影响因子: --
作者:
Weidong Cui;Jayanthkumar Kannan;Helen J. Wang
通讯作者: Weidong Cui;Jayanthkumar Kannan;Helen J. Wang
DOI: 10.1145/1177080.1177123
发表时间: 2006-10
期刊: --
影响因子: --
作者:
Justin Ma;Kirill Levchenko;C. Kreibich;S. Savage;G. Voelker
通讯作者: Justin Ma;Kirill Levchenko;C. Kreibich;S. Savage;G. Voelker
DOI: 10.1109/noms.2008.4575130
发表时间: 2008-04
期刊: NOMS 2008 - 2008 IEEE Network Operations and Management Symposium
影响因子: --
作者:
Byungchul Park;Young J. Won;Myung-Sup Kim;J. W. Hong
通讯作者: Byungchul Park;Young J. Won;Myung-Sup Kim;J. W. Hong
DOI: 10.1145/1080173.1080183
发表时间: 2005-08
期刊: --
影响因子: --
作者:
P. Haffner;S. Sen;Oliver Spatscheck;Dongmei Wang
通讯作者: P. Haffner;S. Sen;Oliver Spatscheck;Dongmei Wang
DOI: --
发表时间: 2013-05
期刊: 2013 IFIP Networking Conference
影响因子: --
作者:
A. Tongaonkar;Ram Keralapura;A. Nucci
通讯作者: A. Tongaonkar;Ram Keralapura;A. Nucci