Mining Pure Patterns in Texts

Mining Pure Patterns in Texts
复制标题

挖掘文本中的纯模式

DOI:
10.1109/iiai-aai.2012.75
复制
发表时间:
2012
期刊:
Proceedings of the 2012 IIAI International Conference on Advanced Applied Informatics
影响因子:
--
通讯作者:
Kensuke Baba and Daisuke Ikeda
Kensuke Baba and Daisuke Ikeda
中科院分区:
--
文献类型:
--
作者:
Yasuhiro Yamada;Tetsuya Nakatoh;Kensuke Baba and Daisuke Ikeda

文献摘要

相似文献

我们在此研究从给定字符串作为文本中查找异常模式。在本文中,模式被表示为字符串的子字符串。关于模式频率的自然假设是模式长度越短,模式频率越大。如果模式的所有子串的频率与模式的频率相同,我们将模式定义为纯模式。这意味着子字符串仅出现在字符串的模式内。这种情况与自然假设相反。本文提出了三种用于量化模式纯度的统计量,即概率、熵和差异,它们是根据模式及其子串的频率计算的。使用 DNA 序列的实验表明,大概率的模式与序列的特征相对应。
We herein investigate finding unusual patterns from a given string as a text. In the present paper, the pattern is expressed as a sub string of the string. The natural assumption with respect to the frequency of a pattern is that the shorter the length of the pattern, the larger the frequency of the pattern. We define a pattern to be pure if the frequencies of all of the sub strings of the pattern are the same as the frequency of the pattern. This means that the sub strings appear only within the pattern in the string. This condition is in contrast to the natural assumption. The present paper proposes three statistics for quantifying the purity of a pattern, i.e., probability, entropy, and difference, which are calculated based on the frequency of the pattern and its sub strings. Experiments using DNA sequences reveal that patterns with large probability correspond to the features of the sequences.