Mining Pure Patterns in Texts
Mining Pure Patterns in Texts
复制标题
挖掘文本中的纯模式
DOI:
10.1109/iiai-aai.2012.75
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Kensuke Baba and Daisuke Ikeda
中科院分区:
文献类型:
--
作者:
Yasuhiro Yamada;Tetsuya Nakatoh;Kensuke Baba and Daisuke Ikeda
We herein investigate finding unusual patterns from a given string as a text. In the present paper, the pattern is expressed as a sub string of the string. The natural assumption with respect to the frequency of a pattern is that the shorter the length of the pattern, the larger the frequency of the pattern. We define a pattern to be pure if the frequencies of all of the sub strings of the pattern are the same as the frequency of the pattern. This means that the sub strings appear only within the pattern in the string. This condition is in contrast to the natural assumption. The present paper proposes three statistics for quantifying the purity of a pattern, i.e., probability, entropy, and difference, which are calculated based on the frequency of the pattern and its sub strings. Experiments using DNA sequences reveal that patterns with large probability correspond to the features of the sequences.