Finding Peculiar Compositions of Two Frequent Strings with Background Texts
Finding Peculiar Compositions of Two Frequent Strings with Background Texts
复制标题
寻找带有背景文本的两个频繁字符串的特殊组合
DOI:
10.1007/s10115-013-0688-9
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Daisuke Ikeda and Einoshin Suzuki
中科院分区:
文献类型:
--
作者:
Shimpei Yotsukura;Toshiaki Omori;Kenji Nagata and Masato Okada;Kengo Saito and Toshiharu Sugawara;Daisuke Ikeda and Einoshin Suzuki
We consider mining unusual patterns from a setof target texts. A typical method outputs unusual patterns if their observed frequencies are far from their expectation estimated under an assumed probabilistic model. However, it is difficult for the method to deal with the zero frequency and thus it suffers from data sparseness. We employ another setofbackgroundtexts to define acompositionto bepeculiarif bothandare more frequent inthan inand converselyis more frequent in.is unusual becauseandare infrequent inwhileis unexpectedly frequent compared toin. To find frequent subpatterns and infrequent patterns simultaneously, we develop a fast algorithm using the suffix tree and show that it scales almost linearly under practical settings of parameters. Experiments using DNA sequences show that found peculiar compositions basically appear in rRNA while patterns found by an existing method seem not to relate to specific biological functions. We also show that discovered patterns have similar lengths at which the distribution of frequencies of fixed length substrings begins to skew. This fact explains why our method can find long peculiar compositions.