Finding Peculiar Compositions of Two Frequent Strings with Background Texts

Finding Peculiar Compositions of Two Frequent Strings with Background Texts
复制标题

寻找带有背景文本的两个频繁字符串的特殊组合

DOI:
10.1007/s10115-013-0688-9
复制
发表时间:
2013
期刊:
Journal of Knowledge and Information Systems
影响因子:
--
通讯作者:
Daisuke Ikeda and Einoshin Suzuki
Daisuke Ikeda and Einoshin Suzuki
中科院分区:
--
文献类型:
--
作者:
Shimpei Yotsukura;Toshiaki Omori;Kenji Nagata and Masato Okada;Kengo Saito and Toshiharu Sugawara;Daisuke Ikeda and Einoshin Suzuki

文献摘要

相似文献

我们考虑从一组目标文本中挖掘不寻常的模式。一个典型的方法输出不寻常的模式,如果他们观察到的频率是远离他们的期望估计在一个假设的概率模型。但是,该方法难以处理零频率,从而受到数据稀疏性的影响。我们使用另一组背景文本来定义一个特殊的组合,如果和都比中更频繁,反之in.is,因为和不频繁,而与中相比是意外频繁。为了同时找到频繁子模式和不频繁模式,我们开发了一个快速算法,使用后缀树,并表明它的规模几乎是线性的参数设置下的实际。使用DNA序列的实验表明,所发现的特殊成分基本上出现在rRNA中,而通过现有方法发现的模式似乎与特定的生物学功能无关。我们还表明,发现的模式具有相似的长度,在固定长度的子串的频率分布开始歪斜。这一事实解释了为什么我们的方法可以找到长的特殊成分。
We consider mining unusual patterns from a setof target texts. A typical method outputs unusual patterns if their observed frequencies are far from their expectation estimated under an assumed probabilistic model. However, it is difficult for the method to deal with the zero frequency and thus it suffers from data sparseness. We employ another setofbackgroundtexts to define acompositionto bepeculiarif bothandare more frequent inthan inand converselyis more frequent in.is unusual becauseandare infrequent inwhileis unexpectedly frequent compared toin. To find frequent subpatterns and infrequent patterns simultaneously, we develop a fast algorithm using the suffix tree and show that it scales almost linearly under practical settings of parameters. Experiments using DNA sequences show that found peculiar compositions basically appear in rRNA while patterns found by an existing method seem not to relate to specific biological functions. We also show that discovered patterns have similar lengths at which the distribution of frequencies of fixed length substrings begins to skew. This fact explains why our method can find long peculiar compositions.