Detection of Peculiar Examples using LOF and One Class SVM

Detection of Peculiar Examples using LOF and One Class SVM
复制标题

DOI:
--
复制
发表时间:
2010-05
期刊:
--
影响因子:
--
通讯作者:
Hiroyuki Shinnou;Minoru Sasaki
Hiroyuki Shinnou;Minoru Sasaki
中科院分区:
其他
文献类型:
--
作者:
Hiroyuki Shinnou;Minoru Sasaki

文献摘要

相似文献

本文提出了一种从语料库中检测目标词特殊实例的方法。本文把以下几个例子作为特例:(1)例子中的目标词的意义是新的,(2)例子中的目标词构成的复合词是新的或很有技术性的。特殊的例子被认为是一个离群值在给定的例子集。因此,我们可以将数据挖掘领域提出的许多方法应用到我们的任务中。本文将数据挖掘领域中具有代表性的离群点检测方法--基于密度的方法、局部离群点因子(LOF)方法和单类支持向量机方法进行联合收割机。实验中,我们以《北京世界新闻》白皮书文本为语料,以10个名词性词语为目标词。该方法提高了LOF和单类SVM的查准率和查全率。我们证明了我们的方法可以检测到新的含义,通过使用名词'midori(绿色树'。两个样本的相似性度量不足是导致漏检和误检的主要原因。今后,我们必须改进它。
This paper proposes the method to detect peculiar examples of the target word from a corpus. In this paper we regard following examples as peculiar examples: (1) a meaning of the target word in the example is new, (2) a compound word consisting of the target word in the example is new or very technical. The peculiar example is regarded as an outlier in the given example set. Therefore we can apply many methods proposed in the data mining domain to our task. In this paper, we propose the method to combine the density based method, Local Outlier Factor (LOF), and One Class SVM, which are representative outlier detection methods in the data mining domain. In the experiment, we use the Whitepaper text in BCCWJ as the corpus, and 10 noun words as target words. Our method improved precision and recall of LOF and One Class SVM. And we show that our method can detect new meanings by using the noun `midori (green)'. The main reason of un-detections and wrong detection is that similarity measure of two examples is inadequacy. In future, we must improve it.