Beyond tf-idf and Cosine Distance in Documents Dissimilarity Measure

Beyond tf-idf and Cosine Distance in Documents Dissimilarity Measure
复制标题

DOI:
10.1007/978-3-319-28940-3_33
复制
发表时间:
2015-12
期刊:
--
影响因子:
--
通讯作者:
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio
中科院分区:
其他
文献类型:
--
作者:
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio

文献摘要

相似文献

在向量空间模型中,采用不同类型的词项加权方案来调整词袋文档向量,以提高应用最广泛的余弦距离的性能。即使余弦距离与一些术语加权方案在某些数据集中产生更可靠的(不)相似性度量,但由于术语加权方案的基本假设,它在其他数据集中可能表现不佳。在本文中,我们认为,明确调整的词袋文档向量使用长期加权是不需要的,如果一个依赖于数据的相异度测量称为相异度。我们在文档检索任务中的实证结果表明,在四个广泛使用的基准文档集合中,最简单的二进制词袋表示比最先进的术语加权方案的余弦距离更好或具有竞争力。
In vector space model, different types of term weighting schemes are used to adjust bag-of-words document vectors in order to improve the performance of the most widely used cosine distance. Even though the cosine distance with some term weighting schemes result in more reliable (dis)similarity measure in some data sets, it may not perform well in others because of the underlying assumptions of the term weighting schemes. In this paper, we argue that the explicit adjustment of bag-of-words document vectors using term weighting is not required if a data-dependent dissimilarity measure called-dissimilarity is used. Our empirical result in document retrieval task reveals thatwith the simplest binary bag-of-words representation is either better or competitive to the cosine distance with the best performing state-of-the-art term weighting scheme in four widely used benchmark document collections.