Duplicate Detection in the Reuters Collection

Duplicate Detection in the Reuters Collection
复制标题

路透社集合中的重复检测

DOI:
--
复制
发表时间:
1997
期刊:
--
影响因子:
--
通讯作者:
M. Sanderson
M. Sanderson
中科院分区:
--
文献类型:
--
作者:
M. Sanderson

文献摘要

被引文献

相似文献

在对路透社的藏品进行一些实验时, 其中包含了大量的文件, (见图1)。进行了一项简短的研究,试图找出有多少 有这样的文件。这项研究的结果表明, 重复文件并不像最初想象的那么简单。 本报告的内容如下。重复检测技术研究进展 研究将提出,其次是方法和结果的描述, 在这里进行的重复检测工作。此外,还有一个附录, 发现的各种类型重复的文档ID。
While conducting some experiments with the Reuters collection, it was discovered that contained within it were a number of documents that were exact duplicates of each other (see Figure 1). A short study was conducted to try to discover how many such documents there were. The results of this study revealed that the notion of a duplicate document was not as simple as first thought. The contents of this report are as follows. A brief review of previous duplicate detection research will be presented, followed by a description of the methods and results of the duplicate detection work conducted here. In addition, there is an appendix holding the document ids of the various types of duplicate found.