Duplicate Detection in the Reuters Collection
Duplicate Detection in the Reuters Collection
复制标题
路透社集合中的重复检测
DOI:
--
复制
发表时间:
1997
期刊:
影响因子:
--
通讯作者:
M. Sanderson
中科院分区:
文献类型:
--
作者:
M. Sanderson
While conducting some experiments with the Reuters collection, it was discovered
that contained within it were a number of documents that were exact duplicates of
each other (see Figure 1). A short study was conducted to try to discover how many
such documents there were. The results of this study revealed that the notion of a
duplicate document was not as simple as first thought.
The contents of this report are as follows. A brief review of previous duplicate detection
research will be presented, followed by a description of the methods and results of
the duplicate detection work conducted here. In addition, there is an appendix holding
the document ids of the various types of duplicate found.