Fast Plagiarism Detection Based on Simple Document Similarity

Fast Plagiarism Detection Based on Simple Document Similarity
复制标题

基于简单文档相似度的快速抄袭检测

DOI:
10.1109/icdim.2017.8244662
复制
发表时间:
2017
期刊:
Proc. the Twelfth International Conference on Digital Information Management
影响因子:
--
通讯作者:
Kensuke Baba
Kensuke Baba
中科院分区:
--
文献类型:
--
作者:
Arief MAULANA;Kazumi SAITO;Tetsuo IKEDA;Hiroaki YUZE;Kensuke Baba

文献摘要

相似文献

大量文档中的剽窃检测需要有效的方法。针对“复制粘贴”式剽窃行为,提出了一种基于近似字符串匹配的剽窃检测算法,并对算法的实现进行了速度改进。该算法通过两种近似方法省略了大部分计算量,但近似方法对检测精度的影响是可以接受的。通过对一组数据进行实验,评估了改进算法在处理时间和精度上的效果。实验结果表明,改进后的算法在准确率降低6.4%的情况下,处理时间减少到原来的1/20左右。
Plagiarism detection in a large number of documents requires efficient methods. This paper proposes a plagiarism detection algorithm based on approximate string matching to be specified in “copy and paste”-type plagiarisms, and a speed improvement to an implementation of the algorithm. Most of the computations required in the algorithm are omitted by two kinds of approximations of the output used for plagiarism detection, while the decrease of accuracy caused by the approximations is acceptable. The effect of the improvement on the processing time and accuracy of the algorithm is evaluated by conducting experiments with a data set. The experimental results show that the improvement can reduce the processing time to approximately one-twentieth for a 6.4% decrease of the accuracy from those for the normal implementation of the algorithm.