Plagiarism Detection Without Reference Collections
Plagiarism Detection Without Reference Collections
复制标题
没有参考文献集的抄袭检测
DOI:
10.1007/978-3-540-70981-7_40
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
Marion Kulig
中科院分区:
文献类型:
--
作者:
Sven Meyer zu Eissen;Benno Stein;Marion Kulig
Current research in the field of automatic plagiarism detection for text documents focuses on the development of algorithms that compare suspicious documents against potential original documents. Although recent approaches perform well in identifying copied or even modified passages ([Brin et al. (1995), Stein (2005)]), they assume a closed world where a reference collection must be given (Finkel (2002)). Recall that a human reader can identify suspicious passages within a document without having a library of potential original documents in mind.This raises the question whether plagiarized passages within a document can be detected automatically if no reference is given, e. g. if the plagiarized passages stem from a book that is not available in digital form. This paper contributes right here; it proposes a method to identify potentially plagiarized passages by analyzing a single document with respect to changes in writing style. Such passages then can be used as a starting point for an Internet search for potential sources. As well as that, such passages can be preselected for inspection by a human referee. Among others, we will present new style features that can be computed efficiently and which provide highly discriminative information: Our experiments, which base on a test corpus that will be published, show encouraging results.