Multilingual Plagiarism Detection
Multilingual Plagiarism Detection
复制标题
多语言抄袭检测
DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
L. Agarwal
中科院分区:
文献类型:
--
作者:
Menna Mostafa;L. Agarwal
Cross lingual plagiarism detection has recently caught attention due to copy-right violations occurring in many fields such as education, journalism, scientific research, literature, screenplays, etc, where an author would translate an article in language L1 into language L2 and then either publish/submit it or change some of the sentences to suit his/her motivations. Therefore, the need for a robust method for cross lingual plagiarism detection arises. Most of the existing work on cross lingual plagiarism detection uses machine translation to translate the suspect document in L1 into L2 and then search for similar documents in L2. However, we argue that this approach suffers from the following limitations: (1) machine translation does not capture different writing styles that differ from field to another, (2) online machine translation that allows anonymous users to suggest better translations which suffers from tampering and incorrect suggestions, (3) the limited ability to identify different types of plagiarisms, for example two articles describing an accident might be labeled as plagiarized although they originated from different sources. Therefore, we propose an approach that will attempt to remedy the above three limitations by using machine learning and crowd sourcing techniques.