Authorship Attribution Using Small Sets of Frequent Part-of-Speech Skip-grams

Authorship Attribution Using Small Sets of Frequent Part-of-Speech Skip-grams
复制标题

DOI:
--
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Yao Jean Marc Pokou;Philippe Fournier-Viger;C. Moghrabi
Yao Jean Marc Pokou;Philippe Fournier-Viger;C. Moghrabi
中科院分区:
其他
文献类型:
--
作者:
Yao Jean Marc Pokou;Philippe Fournier-Viger;C. Moghrabi

文献摘要

被引文献

相似文献

计算机支持的作者归属提供了用于提取风格特征的工具,这些特征可以帮助验证或识别文本文档的作者。在许多情况下,找到文档的作者是非常重要的,例如在刑事调查期间检测剽窃以保护版权和法医支持。本文旨在探讨小说的文体特征,以准确地刻画作者的作品。特别是,使用部分的语音跳跃的克和内部的前k顺序模式挖掘算法被认为是作者归属的任务。一项研究使用了由10位作者撰写的30篇文本,包括2,615,856个单词和99,903个句子,证实了文本中的词性跳跃图的挖掘有助于作者身份的推断。
Computer-supported authorship attribution provides tools for extracting stylistic features that can help verify or identify the author of text documents. In many situations finding the author of a document is very important, such as the detection of plagiarism for protecting copy-rights and forensic support during criminal investigations. Thispaper, thus explores a novel stylistic feature with the aim of accurately characterizing an author’s work. In particular, the use of part-of-speech skip-grams and an in-house top-k sequential pattern mining algorithm are considered for the task of authorship attribution. A study using a collection of of 30 texts, written by 10 authors, consisting of 2 , 615 , 856 words and 99 , 903 sentences, confirms that mining part-of-speech skip-grams in texts facilitates authorship inference.