N-GRAM-BASED AUTHOR PROFILES FOR AUTHORSHIP ATTRIBUTION

N-GRAM-BASED AUTHOR PROFILES FOR AUTHORSHIP ATTRIBUTION
复制标题

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
Vlado Ke;Fuchun Peng;N. Cercone;Calvin Thomas
Vlado Ke;Fuchun Peng;N. Cercone;Calvin Thomas
中科院分区:
其他
文献类型:
--
作者:
Vlado Ke;Fuchun Peng;N. Cercone;Calvin Thomas

文献摘要

被引文献

相似文献

我们提出了一种基于字符级 n-gram 作者概况的计算机辅助作者归属的新颖方法,该方法的灵感来自 1976 年几乎被遗忘的开创性方法。现有的自动作者归属方法隐式地将作者概况构建为特征权重向量、语言模型或类似方法。我们的方法基于字节级 n-gram,它与语言无关,并且生成的作者简介的大小有限。在英语、希腊语和中文数据上进行的实验证明了该方法的有效性和语言独立性。结果的准确性达到当前最先进方法的水平,在某些情况下甚至更高。
We present a novel method for computer-assisted authorship attribution based on characterlevel n-gram author proles, which is motivated by an almost-forgotten, pioneering method in 1976. The existing approaches to automated authorship attribution implicitly build author proles as vectors of feature weights, as language models, or similar. Our approach is based on byte-level n-grams, it is language independent, and the generated author proles are limited in size. The eectiveness of the approach and language independence are demonstrated in experiments performed on English, Greek, and Chinese data. The accuracy of the results is at the level of the current state of the art approaches or higher in some cases.