Domain Independent Authorship Attribution without Domain Adaptation

Domain Independent Authorship Attribution without Domain Adaptation
复制标题

DOI:
--
复制
发表时间:
2011-09
期刊:
--
影响因子:
--
通讯作者:
R. Menon;Yejin Choi
R. Menon;Yejin Choi
中科院分区:
其他
文献类型:
--
作者:
R. Menon;Yejin Choi

文献摘要

被引文献

相似文献

自动作者归属,就其性质而言,如果是域(即,主题和/或体裁)独立。也就是说,需要作者归属的许多真实的世界问题可能不具有容易获得的域内训练数据。然而,大多数基于机器学习技术的先前工作仅关注域内文本的作者归属。在本文中,我们提出了综合评价的各种文体技术的跨域作者归属。从实验的基础上项目古滕贝格图书档案,我们发现,非常简单的技术的基础上停用词是令人惊讶的鲁棒性对域的变化,基本上摆脱了需要域适应时,提供了大量的数据。
Automatic authorship attribution, by its nature, is much more advantageous if it is domain (i.e., topic and/or genre) independent. That is, many real world problems that require authorship attribution may not have in-domain training data readily available. However, most previous work based on machine learning techniques focused only on in-domain text for authorship attribution. In this paper, we present comprehensive evaluation of various stylometric techniques for cross-domain authorship attribution. From the experiments based on the Project Gutenberg book archive, we discover that extremely simple techniques based on stopwords are surprisingly robust against domain change, essentially ridding the need for domain adaptation when supplied with a large amount of data.