Statistical Stylometrics and the Marlowe-Shakespeare Authorship Debate

Statistical Stylometrics and the Marlowe-Shakespeare Authorship Debate
复制标题

统计风格计量学和马洛-莎士比亚作者之争

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Omran Ehmoda
Omran Ehmoda
中科院分区:
--
文献类型:
--
作者:
Neal P. Fox;Omran Ehmoda

文献摘要

被引文献

相似文献

I.历史学家、文学学者、心理学家和最近的计算语言学家一直在寻找一种可靠的方法来分析文本以确定其作者的身份。至少从世纪后期开始(见门登霍尔,1887;马斯科尔,1888 a/B),研究作者身份的一种工具就是从文献中提取统计趋势,并对这些数据进行比较,以便对文献进行适当的分组。寻求这样一种方法的基础是一个关键的假设,即一个作者使用书面语言所固有的一些统计上可量化的特征或一组特征可以被孤立出来,这些特征在该作者的作品中是一致的,但在不同的作者之间是不同的。因此,这些特征可以被用作一种指纹,以区分不同作者的作品,并识别匿名出版作品的可能作者。这种奋进通常被称为stylometry。随着近几十年来自然语言处理和机器学习技术的出现,该领域取得了很大的进步,许多研究人员已经解决了许多具体的问题。
I. Authorship Attribution Paradigm Historians, literary scholars, psychologists, and – more recently – computational linguists have long sought a reliable methodology for analyzing texts to determine the identity of their author. Since at least the late nineteenth century (see Mendenhall, 1887; Mascol, 1888a/b), one tool used in the investigation of authorship has been the extraction of statistical tendencies from the documents and comparison of these data in order to group the documents appropriately. Underlying the search for such a methodology is the critical assumption that some statistically quantifiable characteristic or set of characteristics inherent in a single author’s use of written language could be isolated that would be consistent across works by that author, but differ between different authors. Thus, the feature(s) could be used as a sort of fingerprint to distinguish between works by distinct authors and to identify the likely author of an anonymously published work. This endeavor is generally referred to as stylometry. With the advent of natural language processing and machine learning techniques in recent decades, the field has advanced greatly, and numerous researchers have addressed many specific