Function Words in Authorship Attribution Studies

Function Words in Authorship Attribution Studies
复制标题

作者归属研究中的虚词

DOI:
10.1093/llc/fql048
复制
发表时间:
2007
期刊:
Lit. Linguistic Comput.
影响因子:
--
通讯作者:
Javier Calle Martín
Javier Calle Martín
中科院分区:
--
文献类型:
--
作者:
A. M. García;Javier Calle Martín

文献摘要

被引文献

相似文献

在过去的几十年里,寻找一个可靠的表达来衡量作者的词汇丰富程度已经成为许多统计学家试图解决一些有争议的作者归属问题的圣杯。最大的努力已经致力于找到一个公式,该公式基于对标记、词类型、最频繁词、hapax legomena、hapax dislegomena等的计算,这样它就能成功地描述一个文本,而与其长度无关。在这条线上,Yule的K和Zipf的Z似乎被学者们普遍接受为通过计算内容和功能词来衡量词汇重复和词汇丰富度的可靠措施。1考虑到后者的更高频率,当在p.c.a.中孤立计算时,它们被证明是更可靠的标识符。和基于Delta的归因研究,他们对前者的比率也衡量了文本的功能密度。在本文中,我们的目标是表明,每个常数用于测量一个特定的功能,因此,他们被认为是相辅相成的,因为一个所谓的丰富的文本(在其引理)不一定要以其低功能密度为特征,反之亦然。为此,西撒克逊福音(WSG)和阿波罗尼乌斯的轮胎(AoT)的注释语料库已被使用沿着与一个巨大的原始语料库。
The search for a reliable expression to measure an author's lexical richness has constituted many statisticians' holy grail over the last decades in their attempt to solve some controversial authorship attributions. The greatest effort has been devoted to find a formula grounded on the computation of tokens, word-types, most-frequent-word(s), hapax legomena, hapax dislegomena, etc., such that it would characterize a text successfully, independent of its length. In this line, Yule's K and Zipf 's Z seem to be generally accepted by scholars as reliable measures of lexical repetition and lexical richness by computing content and function words altogether. 1 Given the latter's higher frequency, they prove to be more reliable identifiers when isolatedly computed in p.c.a. and Delta-based attribution studies, and their rate to the former also measures the functional density of a text. In this paper, we aim to show that each constant serves to measure a specific feature and, as such, they are thought to complement one another since a supposedly rich text (in terms of its lemmas) does necessarily have to characterize by its low functional density, and vice versa. For this purpose, an annotated corpus of the West Saxon Gospels (WSG) and Apollonius of Tyre (AoT) has been used along with a huge raw corpus.