Modeling the scholars: Detecting intertextuality through enhanced word-level n-gram matching
Modeling the scholars: Detecting intertextuality through enhanced word-level n-gram matching
复制标题
学者建模:通过增强的词级 n 元语法匹配检测互文性
DOI:
10.1093/llc/fqu014
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Sarah L. Jacobson
中科院分区:
文献类型:
--
作者:
Christopher W. Forstall;Neil Coffee;Thomas Buck;Katherine Roache;Sarah L. Jacobson
The study of intertextuality, or how authors make artistic use of other texts in their works, has a long tradition, and has in recent years benefited from a variety of applications of digital methods. This article describes an approach for detecting the sorts of intertexts that literary scholars have found most meaningful, as embodied in the free Tesserae website . Tests of Tesserae Versions 1 and 2 showed that word-level n-gram matching could recall a majority of parallels identified by scholarly commentators in a benchmark set. But these versions lacked precision, so that the meaningful parallels could be found only among long lists of those that were not meaningful. The Version 3 search described here adds a second stage scoring system that sorts the found parallels by a formula accounting for word frequency and phrase density. Testing against a benchmark set of intertexts in Latin epic poetry shows that the scoring system overall succeeds in ranking parallels of greater significance more highly, allowing site users to find meaningful parallels more quickly. Users can also choose to adjust both recall and precision by focusing only on results above given score levels. As a theoretical matter, these tests establish that lemma identity, word frequency, and phrase density are important constituents of what make a phrase parallel a meaningful intertext.