Modeling the scholars: Detecting intertextuality through enhanced word-level n-gram matching

Modeling the scholars: Detecting intertextuality through enhanced word-level n-gram matching
复制标题

学者建模:通过增强的词级 n 元语法匹配检测互文性

DOI:
10.1093/llc/fqu014
复制
发表时间:
2015
期刊:
Digit. Scholarsh. Humanit.
影响因子:
--
通讯作者:
Sarah L. Jacobson
Sarah L. Jacobson
中科院分区:
--
文献类型:
--
作者:
Christopher W. Forstall;Neil Coffee;Thomas Buck;Katherine Roache;Sarah L. Jacobson

文献摘要

被引文献

相似文献

互文性研究,即作者如何在作品中艺术地利用其他文本的研究,有着悠久的传统,近年来得益于数字方法的各种应用。本文描述了一种方法来检测的互文本,文学学者发现最有意义的,体现在免费的Tesserae网站。Tesserae版本1和2的测试表明,词级n-gram匹配可以回忆起学术评论员在基准集中识别的大部分相似之处。但这些版本缺乏精确性,因此只有在一长串没有意义的相似之处中才能找到有意义的相似之处。这里描述的版本3搜索增加了第二阶段评分系统,该系统通过考虑词频和短语密度的公式对找到的相似项进行排序。对拉丁史诗互文的一组基准测试表明,评分系统总体上成功地将更重要的相似之处排在更高的位置,使网站用户能够更快地找到有意义的相似之处。用户还可以选择只关注高于给定分数水平的结果来调整召回率和精确率。作为一个理论问题,这些测试建立了词元身份,词频和短语密度是什么使一个短语平行有意义的互文的重要组成部分。
The study of intertextuality, or how authors make artistic use of other texts in their works, has a long tradition, and has in recent years benefited from a variety of applications of digital methods. This article describes an approach for detecting the sorts of intertexts that literary scholars have found most meaningful, as embodied in the free Tesserae website . Tests of Tesserae Versions 1 and 2 showed that word-level n-gram matching could recall a majority of parallels identified by scholarly commentators in a benchmark set. But these versions lacked precision, so that the meaningful parallels could be found only among long lists of those that were not meaningful. The Version 3 search described here adds a second stage scoring system that sorts the found parallels by a formula accounting for word frequency and phrase density. Testing against a benchmark set of intertexts in Latin epic poetry shows that the scoring system overall succeeds in ranking parallels of greater significance more highly, allowing site users to find meaningful parallels more quickly. Users can also choose to adjust both recall and precision by focusing only on results above given score levels. As a theoretical matter, these tests establish that lemma identity, word frequency, and phrase density are important constituents of what make a phrase parallel a meaningful intertext.