The Remaking of Reading : Data Mining and the Digital Humanities

The Remaking of Reading : Data Mining and the Digital Humanities
复制标题

阅读的重塑:数据挖掘和数字人文

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
M. Kirschenbaum
M. Kirschenbaum
中科院分区:
--
文献类型:
--
作者:
M. Kirschenbaum

文献摘要

被引文献

相似文献

本文讨论了数据挖掘在文学批评这一看似不可能的领域中的应用。虽然底层技术是传统的朴素贝叶斯、支持向量机文学批评和更普遍的“数字人文学科”,但与其他领域的不同之处在于,它们很少在讨论中承认基本事实。相反,数据挖掘和机器学习最好从“挑衅”的角度来理解--异常结果可能会让读者感到惊讶,从而注意到文本中以前认为不重要的某些方面--以及“不阅读”或“远距离阅读”,即在比传统人文主义的“近距离阅读”方法更广泛的语料库中自动搜索模式。当美国国家艺术基金会(National Endowment for the Arts)的一份广为宣传的报告得出阅读本身“处于危险之中”的结论时,大型在线文本收藏(谷歌图书,开放内容联盟)正在以机器可读的形式提供数百万文本。数据挖掘是阅读改造的一部分。
This paper discusses applications of data mining in the seemingly unlikely field of literary criticism. While the underlying techniques are traditional—Naïve Bayes, SVM—literary criticism, and the “digital humanities” more generally, differ from other domains in that they rarely admit ground truth into their discussions. Instead, data mining and machine learning are best understood in terms of “provocation”—the potential for outlier results to surprise a reader into attending to some aspect of a text not previously deemed significant—as well as “notreading” or “distant reading,” the automated search for patterns across a much wider corpus than could be read and assimilated via traditional humanistic methods of “close reading.” At a moment when a widely publicized report by the National Endowment for the Arts concluded reading itself was “at risk,” large online text collections (Google Books, the Open Content Alliance) are making millions of texts available in machine-readable form. Data mining is part of this remaking of reading.