Words, Patterns and Documents: Experiments in Machine Learning and Text Analysis
Words, Patterns and Documents: Experiments in Machine Learning and Text Analysis
复制标题
单词、模式和文档:机器学习和文本分析实验
DOI:
--
复制
发表时间:
2009
影响因子:
0.4
通讯作者:
Mark Olsen
中科院分区:
文献类型:
--
作者:
S. Argamon;Mark Olsen
This introduces the set of papers reflecting initial collaborative work between the ARTFL Project at the University of Chicago and the Linguistic Cognition Laboratory at the Illinois Institute of Technology on the intersection of machine learning, text mining and text analysis. One of the emerging grand challenges for digital humanities in the next decade is to address rapidly expanding repositories of electronic text. A number of efforts, such as Google Book Search and the Bibliothèque numérique européenne, are digitizing the holdings of many of the world's great research libraries. The resulting collections will contain nothing less than, in Gregory Crane's view, the “stored record of humanity” [Crane 2006]. This expansion beyond existing digital collections will be one of at least a couple of orders of magnitude, and will introduce a variety of new problems beyond simply scale, including heterogeneity of content and granularity of objects. The problems posed by the emerging global digital library offer opportunities for collaborative work between scholars in the humanities and computer scientists in many domains, from optical character and page recognition to historical ontologies [Argamon and Olsen 2006]. The papers presented here reflect initial collaborative work between the ARTFL Project at the University of Chicago and the Linguistic Cognition Laboratory at the Illinois Institute of Technology on one subset of the technologies required for a future global digital library: the intersection of machine learning, text mining and text analysis. Traditional models of text analysis in digital humanities have concentrated on searching for a relatively small number of words and reporting results in formats long familiar to humanities scholars, most notably concordances, collocation tables, and word frequency breakdowns. [1] While effective for many types of questions, this approach will not scale effectively beyond collections of a relatively modest size, as result sets for even uncommon groups of words will balloon to a size not readily digestible by humans. Furthermore, this approach does not lend itself to abstract discussions of entire works