Words, Patterns and Documents: Experiments in Machine Learning and Text Analysis

Words, Patterns and Documents: Experiments in Machine Learning and Text Analysis
复制标题

单词、模式和文档:机器学习和文本分析实验

DOI:
--
复制
发表时间:
2009
影响因子:
0.4
通讯作者:
Mark Olsen
Mark Olsen
中科院分区:
--
文献类型:
--
作者:
S. Argamon;Mark Olsen

文献摘要

被引文献

相似文献

这介绍了一套文件,反映了在芝加哥大学的ARTFL项目和语言认知实验室在伊利诺伊理工学院的机器学习,文本挖掘和文本分析的交叉点之间的初步合作工作。数字人文学科在未来十年中面临的重大挑战之一是解决迅速扩大的电子文本库。谷歌图书搜索和欧洲数字图书馆等许多努力正在将世界上许多伟大的研究图书馆的馆藏数字化。在格雷戈里·起重机看来,由此产生的收藏品将包含“人类的储存记录”[起重机2006]。这种超越现有数字收藏的扩展将是至少几个数量级的扩展之一,并将引入各种新的问题,而不仅仅是规模,包括内容的异质性和对象的粒度。新兴的全球数字图书馆带来的问题为人文学者和计算机科学家在许多领域的合作提供了机会,从光学字符和页面识别到历史本体[Argamon和Olsen 2006]。这里介绍的论文反映了芝加哥大学的ARTFL项目和伊利诺伊理工学院的语言认知实验室在未来全球数字图书馆所需技术的一个子集上的初步合作工作:机器学习,文本挖掘和文本分析的交叉。传统的数字人文文本分析模型集中在搜索相对较少的单词,并以人文学者长期熟悉的格式报告结果,最显着的是索引,搭配表和词频分析。[1]虽然对许多类型的问题有效,但这种方法不会有效地扩展到相对适度大小的集合之外,因为即使是不常见的单词组的结果集也会膨胀到人类不易消化的大小。此外,这种方法不适合于对整个作品的抽象讨论
This introduces the set of papers reflecting initial collaborative work between the ARTFL Project at the University of Chicago and the Linguistic Cognition Laboratory at the Illinois Institute of Technology on the intersection of machine learning, text mining and text analysis. One of the emerging grand challenges for digital humanities in the next decade is to address rapidly expanding repositories of electronic text. A number of efforts, such as Google Book Search and the Bibliothèque numérique européenne, are digitizing the holdings of many of the world's great research libraries. The resulting collections will contain nothing less than, in Gregory Crane's view, the “stored record of humanity” [Crane 2006]. This expansion beyond existing digital collections will be one of at least a couple of orders of magnitude, and will introduce a variety of new problems beyond simply scale, including heterogeneity of content and granularity of objects. The problems posed by the emerging global digital library offer opportunities for collaborative work between scholars in the humanities and computer scientists in many domains, from optical character and page recognition to historical ontologies [Argamon and Olsen 2006]. The papers presented here reflect initial collaborative work between the ARTFL Project at the University of Chicago and the Linguistic Cognition Laboratory at the Illinois Institute of Technology on one subset of the technologies required for a future global digital library: the intersection of machine learning, text mining and text analysis. Traditional models of text analysis in digital humanities have concentrated on searching for a relatively small number of words and reporting results in formats long familiar to humanities scholars, most notably concordances, collocation tables, and word frequency breakdowns. [1] While effective for many types of questions, this approach will not scale effectively beyond collections of a relatively modest size, as result sets for even uncommon groups of words will balloon to a size not readily digestible by humans. Furthermore, this approach does not lend itself to abstract discussions of entire works