Dealing with heterogeneous big data when geoparsing historical corpora
Dealing with heterogeneous big data when geoparsing historical corpora
复制标题
地理解析历史语料库时处理异构大数据
DOI:
10.1109/bigdata.2014.7004457
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Daniel Hartmann
中科院分区:
文献类型:
--
作者:
C. J. Rupp;Paul Rayson;I. Gregory;A. Hardie;Amelia Joulain;Daniel Hartmann
It has long been known that `variety' is one of the key challenges and opportunities of big data. This is especially true when we consider the variety of content in historical corpora resulting from large-scale digitisation activities. Collections such as Early English Books Online (EEBO) and the British Library 19th Century Newspapers are extremely large and heterogeneous data sources containing a variety of content in terms of time, location, topic, style and quality. The range of geographical locations referenced in these corpora poses a difficult challenge for state of the art geoparsing tools. In the context of our work on Spatial Humanities analyses, we present our solution for dealing with the variety and scale of these corpora.