Dealing with heterogeneous big data when geoparsing historical corpora

Dealing with heterogeneous big data when geoparsing historical corpora
复制标题

地理解析历史语料库时处理异构大数据

DOI:
10.1109/bigdata.2014.7004457
复制
发表时间:
2014
期刊:
2014 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Daniel Hartmann
Daniel Hartmann
中科院分区:
--
文献类型:
--
作者:
C. J. Rupp;Paul Rayson;I. Gregory;A. Hardie;Amelia Joulain;Daniel Hartmann

文献摘要

被引文献

相似文献

人们早就知道,“多样性”是海量数据的主要挑战和机遇之一。当我们考虑到大规模数字化活动所产生的历史语料库中的各种内容时,这一点尤其正确。早期英语图书在线(EEBO)和大英图书馆19世纪世纪报纸等馆藏是非常大的异构数据源,包含时间,地点,主题,风格和质量方面的各种内容。这些语料库中引用的地理位置的范围对最先进的地理解析工具提出了严峻的挑战。在我们的工作的背景下,空间人文分析,我们提出了我们的解决方案,处理这些语料库的品种和规模。
It has long been known that `variety' is one of the key challenges and opportunities of big data. This is especially true when we consider the variety of content in historical corpora resulting from large-scale digitisation activities. Collections such as Early English Books Online (EEBO) and the British Library 19th Century Newspapers are extremely large and heterogeneous data sources containing a variety of content in terms of time, location, topic, style and quality. The range of geographical locations referenced in these corpora poses a difficult challenge for state of the art geoparsing tools. In the context of our work on Spatial Humanities analyses, we present our solution for dealing with the variety and scale of these corpora.